About this run
Run
- Id
- d7442a5f-fd3b-438a-8816-91f4625f2492
- Prompt
- Incident management process
- Configuration
- study 5x3
- Loaded from
- slow-thinker.keys.json ← slow-thinker.base.json ← slow-thinker.study.json
- Made with
- Process executor 0.1.5
- This page
- Report generator 0.1.16 · Analysis generator 0.1.3
- Status
- completed
- Started
- 2026-09-21 02:48
- Duration
- 18 min 25 s
Council
- Proposers
- 5
- Refinement rounds
- 3
- Voters
- 5
- Selected plan
- opus5_refine_1 · 4 of 5 votes · 35 steps
Plan cost
Analysis LLM analysis
Models of the council
| Model | Thinking |
|---|---|
| deepseek-v4-pro · deepseek/deepseek-v4-pro | thinking on, effort high |
| gpt5.6-sol · openai/gpt-5.6-sol | reasoning effort high |
| grok4.6 · xai/grok-4.6 | reasoning effort high |
| opus5 · anthropic/claude-opus-5 | adaptive thinking, effort high |
| qwen3.8-max · alibaba/qwen3.8-max | thinking on, budget 16.0k tokens |
Analyses of this run
| # | Date | Analyst | Schema | Analysis generator | Calls | Cost | Verdict on the vote | Report |
|---|---|---|---|---|---|---|---|---|
| 1 | 2026-09-21 03:21 | judge · anthropic/claude-opus-5 |
v5 | 0.1.3 slow-thinker 0.1.0 |
8 | 6.49 USD | agrees | this page |
Task
A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit.
Process
Round 0 — initial proposals
All five agents converge on the same skeleton — exec mandate, baseline of the 31 incidents, a 4–5 level severity scale, separated IC/comms/scribe/SME roles, paid on-call, one paging tool, synthetic payment probes, timed status-page updates, mandatory blameless postmortems with tracked actions, pilot-then-waves, SOC 2 evidence as a by-product — so they differ mainly in staffing design, pace, and operational depth. P1 and P4 are the most granular (29–30 steps), P2 is the most payments-literate and the only one with day-1 interim controls, P3 is mid-weight with some internal inconsistencies, and P5 is the thinnest and least dated.
What the proposals share
- Command separated from debugging: every plan defines IC / comms lead / scribe / SMEs, forbids the IC from typing, allows anyone to declare, and lets only the IC downgrade (P1 S6, P2 S5, P3 S4, P4 S7, P5 S3).
- Pay for on-call and page only for owned code as the explicit answer to the pushback, with a service-ownership catalog and criticality tiers as the routing foundation (P1 S4/S9, P2 S3/S6, P3 S4/S5, P4 S5/S18, P5 S4/S5).
- Alert-quality gate plus customer-journey detection: every page needs owner, runbook, severity and SLO linkage; synthetic end-to-end payment probes from both regions; consolidation of the six tools into one pager (P1 S10/S11/S12, P2 S7/S8, P3 S6/S9, P4 S13/S16/S20, P5 S5/S6/S10).
- Mandatory postmortems with enforced action tracking: SEV1/SEV2 always, 3–5 day drafts, single template, named individual owners, due dates in one tracker, escalation of overdue items, and SOC 2 evidence produced from the live process rather than reconstructed.
What separates them
- 24x7 staffing design is the biggest split: P1 S7/S8 uses a 25–35-person Duty IC pool plus tiered team obligations; P2 S5 collapses 28 teams into 8–12 domain rotations of ≥6; P3 S4 adds an SRE Tier 2 that absorbs infra pages and floats a managed overnight triage vendor; P4 S8 puts Layer C teams on business hours only with an EM night list; P5 S4 invokes "follow-the-sun between the two AWS regions", which does not create time-zone coverage.
- Pace and interim cover: P2 S2 installs interim IC staffing, one declaration path and a provisional severity guide within 7 days and triages the 53 open actions immediately; P4 S1 targets live in ~10 weeks and S30 freezes the process before the observation window; P1 runs design 1–6, pilot 7–12, rollout 13–24; P3 ends rollout at week 24 with a month-5 mock audit; P5 gives almost no dates and pilots only at step 11 of 14.
- Payments/ledger rigour: P2 S10 guards against split-brain, duplication and replay, requires reconciliation and backlog processing before "resolved", and separates mitigation from resolution; P1 adds ledger double-entry assurance (S11), a regulator matrix incl. NYDFS/FinCEN/sponsor banks (S16) and an automated SLA-credit workflow (S17); P4 S4 extracts notification clocks from contracts before setting severity triggers; P5 S7 defers regulator timing entirely to legal with no clock.
- How failure and resistance are handled: P1 S28 is a standing risk register with pre-committed fallbacks (volunteer shortfall, comp not approved, pruning a real alert); P4 S26 stops wave expansion if the 60/120-day pulse is red; P2 S17 requires exception records with owners and expiry dates; P3 and P5 have no comparable contingency mechanism.
The calls of this round
29-step program run by a Head of Reliability with exec steering, starting from a forensic re-coding of all 31 incidents and the six-tool alert estate. Builds a service catalog with Tier 0–3, a five-level severity scale with auto-escalation on any ledger-touching incident, a 25–35-person Duty IC rotation plus tiered team on-call, and an on-call readiness bar that blocks a service from paging anyone until runbooks exist. Unusually complete on money and law: stipend ranges, FLSA/NY review, a regulator playbook, and an automated SLA-credit workflow targeting $1.3M → <$400k.
Ends with a month-6 dry-run audit over 15 sampled incidents and a year-two maturity roadmap.
The plan produced
1. Program charter, executive mandate and funding
Convert the CEO's frustration into a named program with one accountable owner, a budget and a deadline that is earlier than the audit.
The charter must state that incident management is a company-level operating process, not a per-team choice. Without that, the 28 teams will opt out.
- Appoint a single Incident Management Program Lead (Head of Reliability/SRE) with direct exec sponsorship from CTO and CEO.
- Form a steering group: CTO, VP Eng, Head of Support/CS, CISO/Compliance, Legal, Finance (SLA credits), HR (on-call pay).
- Set the non-negotiables: one severity scale, one paging tool, one postmortem format, mandatory action tracking, paid on-call.
- Fix the timeline: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready by week 28 (four weeks of buffer before the audit).
- Approve budget lines: tooling (~$150–250k/yr), on-call compensation (~$600k–1M/yr), 2–3 dedicated program FTEs. Anchor it against $1.3M of credits plus incident cost.
2. Forensic baseline of the 31 incidents and the alert estate (after 1)
Before designing anything, rebuild the facts. Re-open all 31 incidents and profile the 3,400 monthly alerts so every later design decision is evidence-based.
This also creates the "before" picture the exec and the auditor will compare against.
- Re-code each incident: trigger, service, detection source (customer vs monitor), timestamps for detect/acknowledge/declare/mitigate/resolve, who led, credits paid, root cause family.
- Quantify the 40% customer-first detections: which signal was missing in each case.
- Classify the two "nobody in charge" incidents minute by minute; use them as the burning-platform story.
- Audit the six alerting tools: volume per tool, per team, per alert rule; identify the top 50 rules that produce most of the 85% noise; find rules with no owner and no runbook.
- Baseline the numbers formally: MTTD 22 min, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed. Freeze them as the reference line.
3. Stakeholder listening tour and resistance map (after 1)
Engineer pushback against "carrying a pager for other teams' code" is the main delivery risk. Treat it as a design input, not an attitude problem.
Run structured interviews across all 28 teams, plus Support, CS and Sales, in two weeks.
- Test the real objection: is it unpaid work, night sleep, unfamiliar code, poor runbooks, or fear of blame? Each has a different fix.
- Collect current informal practices — the 12 teams already on-call are the pilot candidates and the source of veterans.
- Document the promise that answers the objection: you are only paged for services your team owns, plus a trained commander who runs the incident and pulls in others.
- Map influencers and blockers by team; recruit 10–15 credible engineers as a design working group so the process is co-authored, not imposed.
- Survey baseline sentiment (trust in alerts, willingness to be on-call, burnout) to re-measure at 6 and 12 months.
4. Service ownership catalog and criticality tiering (after 2)
You cannot page the right person across 180 services until each service has a named owning team. This is the foundation of both on-call fairness and severity mapping.
Build a machine-readable catalog (Backstage or equivalent) that is the single source of truth for routing.
- One owning team per service, a named engineering manager, a Slack channel, a paging escalation policy, a dependency list.
- Tier services by business impact: Tier 0 (money movement, ledger, auth, shared PostgreSQL cluster), Tier 1 (customer-facing but degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map each Tier 0/1 service to the customer-visible capability it supports (payment initiation, settlement, reporting, onboarding).
- Flag orphan services and cross-team shared components; force an ownership decision for each within 30 days, or schedule decommissioning.
- Publish coverage gaps to the steering group: any Tier 0 service without an owner is an executive escalation.
5. Severity scale and declaration criteria (after 2, 4)
Define a five-level scale with objective, payments-specific triggers so declaration is a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade.
Each level triggers a fixed bundle of response, comms and postmortem obligations.
- SEV1: money movement stopped or incorrect, ledger integrity in doubt, data breach, full region loss, >10% of customers impacted. Triggers: immediate 24x7 page of IC + comms + exec, bridge within 5 min, status page within 15 min, mandatory postmortem, regulator assessment.
- SEV2: severe degradation, settlement at risk of missing a window, single large/strategic customer fully down, SLA breach likely. Triggers: IC paged, status page within 30 min, mandatory postmortem.
- SEV3: partial or workaround-available degradation, no credit exposure. Team-led, business-hours comms, postmortem optional but encouraged.
- SEV4/5: minor or internal-only; ticket-tracked, no paging.
- Add auto-escalation rules: any SEV3 open >2 h, or any incident touching the shared ledger cluster, becomes SEV2 automatically. Include a severity decision tree and 12 worked examples drawn from the 31 real incidents.
6. Incident roles, decision authority and handover rules (after 5)
Solve the "nobody in charge for an hour" failure by making command explicit, transferable and logged.
Define five roles with written responsibilities, entry criteria and explicit authority.
- Incident Commander: owns the incident, not the fix. Authority to declare severity, pull any engineer, approve customer-impacting mitigations, invoke failover and authorise spend. The IC never types in the terminal.
- Communications Lead: owns status page, internal updates, account-manager briefings and the exec summary. Single voice to customers.
- Scribe: maintains the timeline, decisions and open questions; feeds the postmortem and the audit evidence trail.
- Subject-Matter Responders: engineers from owning teams; they investigate and remediate, and report to the IC.
- Executive Liaison (SEV1 only): shields the IC from exec questions and owns regulator/board escalation.
- Rules: the IC role is assumed within 5 minutes of declaration, stated explicitly in the channel ("I am IC"), and any handover is announced and logged. Roles may be combined below SEV2; never at SEV1.
7. 24x7 incident command coverage model (after 4, 6)
Command is staffed by a small, trained, cross-team pool — not by 28 teams individually. This is what makes 24x7 realistic in one New York time zone.
Design a central rotation that scales with a growing certified pool.
- Create a Duty Incident Commander rotation of 25–35 certified volunteers (target ~1 per team, plus managers and senior engineers), giving each person roughly one week per 6–8 months.
- Pair with a Duty Comms Lead rotation (Support/CS leads plus engineering managers, ~15–20 people) and a Scribe pool (rotating, lowest barrier, used as the training entry point).
- Coverage: primary + secondary IC at all times; hard 5-minute acknowledgement SLA with automatic failover to secondary, then to the on-call engineering director.
- Night coverage options to evaluate in writing: US-only rotation with paid night stipend now; a Lisbon/Dublin or APAC follow-the-sun cell as a 12-month option; a 24x7 NOC-style triage desk for first-line detection.
- Eligibility: certification required (S21); commanders are volunteers with manager approval and can step out with 30 days' notice.
8. Team on-call structure, rotations and routing rules (after 4, 6)
Rebuild team on-call around the principle that answers the pushback: you are only paged for code your team owns.
Apply a tiered obligation so 28 teams are not treated identically.
- Tier 0/1 owning teams (expected 14–18 teams): 24x7 primary + secondary, minimum 6 people per rotation, one-week shifts, handover Wednesday mornings.
- Tier 2/3 teams: business-hours on-call with a best-effort out-of-hours escalation path, no night paging.
- Platform/Infrastructure and Database teams: 24x7, since they own the shared PostgreSQL ledger cluster and the Kubernetes/regional layer.
- Rotations under 6 people are merged across teams or backfilled by hiring; no rotation of fewer than 4 is approved.
- Routing: every page resolves through the service catalog to the owning team's escalation policy; cross-team pages are made by the IC, never by an alert.
- Guardrails: maximum one week in four, no on-call in the week after a SEV1 you led, protected recovery time after any night page, and a per-person page budget (see S10).
9. On-call compensation, labour compliance and fairness policy (after 3, 8)
Unpaid on-call is both a retention risk and a legal exposure in New York. Paying for it is the fastest way to convert resistance into participation.
Design the scheme with HR, Legal, Finance and Payroll, and publish it before asking anyone to sign up.
- Base stipend per week on rotation, differentiated by tier: e.g. $800–1,200 for 24x7 Tier 0/1, $300–500 for business-hours rotations, with premiums for holidays and weekends.
- Per-incident payment for out-of-hours activation (e.g. $150 per night page plus hourly beyond one hour) and guaranteed time-off-in-lieu after night work.
- Separate Duty IC stipend, since command is a distinct and heavier burden.
- Verify FLSA exempt/non-exempt treatment, NY State wage rules and overtime exposure for non-exempt staff; document the legal review.
- Budget and model the annual cost; get board/CFO approval as a line item, benchmarked against $1.3M of credits.
- Add non-cash elements: on-call time counted as delivery load (teams reduce sprint commitment by ~15%), incident leadership recognised in promotion criteria, and a public quarterly report of on-call load per team.
10. Alert quality standard and page budget (after 2, 4)
3,400 alerts a month at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a contractual condition of paging someone.
Publish a standard, then enforce it mechanically.
- Every paging alert must have: a named owning team, a documented customer impact, a runbook link, a tested threshold, and a severity mapping. Alerts failing this are demoted to ticket or deleted.
- Page only on symptoms that affect customers (SLO burn rate, error budget, queue depth against settlement deadlines); cause-based CPU/memory alerts become dashboards or tickets.
- Set a page budget: maximum 2 out-of-hours pages per person per week. Breach triggers a mandatory alert-tuning sprint for the owning team and blocks new alert creation.
- Auto-quarantine: any alert that fires more than 5 times a month without action, or has >70% no-action acknowledgements, is silenced automatically and returned to its owner.
- Monthly alert review per team: kill, tune, or keep, with the numbers on screen.
- Target: 3,400 → under 500 pages a month, with actionability above 75% within six months.
11. Detection uplift: SLOs, synthetic journeys and ledger assurance (after 4, 10)
The goal is to stop customers telling you first. Detection must be driven by customer-visible outcomes, not host metrics.
Instrument the money path end to end and alert on it.
- Define SLOs for each Tier 0/1 customer capability: payment initiation success rate, authorisation latency, settlement file timeliness, API availability and reporting freshness. Tie them to the 99.95% contractual SLA with a stricter internal target.
- Deploy synthetic transactions from outside the platform, in both regions, every 60 seconds, covering the full payment lifecycle including a small real-value canary flow where feasible.
- Add ledger assurance checks: continuous double-entry balance reconciliation, replication lag and failover-readiness alarms on the shared PostgreSQL cluster, and settlement-window countdown alerts.
- Build per-customer anomaly detection for the top 100 accounts (volume drop, error spike) so a single-tenant outage is detected before the account manager calls.
- Create an "inbound signal" bridge: any support ticket or account-manager report matching impact keywords auto-creates a triage incident within 5 minutes.
- Track every incident's detection source; make "customer detected first" a reviewed defect with its own follow-up action.
12. Tool consolidation and incident platform implementation (after 5, 7, 8, 10)
Collapse six alerting tools into one paging and incident platform so there is a single queue, a single timeline and a single audit record.
Run a short, time-boxed selection and migrate within the pilot window.
- Select an integrated stack: paging/on-call scheduling plus an incident management layer (e.g. PagerDuty + incident.io/FireHydrant, or a single vendor) and a hosted status page.
- Implement one-command declaration in Slack (
/incident declare) that creates the channel and bridge, pages the Duty IC, sets severity, opens the timeline and starts the clock. - Migrate all monitoring sources to route into the one platform; decommission direct paging from the legacy six and block new integrations that bypass it.
- Automate the evidence trail: timestamps, role assignments, severity changes, comms sent, and postmortem linkage exported for SOC 2.
- Integrate with the service catalog for routing, Jira for actions, Salesforce/CS tooling for affected-customer lists, and Zoom/Slack Huddle for the bridge.
- Hard requirement: the platform must work when AWS in one region is down — verify out-of-band paging (SMS/phone) and a printed/offline fallback runbook.
13. Detection-to-escalation path and the five-minute command rule (after 6, 7, 12)
Write the single path from "something looks wrong" to "someone is in charge", and make it impossible to skip.
The design target is detection to commander in under five minutes, any hour.
- Entry points: automated alert, engineer observation, support ticket, account manager, partner bank, customer-facing SEV hotline. All converge on the same declaration command.
- Anyone in the company may declare up to SEV2; nobody is punished for over-declaring. Publish that rule in writing and repeat it.
- Auto-page ladder: Duty IC (5 min) → secondary IC (5 min) → on-call Director (10 min) → CTO. Same ladder for the owning team's responder.
- Cross-team pull: the IC can page any team's on-call directly, with a 10-minute acknowledgement obligation. This is the reciprocal commitment that makes single-team ownership viable.
- Explicit takeover protocol: if no one claims IC within 5 minutes, the platform assigns it and announces it; the assignee cannot decline, only hand over.
- Define standing severity triggers for immediate regional failover, ledger read-only mode and partner-bank notification, with pre-authorised decision rights so the IC does not wait for an executive.
14. Internal communications protocol (after 6, 12)
Standardise the internal channel so responders, executives and support see the same picture without interrupting the IC.
Separate the working channel from the audience channel.
- One incident channel per incident (auto-created), one bridge, and a read-only broadcast channel for executives, Support and Sales.
- Update cadence by severity: SEV1 every 30 minutes even if nothing has changed; SEV2 every 60 minutes; SEV3 at state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, ETA or next update time, current IC and Comms Lead.
- Exec briefing rule: executives ask questions only to the Executive Liaison; the IC is not interrupted. Publish this as a behavioural expectation signed by the exec team.
- Support/CS enablement: a live affected-customer list and a holding statement within 15 minutes of SEV1/SEV2 so the front line is never guessing.
- Handover protocol for incidents beyond 4 hours: formal IC handover checklist, fatigue rule, and staffing of a second shift.
15. Customer communications and status page policy (after 5, 14)
Customers currently learn of outages from their own monitoring and hear from whoever happens to be around. Replace that with a timed, owned, pre-approved process.
The Comms Lead is the single author; templates remove the need to write under pressure.
- Timing commitments: status page posted within 15 minutes of SEV1 declaration and 30 minutes for SEV2; updates every 30/60 minutes; resolution notice within 30 minutes of mitigation; customer-facing summary within 5 business days for SEV1.
- Pre-approve 12–15 templates with Legal and Comms (degradation, delay in settlement, API errors, security event, third-party failure) so nothing needs legal review mid-incident.
- Subscription-based status page with per-component granularity mapped to the customer capabilities from S11, plus an email/webhook/RSS feed.
- Tiered outreach: top 100 accounts get a direct named call or email from their account manager within 30 minutes of SEV1, with a briefing pack from the Comms Lead; long tail gets the status page and a proactive email.
- Rules of language: state impact and next update time, never speculate on cause, never assign blame to a vendor before facts are confirmed.
- Run a quarterly customer-perception check with the top accounts on whether comms were timely and useful.
16. Regulatory, partner and legal notification playbook (after 5, 15)
In payments, some incidents are reportable and the clock starts at detection. Build the assessment into the process so it is never an afterthought.
Work with Legal, Compliance and the CISO to produce a decision tree and contact matrix.
- Map obligations: NYDFS Part 500 (72-hour cybersecurity event notification), state breach laws, GLBA/FTC Safeguards, PCI DSS if card data is in scope, sponsor-bank and card-network contractual notice windows, and any FinCEN/OFAC implications.
- Add a mandatory regulatory-assessment checkpoint to every SEV1 and every security-related SEV2, owned by the Executive Liaison, completed within 2 hours of declaration and recorded even when the answer is "not reportable".
- Build the contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel, with 24x7 numbers and named backups.
- Pre-draft notification letters and hold them under legal privilege review.
- Check customer contracts for bespoke notification SLAs (often 1–4 hours for enterprise accounts) and encode them in the customer tiering.
- Test the playbook once per quarter as part of the simulation programme.
17. SLA credit and financial impact workflow (after 5, 15)
Link incidents to money so severity, credits and prioritisation stay consistent — and so Finance stops being surprised.
Make credit calculation an automated output of the incident record, not a negotiation.
- Define the availability measurement method per contract, per component, and agree it with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and the per-capability telemetry; generate a proposed credit schedule within 5 business days of resolution.
- Decide the posture: proactive credits for the top tier (reputation upside) versus claims-based for the rest; document the approval chain.
- Track credits per incident and per root-cause family; feed a quarterly report showing which reliability investments would have prevented which credits.
- Set a target: reduce credits from $1.3M to under $400k in year one, and use that delta as the ongoing business case.
18. Postmortem standard and blameless review forum (after 5, 6)
Replace "some incidents, various formats" with a mandatory, single-format, blameless process with fixed deadlines.
The discipline is in the deadlines and the forum, not in the template.
- Mandatory for: every SEV1 and SEV2, every incident where a customer detected it first, every incident over 2 hours, every repeat of a known cause, and every near-miss involving the ledger. Optional but templated for SEV3.
- Fixed timeline: draft within 3 business days, peer review within 5, published company-wide within 10. The IC owns delivery; the owning team's manager is accountable.
- One template: timeline, customer and financial impact, detection analysis (why not sooner), response analysis (why mitigation took as long as it did), contributing factors, what went well, action items with owner and due date.
- Blameless rules in writing: describe systems and decisions in the context available at the time; no individual named as a cause; HR and management commit that postmortems are never used in performance reviews.
- Weekly 60-minute Incident Review Board: reviews all postmortems from the prior week, challenges quality, ratifies severity, and approves or rejects action items. Attendance by engineering directors is mandatory.
- Publish a searchable postmortem library and a quarterly "top five recurring causes" analysis.
19. Action item ownership and tracking system (after 12, 18)
11 of 64 actions closed is the single clearest symptom of a process nobody enforces. Give actions the same status as customer commitments.
Track them where engineering work already lives, with visible escalation.
- Every action gets: a named individual owner (not a team), a priority class, a due date and a Jira ticket auto-created from the postmortem.
- Priority classes with hard SLAs: P0 prevents recurrence of a SEV1, due in 30 days; P1 in 60 days; P2 in 90 days. P0s are committed into the next sprint before any roadmap work.
- Capacity rule: teams reserve a standing 15–20% of sprint capacity for reliability and incident actions. Without reserved capacity, the actions will not land.
- Escalation ladder for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Monthly reporting of closure rate by team in the engineering leadership review; include it in manager performance objectives.
- Target: 90% of P0/P1 actions closed on time within two quarters.
20. Runbooks, major-incident playbooks and the on-call readiness bar (after 4, 8)
Nobody can respond well to unfamiliar systems at 3 a.m. without runbooks — and poor runbooks are a real part of the pager resistance.
Define a minimum readiness bar that a service must meet before it is allowed to page anyone.
- Readiness checklist per Tier 0/1 service: current architecture diagram, dependency map, dashboard link, alert-to-runbook mapping, rollback procedure, feature-flag kill switches, escalation contacts, and a data-loss/latency impact statement.
- Write major-incident playbooks for the top failure modes derived from S2: shared PostgreSQL ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank/third-party outage, settlement-window breach, and suspected security compromise.
- Prioritise the shared ledger: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and a clearly stated data-loss tolerance (RPO/RTO) signed off by the exec.
- Runbooks must be tested at least twice a year in a drill; untested runbooks are marked stale in the catalog.
- Enforcement: a service without readiness sign-off cannot create paging alerts, and the gap is reported to its director.
21. Training, certification and the commander academy (after 6, 13, 14, 18)
Command is a skill, not a title. Build a certification path so 24x7 coverage is staffed by people who have practised.
Use a tiered curriculum with real assessment.
- Scribe (2 hours): timeline discipline and tooling. The entry point for everyone.
- Responder (half a day): severity scale, declaration, escalation, runbook use, comms hygiene. Mandatory for every engineer joining an on-call rotation.
- Incident Commander (two days plus shadowing): command presence, delegation, decision-making under uncertainty, severity calls, handover, exec management. Requires two shadowed incidents and one simulated SEV1 before certification.
- Communications Lead (one day): status-page writing, customer tiering, legal boundaries, regulator triggers.
- Certification is valid 12 months and renewed via a simulation; the register of certified people is an audit artefact.
- Add on-call onboarding per team: a new joiner shadows two shifts before holding primary, and never holds primary in their first 90 days.
22. Simulation programme: game days, drills and wheel of misfortune (after 12, 20, 21)
The process must be rehearsed before it meets a real SEV1. Simulations also build the commander pool and expose runbook gaps cheaply.
Run a standing calendar rather than one-off exercises.
- Monthly 60-minute tabletop ("wheel of misfortune") per engineering group, using a real past incident from the 31.
- Quarterly full-scale game day in production or a production-like environment: regional failover, ledger replica promotion, dependency failure, with the whole role structure activated and timed.
- Twice-yearly unannounced paging drill to measure real acknowledgement times at night.
- One security-incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Every exercise produces a lightweight postmortem and action items in the same system as real incidents.
- Measure and publish drill metrics: time to IC, time to first status update, time to correct mitigation decision.
23. Pilot with wave 0 teams (after 9, 11, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Run a six-week pilot with tight measurement and a public verdict.
- Select 5–6 teams: core payments, ledger/database, platform/Kubernetes, API gateway, plus two of the 12 teams already on-call.
- Activate the full stack for them: new severity scale, Duty IC rotation, single paging tool, alert budget, status-page policy, mandatory postmortems, paid on-call.
- Hold a weekly pilot retro; expect and document 20–40 process defects, and fix them in the standard before rollout.
- Validate the hard questions: does the 5-minute IC rule hold at 3 a.m.? Do cross-team pulls get answered? Is the severity tree unambiguous?
- Exit criteria: MTTD under 10 minutes for pilot services, IC assigned within 5 minutes in 95% of incidents, page volume down 50%, all postmortems on time, positive on-call sentiment.
- Publish a one-page pilot result to the whole company — this is the main adoption argument for the remaining teams.
24. Metrics, dashboards and the review cadence (after 12, 18, 23)
Instrument the process itself, so improvement is visible and the audit has evidence of monitoring and review.
Define a small set of metrics with owners and a fixed meeting rhythm.
- Response metrics: MTTD, time to declare, time to IC assigned, MTTA, MTTM, MTTR, incidents per month by severity, % detected by customers first.
- Quality metrics: page volume per person per week, alert actionability rate, page budget breaches, postmortem on-time rate, action closure rate and ageing.
- Business metrics: SLA credits paid, availability against 99.95% per capability, error-budget consumption, repeat-incident rate.
- People metrics: on-call load distribution across teams, out-of-hours pages per person, on-call sentiment and attrition among on-call staff.
- Cadence: weekly Incident Review Board (postmortems and actions), monthly Reliability Review (metrics per team, alert hygiene, on-call load), quarterly Executive/Board review (credits, trends, investment asks), annual policy review.
- Every metric gets a target and a named owner; dashboards are self-serve and public inside the company.
25. Wave rollout across all 28 teams with readiness gates (after 23, 24)
Roll out in four waves of six to eight teams, every three weeks, ordered by criticality. Each wave passes an explicit gate rather than a deadline.
Gates keep quality high and make the standard credible.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with 6+ certified responders, one IC candidate nominated, one drill passed.
- Assign each wave a named coach from the program team for three weeks of hands-on support.
- Gate criteria are checked and signed by the director; teams that fail are re-scheduled, not waived.
- Freeze legacy tooling per wave: after onboarding, the old alerting paths are disabled, not left as a fallback.
- Publish a live adoption scoreboard by team so progress is social, not administrative.
26. SOC 2 control mapping, evidence automation and internal dry-run audit (after 18, 19, 25)
Design the process so audit evidence is a by-product of doing the work, then test it before the auditors do.
Engage the auditor early to confirm the interpretation of controls.
- Map the process to the Trust Services Criteria: CC7.3 and CC7.4 (incident identification, response, recovery), CC7.2 (monitoring), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), plus availability criteria A1.2.
- Produce and approve formal policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard — versioned, signed, annually reviewed.
- Automate evidence: incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, postmortem library, action-item closure reports, training and certification register, drill records.
- Confirm the observation window with the auditor and ensure the process is operating for a minimum of three months before fieldwork.
- Run an internal dry-run audit at month six: sample 15 incidents and walk the full evidence chain; fix gaps with 8 weeks to spare.
- Keep a remediation log for any incident where the process was not followed, with the corrective action — auditors respond better to documented exceptions than to a claim of perfection.
27. Change management, incentives and communications campaign (after 3, 9, 23)
Run this in parallel from day one. The process will be judged by engineers on fairness, and by executives on visible results.
Communicate the deal explicitly and repeatedly.
- The deal in one sentence: you are paid for on-call, you are paged only for what you own, a trained commander runs the incident, and your postmortem actions get real sprint capacity.
- Launch communications: CTO all-hands, per-team roadshows, a one-page process card for laptops, an internal wiki hub and a Slack support channel with a 4-hour answer SLA.
- Recognition: incident-response contribution in promotion criteria and performance frameworks, quarterly awards for best postmortem and biggest alert-noise reduction, public thanks after every SEV1.
- Manager accountability: adoption, alert hygiene, action closure and on-call load in each engineering manager's quarterly objectives.
- Handle the exceptions: a written path for engineers who cannot do nights (caring responsibilities, health), covered by stipended volunteers elsewhere.
- Track sentiment quarterly and publish the results, including bad news, to keep credibility.
28. Program risk register and contingency planning (after 1)
Name the ways this program fails and pre-commit the response. Review it monthly in the steering group.
The main risks are predictable.
- Volunteer shortfall for the IC pool: contingency is to make command a rostered duty for engineering managers and senior engineers until the pool reaches 25.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but do not launch mandatory night on-call without some compensation.
- Tool migration slipping: keep the single-queue requirement and cut scope on the incident-management layer, not on paging consolidation.
- Alert pruning causing a missed incident: prune from paging to ticket first, observe for 30 days, then delete; keep a recovery path.
- Burnout or attrition among the 12 experienced on-call teams: monitor load weekly and cap individual page counts.
- A major SEV1 mid-rollout: pre-agree that the program lead becomes a full-time responder and the wave schedule slips by one wave, with the steering group informed the same day.
29. Continuous improvement, maturity roadmap and post-audit sustainability (after 24, 25, 26)
Protect against the classic failure: the process decays once the audit is passed. Build the second-year plan before the first year ends.
Set a maturity model and a forward roadmap with owners.
- Quarterly process retrospective with the IC pool: what in the process itself slowed us down, what needs simplifying, is the severity scale calibrated?
- Re-baseline targets every six months; a process that hits all targets is under-ambitious.
- Year-two roadmap candidates: follow-the-sun coverage cell, automated mitigation and self-healing for the top three recurring causes, error-budget policy that gates releases, per-customer real-time impact reporting, and blast-radius reduction for the shared ledger cluster (the largest single structural risk).
- Move from lagging metrics (MTTR) to leading ones (error-budget burn, near-miss rate, drill performance).
- Make the annual policy review, certification renewal and drill calendar permanent calendar items owned by the Head of Reliability, independent of the audit cycle.
- Report to the board quarterly on availability, credits and incident trends so the process keeps executive attention after SOC 2 is signed.
- Median time to detect falls from 22 minutes to under 5 minutes within 9 months.
- Customer-first detection falls from 40% of incidents to under 10% within 6 months and under 5% within 12.
- Median time to mitigate falls from 3 h 10 min to under 60 minutes within 12 months.
- An Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes.
- Status page updated within 15 minutes of SEV1 declaration and 30 minutes of SEV2 in 95% of cases.
- Monthly alert volume falls from 3,400 to under 500 pages, with actionability above 75%; out-of-hours pages under 2 per person per week.
- All six legacy alerting tools consolidated into one paging platform, legacy paging paths disabled, by week 16.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by their due date.
- SLA credits fall from $1.3M to under $400k in the first 12 months.
- Customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar.
- All 28 teams onboarded by week 24, with 24x7 rotations of 6+ certified responders for every Tier 0/1 team.
- 30+ certified Incident Commanders and 20+ certified Communications Leads, giving 24x7 primary and secondary command cover.
- Paid on-call policy approved by HR, Legal and Finance and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among engineers on rotation.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked action items.
[SYSTEM] You are an expert assistant in complex project planning. Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform. Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible. Take as much time and space as you need to think through all aspects of the problem. After your thorough analysis, answer with the plan in the requested structure. Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text. [HUMAN] Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code". Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit." Generate an initial, detailed, and well-structured plan proposal. The plan has these parts: - "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none). - "estimated_complexity": "low", "medium" or "high". - "success_metrics": clear and measurable success metrics, one per line as a bullet list. Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Program charter, executive mandate and funding", "description": "Convert the CEO's frustration into a named program with one accountable owner, a budget and a deadline that is earlier than the audit.\n\nThe charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, the 28 teams will opt out.\n\n- Appoint a single Incident Management Program Lead (Head of Reliability/SRE) with direct exec sponsorship from CTO and CEO.\n- Form a steering group: CTO, VP Eng, Head of Support/CS, CISO/Compliance, Legal, Finance (SLA credits), HR (on-call pay).\n- Set the non-negotiables: one severity scale, one paging tool, one postmortem format, mandatory action tracking, paid on-call.\n- Fix the timeline: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready by week 28 (four weeks of buffer before the audit).\n- Approve budget lines: tooling (~$150–250k/yr), on-call compensation (~$600k–1M/yr), 2–3 dedicated program FTEs. Anchor it against $1.3M of credits plus incident cost.", "dependencies": []}, {"step_id": "S2", "title": "Forensic baseline of the 31 incidents and the alert estate", "description": "Before designing anything, rebuild the facts. Re-open all 31 incidents and profile the 3,400 monthly alerts so every later design decision is evidence-based.\n\nThis also creates the \"before\" picture the exec and the auditor will compare against.\n\n- Re-code each incident: trigger, service, detection source (customer vs monitor), timestamps for detect/acknowledge/declare/mitigate/resolve, who led, credits paid, root cause family.\n- Quantify the 40% customer-first detections: which signal was missing in each case.\n- Classify the two \"nobody in charge\" incidents minute by minute; use them as the burning-platform story.\n- Audit the six alerting tools: volume per tool, per team, per alert rule; identify the top 50 rules that produce most of the 85% noise; find rules with no owner and no runbook.\n- Baseline the numbers formally: MTTD 22 min, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed. Freeze them as the reference line.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Stakeholder listening tour and resistance map", "description": "Engineer pushback against \"carrying a pager for other teams' code\" is the main delivery risk. Treat it as a design input, not an attitude problem.\n\nRun structured interviews across all 28 teams, plus Support, CS and Sales, in two weeks.\n\n- Test the real objection: is it unpaid work, night sleep, unfamiliar code, poor runbooks, or fear of blame? Each has a different fix.\n- Collect current informal practices — the 12 teams already on-call are the pilot candidates and the source of veterans.\n- Document the promise that answers the objection: **you are only paged for services your team owns**, plus a trained commander who runs the incident and pulls in others.\n- Map influencers and blockers by team; recruit 10–15 credible engineers as a design working group so the process is co-authored, not imposed.\n- Survey baseline sentiment (trust in alerts, willingness to be on-call, burnout) to re-measure at 6 and 12 months.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Service ownership catalog and criticality tiering", "description": "You cannot page the right person across 180 services until each service has a named owning team. This is the foundation of both on-call fairness and severity mapping.\n\nBuild a machine-readable catalog (Backstage or equivalent) that is the single source of truth for routing.\n\n- One owning team per service, a named engineering manager, a Slack channel, a paging escalation policy, a dependency list.\n- Tier services by business impact: Tier 0 (money movement, ledger, auth, shared PostgreSQL cluster), Tier 1 (customer-facing but degradable), Tier 2 (internal/batch), Tier 3 (non-critical).\n- Map each Tier 0/1 service to the customer-visible capability it supports (payment initiation, settlement, reporting, onboarding).\n- Flag orphan services and cross-team shared components; force an ownership decision for each within 30 days, or schedule decommissioning.\n- Publish coverage gaps to the steering group: any Tier 0 service without an owner is an executive escalation.", "dependencies": ["S2"]}, {"step_id": "S5", "title": "Severity scale and declaration criteria", "description": "Define a five-level scale with objective, payments-specific triggers so declaration is a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade.\n\nEach level triggers a fixed bundle of response, comms and postmortem obligations.\n\n- **SEV1**: money movement stopped or incorrect, ledger integrity in doubt, data breach, full region loss, >10% of customers impacted. Triggers: immediate 24x7 page of IC + comms + exec, bridge within 5 min, status page within 15 min, mandatory postmortem, regulator assessment.\n- **SEV2**: severe degradation, settlement at risk of missing a window, single large/strategic customer fully down, SLA breach likely. Triggers: IC paged, status page within 30 min, mandatory postmortem.\n- **SEV3**: partial or workaround-available degradation, no credit exposure. Team-led, business-hours comms, postmortem optional but encouraged.\n- **SEV4/5**: minor or internal-only; ticket-tracked, no paging.\n- Add auto-escalation rules: any SEV3 open >2 h, or any incident touching the shared ledger cluster, becomes SEV2 automatically. Include a severity decision tree and 12 worked examples drawn from the 31 real incidents.", "dependencies": ["S2", "S4"]}, {"step_id": "S6", "title": "Incident roles, decision authority and handover rules", "description": "Solve the \"nobody in charge for an hour\" failure by making command explicit, transferable and logged.\n\nDefine five roles with written responsibilities, entry criteria and explicit authority.\n\n- **Incident Commander**: owns the incident, not the fix. Authority to declare severity, pull any engineer, approve customer-impacting mitigations, invoke failover and authorise spend. The IC never types in the terminal.\n- **Communications Lead**: owns status page, internal updates, account-manager briefings and the exec summary. Single voice to customers.\n- **Scribe**: maintains the timeline, decisions and open questions; feeds the postmortem and the audit evidence trail.\n- **Subject-Matter Responders**: engineers from owning teams; they investigate and remediate, and report to the IC.\n- **Executive Liaison** (SEV1 only): shields the IC from exec questions and owns regulator/board escalation.\n- Rules: the IC role is assumed within 5 minutes of declaration, stated explicitly in the channel (\"I am IC\"), and any handover is announced and logged. Roles may be combined below SEV2; never at SEV1.", "dependencies": ["S5"]}, {"step_id": "S7", "title": "24x7 incident command coverage model", "description": "Command is staffed by a small, trained, cross-team pool — not by 28 teams individually. This is what makes 24x7 realistic in one New York time zone.\n\nDesign a central rotation that scales with a growing certified pool.\n\n- Create a **Duty Incident Commander** rotation of 25–35 certified volunteers (target ~1 per team, plus managers and senior engineers), giving each person roughly one week per 6–8 months.\n- Pair with a Duty Comms Lead rotation (Support/CS leads plus engineering managers, ~15–20 people) and a Scribe pool (rotating, lowest barrier, used as the training entry point).\n- Coverage: primary + secondary IC at all times; hard 5-minute acknowledgement SLA with automatic failover to secondary, then to the on-call engineering director.\n- Night coverage options to evaluate in writing: US-only rotation with paid night stipend now; a Lisbon/Dublin or APAC follow-the-sun cell as a 12-month option; a 24x7 NOC-style triage desk for first-line detection.\n- Eligibility: certification required (S21); commanders are volunteers with manager approval and can step out with 30 days' notice.", "dependencies": ["S6", "S4"]}, {"step_id": "S8", "title": "Team on-call structure, rotations and routing rules", "description": "Rebuild team on-call around the principle that answers the pushback: **you are only paged for code your team owns**.\n\nApply a tiered obligation so 28 teams are not treated identically.\n\n- Tier 0/1 owning teams (expected 14–18 teams): 24x7 primary + secondary, minimum 6 people per rotation, one-week shifts, handover Wednesday mornings.\n- Tier 2/3 teams: business-hours on-call with a best-effort out-of-hours escalation path, no night paging.\n- Platform/Infrastructure and Database teams: 24x7, since they own the shared PostgreSQL ledger cluster and the Kubernetes/regional layer.\n- Rotations under 6 people are merged across teams or backfilled by hiring; no rotation of fewer than 4 is approved.\n- Routing: every page resolves through the service catalog to the owning team's escalation policy; cross-team pages are made by the IC, never by an alert.\n- Guardrails: maximum one week in four, no on-call in the week after a SEV1 you led, protected recovery time after any night page, and a per-person page budget (see S10).", "dependencies": ["S4", "S6"]}, {"step_id": "S9", "title": "On-call compensation, labour compliance and fairness policy", "description": "Unpaid on-call is both a retention risk and a legal exposure in New York. Paying for it is the fastest way to convert resistance into participation.\n\nDesign the scheme with HR, Legal, Finance and Payroll, and publish it before asking anyone to sign up.\n\n- Base stipend per week on rotation, differentiated by tier: e.g. $800–1,200 for 24x7 Tier 0/1, $300–500 for business-hours rotations, with premiums for holidays and weekends.\n- Per-incident payment for out-of-hours activation (e.g. $150 per night page plus hourly beyond one hour) and guaranteed time-off-in-lieu after night work.\n- Separate Duty IC stipend, since command is a distinct and heavier burden.\n- Verify FLSA exempt/non-exempt treatment, NY State wage rules and overtime exposure for non-exempt staff; document the legal review.\n- Budget and model the annual cost; get board/CFO approval as a line item, benchmarked against $1.3M of credits.\n- Add non-cash elements: on-call time counted as delivery load (teams reduce sprint commitment by ~15%), incident leadership recognised in promotion criteria, and a public quarterly report of on-call load per team.", "dependencies": ["S8", "S3"]}, {"step_id": "S10", "title": "Alert quality standard and page budget", "description": "3,400 alerts a month at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a contractual condition of paging someone.\n\nPublish a standard, then enforce it mechanically.\n\n- Every paging alert must have: a named owning team, a documented customer impact, a runbook link, a tested threshold, and a severity mapping. Alerts failing this are demoted to ticket or deleted.\n- Page only on symptoms that affect customers (SLO burn rate, error budget, queue depth against settlement deadlines); cause-based CPU/memory alerts become dashboards or tickets.\n- Set a **page budget**: maximum 2 out-of-hours pages per person per week. Breach triggers a mandatory alert-tuning sprint for the owning team and blocks new alert creation.\n- Auto-quarantine: any alert that fires more than 5 times a month without action, or has >70% no-action acknowledgements, is silenced automatically and returned to its owner.\n- Monthly alert review per team: kill, tune, or keep, with the numbers on screen.\n- Target: 3,400 → under 500 pages a month, with actionability above 75% within six months.", "dependencies": ["S2", "S4"]}, {"step_id": "S11", "title": "Detection uplift: SLOs, synthetic journeys and ledger assurance", "description": "The goal is to stop customers telling you first. Detection must be driven by customer-visible outcomes, not host metrics.\n\nInstrument the money path end to end and alert on it.\n\n- Define SLOs for each Tier 0/1 customer capability: payment initiation success rate, authorisation latency, settlement file timeliness, API availability and reporting freshness. Tie them to the 99.95% contractual SLA with a stricter internal target.\n- Deploy synthetic transactions from outside the platform, in both regions, every 60 seconds, covering the full payment lifecycle including a small real-value canary flow where feasible.\n- Add ledger assurance checks: continuous double-entry balance reconciliation, replication lag and failover-readiness alarms on the shared PostgreSQL cluster, and settlement-window countdown alerts.\n- Build per-customer anomaly detection for the top 100 accounts (volume drop, error spike) so a single-tenant outage is detected before the account manager calls.\n- Create an \"inbound signal\" bridge: any support ticket or account-manager report matching impact keywords auto-creates a triage incident within 5 minutes.\n- Track every incident's detection source; make \"customer detected first\" a reviewed defect with its own follow-up action.", "dependencies": ["S10", "S4"]}, {"step_id": "S12", "title": "Tool consolidation and incident platform implementation", "description": "Collapse six alerting tools into one paging and incident platform so there is a single queue, a single timeline and a single audit record.\n\nRun a short, time-boxed selection and migrate within the pilot window.\n\n- Select an integrated stack: paging/on-call scheduling plus an incident management layer (e.g. PagerDuty + incident.io/FireHydrant, or a single vendor) and a hosted status page.\n- Implement one-command declaration in Slack (`/incident declare`) that creates the channel and bridge, pages the Duty IC, sets severity, opens the timeline and starts the clock.\n- Migrate all monitoring sources to route into the one platform; decommission direct paging from the legacy six and block new integrations that bypass it.\n- Automate the evidence trail: timestamps, role assignments, severity changes, comms sent, and postmortem linkage exported for SOC 2.\n- Integrate with the service catalog for routing, Jira for actions, Salesforce/CS tooling for affected-customer lists, and Zoom/Slack Huddle for the bridge.\n- Hard requirement: the platform must work when AWS in one region is down — verify out-of-band paging (SMS/phone) and a printed/offline fallback runbook.", "dependencies": ["S5", "S7", "S8", "S10"]}, {"step_id": "S13", "title": "Detection-to-escalation path and the five-minute command rule", "description": "Write the single path from \"something looks wrong\" to \"someone is in charge\", and make it impossible to skip.\n\nThe design target is detection to commander in under five minutes, any hour.\n\n- Entry points: automated alert, engineer observation, support ticket, account manager, partner bank, customer-facing SEV hotline. All converge on the same declaration command.\n- Anyone in the company may declare up to SEV2; nobody is punished for over-declaring. Publish that rule in writing and repeat it.\n- Auto-page ladder: Duty IC (5 min) → secondary IC (5 min) → on-call Director (10 min) → CTO. Same ladder for the owning team's responder.\n- Cross-team pull: the IC can page any team's on-call directly, with a 10-minute acknowledgement obligation. This is the reciprocal commitment that makes single-team ownership viable.\n- Explicit takeover protocol: if no one claims IC within 5 minutes, the platform assigns it and announces it; the assignee cannot decline, only hand over.\n- Define standing severity triggers for immediate regional failover, ledger read-only mode and partner-bank notification, with pre-authorised decision rights so the IC does not wait for an executive.", "dependencies": ["S6", "S7", "S12"]}, {"step_id": "S14", "title": "Internal communications protocol", "description": "Standardise the internal channel so responders, executives and support see the same picture without interrupting the IC.\n\nSeparate the working channel from the audience channel.\n\n- One incident channel per incident (auto-created), one bridge, and a read-only broadcast channel for executives, Support and Sales.\n- Update cadence by severity: SEV1 every 30 minutes even if nothing has changed; SEV2 every 60 minutes; SEV3 at state change.\n- Fixed update template: what is happening, customer impact in plain language, what we are doing, ETA or next update time, current IC and Comms Lead.\n- Exec briefing rule: executives ask questions only to the Executive Liaison; the IC is not interrupted. Publish this as a behavioural expectation signed by the exec team.\n- Support/CS enablement: a live affected-customer list and a holding statement within 15 minutes of SEV1/SEV2 so the front line is never guessing.\n- Handover protocol for incidents beyond 4 hours: formal IC handover checklist, fatigue rule, and staffing of a second shift.", "dependencies": ["S6", "S12"]}, {"step_id": "S15", "title": "Customer communications and status page policy", "description": "Customers currently learn of outages from their own monitoring and hear from whoever happens to be around. Replace that with a timed, owned, pre-approved process.\n\nThe Comms Lead is the single author; templates remove the need to write under pressure.\n\n- Timing commitments: status page posted within 15 minutes of SEV1 declaration and 30 minutes for SEV2; updates every 30/60 minutes; resolution notice within 30 minutes of mitigation; customer-facing summary within 5 business days for SEV1.\n- Pre-approve 12–15 templates with Legal and Comms (degradation, delay in settlement, API errors, security event, third-party failure) so nothing needs legal review mid-incident.\n- Subscription-based status page with per-component granularity mapped to the customer capabilities from S11, plus an email/webhook/RSS feed.\n- Tiered outreach: top 100 accounts get a direct named call or email from their account manager within 30 minutes of SEV1, with a briefing pack from the Comms Lead; long tail gets the status page and a proactive email.\n- Rules of language: state impact and next update time, never speculate on cause, never assign blame to a vendor before facts are confirmed.\n- Run a quarterly customer-perception check with the top accounts on whether comms were timely and useful.", "dependencies": ["S5", "S14"]}, {"step_id": "S16", "title": "Regulatory, partner and legal notification playbook", "description": "In payments, some incidents are reportable and the clock starts at detection. Build the assessment into the process so it is never an afterthought.\n\nWork with Legal, Compliance and the CISO to produce a decision tree and contact matrix.\n\n- Map obligations: NYDFS Part 500 (72-hour cybersecurity event notification), state breach laws, GLBA/FTC Safeguards, PCI DSS if card data is in scope, sponsor-bank and card-network contractual notice windows, and any FinCEN/OFAC implications.\n- Add a mandatory regulatory-assessment checkpoint to every SEV1 and every security-related SEV2, owned by the Executive Liaison, completed within 2 hours of declaration and recorded even when the answer is \"not reportable\".\n- Build the contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel, with 24x7 numbers and named backups.\n- Pre-draft notification letters and hold them under legal privilege review.\n- Check customer contracts for bespoke notification SLAs (often 1–4 hours for enterprise accounts) and encode them in the customer tiering.\n- Test the playbook once per quarter as part of the simulation programme.", "dependencies": ["S5", "S15"]}, {"step_id": "S17", "title": "SLA credit and financial impact workflow", "description": "Link incidents to money so severity, credits and prioritisation stay consistent — and so Finance stops being surprised.\n\nMake credit calculation an automated output of the incident record, not a negotiation.\n\n- Define the availability measurement method per contract, per component, and agree it with Legal and Finance.\n- Auto-compute affected minutes per customer from the incident record and the per-capability telemetry; generate a proposed credit schedule within 5 business days of resolution.\n- Decide the posture: proactive credits for the top tier (reputation upside) versus claims-based for the rest; document the approval chain.\n- Track credits per incident and per root-cause family; feed a quarterly report showing which reliability investments would have prevented which credits.\n- Set a target: reduce credits from $1.3M to under $400k in year one, and use that delta as the ongoing business case.", "dependencies": ["S5", "S15"]}, {"step_id": "S18", "title": "Postmortem standard and blameless review forum", "description": "Replace \"some incidents, various formats\" with a mandatory, single-format, blameless process with fixed deadlines.\n\nThe discipline is in the deadlines and the forum, not in the template.\n\n- Mandatory for: every SEV1 and SEV2, every incident where a customer detected it first, every incident over 2 hours, every repeat of a known cause, and every near-miss involving the ledger. Optional but templated for SEV3.\n- Fixed timeline: draft within 3 business days, peer review within 5, published company-wide within 10. The IC owns delivery; the owning team's manager is accountable.\n- One template: timeline, customer and financial impact, detection analysis (why not sooner), response analysis (why mitigation took as long as it did), contributing factors, what went well, action items with owner and due date.\n- Blameless rules in writing: describe systems and decisions in the context available at the time; no individual named as a cause; HR and management commit that postmortems are never used in performance reviews.\n- Weekly 60-minute Incident Review Board: reviews all postmortems from the prior week, challenges quality, ratifies severity, and approves or rejects action items. Attendance by engineering directors is mandatory.\n- Publish a searchable postmortem library and a quarterly \"top five recurring causes\" analysis.", "dependencies": ["S5", "S6"]}, {"step_id": "S19", "title": "Action item ownership and tracking system", "description": "11 of 64 actions closed is the single clearest symptom of a process nobody enforces. Give actions the same status as customer commitments.\n\nTrack them where engineering work already lives, with visible escalation.\n\n- Every action gets: a named individual owner (not a team), a priority class, a due date and a Jira ticket auto-created from the postmortem.\n- Priority classes with hard SLAs: P0 prevents recurrence of a SEV1, due in 30 days; P1 in 60 days; P2 in 90 days. P0s are committed into the next sprint before any roadmap work.\n- Capacity rule: teams reserve a standing 15–20% of sprint capacity for reliability and incident actions. Without reserved capacity, the actions will not land.\n- Escalation ladder for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.\n- Monthly reporting of closure rate by team in the engineering leadership review; include it in manager performance objectives.\n- Target: 90% of P0/P1 actions closed on time within two quarters.", "dependencies": ["S18", "S12"]}, {"step_id": "S20", "title": "Runbooks, major-incident playbooks and the on-call readiness bar", "description": "Nobody can respond well to unfamiliar systems at 3 a.m. without runbooks — and poor runbooks are a real part of the pager resistance.\n\nDefine a minimum readiness bar that a service must meet before it is allowed to page anyone.\n\n- Readiness checklist per Tier 0/1 service: current architecture diagram, dependency map, dashboard link, alert-to-runbook mapping, rollback procedure, feature-flag kill switches, escalation contacts, and a data-loss/latency impact statement.\n- Write major-incident playbooks for the top failure modes derived from S2: shared PostgreSQL ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank/third-party outage, settlement-window breach, and suspected security compromise.\n- Prioritise the shared ledger: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and a clearly stated data-loss tolerance (RPO/RTO) signed off by the exec.\n- Runbooks must be tested at least twice a year in a drill; untested runbooks are marked stale in the catalog.\n- Enforcement: a service without readiness sign-off cannot create paging alerts, and the gap is reported to its director.", "dependencies": ["S4", "S8"]}, {"step_id": "S21", "title": "Training, certification and the commander academy", "description": "Command is a skill, not a title. Build a certification path so 24x7 coverage is staffed by people who have practised.\n\nUse a tiered curriculum with real assessment.\n\n- **Scribe** (2 hours): timeline discipline and tooling. The entry point for everyone.\n- **Responder** (half a day): severity scale, declaration, escalation, runbook use, comms hygiene. Mandatory for every engineer joining an on-call rotation.\n- **Incident Commander** (two days plus shadowing): command presence, delegation, decision-making under uncertainty, severity calls, handover, exec management. Requires two shadowed incidents and one simulated SEV1 before certification.\n- **Communications Lead** (one day): status-page writing, customer tiering, legal boundaries, regulator triggers.\n- Certification is valid 12 months and renewed via a simulation; the register of certified people is an audit artefact.\n- Add on-call onboarding per team: a new joiner shadows two shifts before holding primary, and never holds primary in their first 90 days.", "dependencies": ["S6", "S13", "S14", "S18"]}, {"step_id": "S22", "title": "Simulation programme: game days, drills and wheel of misfortune", "description": "The process must be rehearsed before it meets a real SEV1. Simulations also build the commander pool and expose runbook gaps cheaply.\n\nRun a standing calendar rather than one-off exercises.\n\n- Monthly 60-minute tabletop (\"wheel of misfortune\") per engineering group, using a real past incident from the 31.\n- Quarterly full-scale game day in production or a production-like environment: regional failover, ledger replica promotion, dependency failure, with the whole role structure activated and timed.\n- Twice-yearly unannounced paging drill to measure real acknowledgement times at night.\n- One security-incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.\n- Every exercise produces a lightweight postmortem and action items in the same system as real incidents.\n- Measure and publish drill metrics: time to IC, time to first status update, time to correct mitigation decision.", "dependencies": ["S21", "S12", "S20"]}, {"step_id": "S23", "title": "Pilot with wave 0 teams", "description": "Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.\n\nRun a six-week pilot with tight measurement and a public verdict.\n\n- Select 5–6 teams: core payments, ledger/database, platform/Kubernetes, API gateway, plus two of the 12 teams already on-call.\n- Activate the full stack for them: new severity scale, Duty IC rotation, single paging tool, alert budget, status-page policy, mandatory postmortems, paid on-call.\n- Hold a weekly pilot retro; expect and document 20–40 process defects, and fix them in the standard before rollout.\n- Validate the hard questions: does the 5-minute IC rule hold at 3 a.m.? Do cross-team pulls get answered? Is the severity tree unambiguous?\n- Exit criteria: MTTD under 10 minutes for pilot services, IC assigned within 5 minutes in 95% of incidents, page volume down 50%, all postmortems on time, positive on-call sentiment.\n- Publish a one-page pilot result to the whole company — this is the main adoption argument for the remaining teams.", "dependencies": ["S22", "S9", "S11"]}, {"step_id": "S24", "title": "Metrics, dashboards and the review cadence", "description": "Instrument the process itself, so improvement is visible and the audit has evidence of monitoring and review.\n\nDefine a small set of metrics with owners and a fixed meeting rhythm.\n\n- Response metrics: MTTD, time to declare, time to IC assigned, MTTA, MTTM, MTTR, incidents per month by severity, % detected by customers first.\n- Quality metrics: page volume per person per week, alert actionability rate, page budget breaches, postmortem on-time rate, action closure rate and ageing.\n- Business metrics: SLA credits paid, availability against 99.95% per capability, error-budget consumption, repeat-incident rate.\n- People metrics: on-call load distribution across teams, out-of-hours pages per person, on-call sentiment and attrition among on-call staff.\n- Cadence: weekly Incident Review Board (postmortems and actions), monthly Reliability Review (metrics per team, alert hygiene, on-call load), quarterly Executive/Board review (credits, trends, investment asks), annual policy review.\n- Every metric gets a target and a named owner; dashboards are self-serve and public inside the company.", "dependencies": ["S12", "S18", "S23"]}, {"step_id": "S25", "title": "Wave rollout across all 28 teams with readiness gates", "description": "Roll out in four waves of six to eight teams, every three weeks, ordered by criticality. Each wave passes an explicit gate rather than a deadline.\n\nGates keep quality high and make the standard credible.\n\n- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.\n- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with 6+ certified responders, one IC candidate nominated, one drill passed.\n- Assign each wave a named coach from the program team for three weeks of hands-on support.\n- Gate criteria are checked and signed by the director; teams that fail are re-scheduled, not waived.\n- Freeze legacy tooling per wave: after onboarding, the old alerting paths are disabled, not left as a fallback.\n- Publish a live adoption scoreboard by team so progress is social, not administrative.", "dependencies": ["S23", "S24"]}, {"step_id": "S26", "title": "SOC 2 control mapping, evidence automation and internal dry-run audit", "description": "Design the process so audit evidence is a by-product of doing the work, then test it before the auditors do.\n\nEngage the auditor early to confirm the interpretation of controls.\n\n- Map the process to the Trust Services Criteria: CC7.3 and CC7.4 (incident identification, response, recovery), CC7.2 (monitoring), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), plus availability criteria A1.2.\n- Produce and approve formal policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard — versioned, signed, annually reviewed.\n- Automate evidence: incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, postmortem library, action-item closure reports, training and certification register, drill records.\n- Confirm the observation window with the auditor and ensure the process is operating for a minimum of three months before fieldwork.\n- Run an internal dry-run audit at month six: sample 15 incidents and walk the full evidence chain; fix gaps with 8 weeks to spare.\n- Keep a remediation log for any incident where the process was not followed, with the corrective action — auditors respond better to documented exceptions than to a claim of perfection.", "dependencies": ["S18", "S19", "S25"]}, {"step_id": "S27", "title": "Change management, incentives and communications campaign", "description": "Run this in parallel from day one. The process will be judged by engineers on fairness, and by executives on visible results.\n\nCommunicate the deal explicitly and repeatedly.\n\n- The deal in one sentence: **you are paid for on-call, you are paged only for what you own, a trained commander runs the incident, and your postmortem actions get real sprint capacity**.\n- Launch communications: CTO all-hands, per-team roadshows, a one-page process card for laptops, an internal wiki hub and a Slack support channel with a 4-hour answer SLA.\n- Recognition: incident-response contribution in promotion criteria and performance frameworks, quarterly awards for best postmortem and biggest alert-noise reduction, public thanks after every SEV1.\n- Manager accountability: adoption, alert hygiene, action closure and on-call load in each engineering manager's quarterly objectives.\n- Handle the exceptions: a written path for engineers who cannot do nights (caring responsibilities, health), covered by stipended volunteers elsewhere.\n- Track sentiment quarterly and publish the results, including bad news, to keep credibility.", "dependencies": ["S3", "S9", "S23"]}, {"step_id": "S28", "title": "Program risk register and contingency planning", "description": "Name the ways this program fails and pre-commit the response. Review it monthly in the steering group.\n\nThe main risks are predictable.\n\n- **Volunteer shortfall for the IC pool**: contingency is to make command a rostered duty for engineering managers and senior engineers until the pool reaches 25.\n- **Compensation not approved in time**: fall back to time-off-in-lieu plus a phased stipend, but do not launch mandatory night on-call without some compensation.\n- **Tool migration slipping**: keep the single-queue requirement and cut scope on the incident-management layer, not on paging consolidation.\n- **Alert pruning causing a missed incident**: prune from paging to ticket first, observe for 30 days, then delete; keep a recovery path.\n- **Burnout or attrition among the 12 experienced on-call teams**: monitor load weekly and cap individual page counts.\n- **A major SEV1 mid-rollout**: pre-agree that the program lead becomes a full-time responder and the wave schedule slips by one wave, with the steering group informed the same day.", "dependencies": ["S1"]}, {"step_id": "S29", "title": "Continuous improvement, maturity roadmap and post-audit sustainability", "description": "Protect against the classic failure: the process decays once the audit is passed. Build the second-year plan before the first year ends.\n\nSet a maturity model and a forward roadmap with owners.\n\n- Quarterly process retrospective with the IC pool: what in the process itself slowed us down, what needs simplifying, is the severity scale calibrated?\n- Re-baseline targets every six months; a process that hits all targets is under-ambitious.\n- Year-two roadmap candidates: follow-the-sun coverage cell, automated mitigation and self-healing for the top three recurring causes, error-budget policy that gates releases, per-customer real-time impact reporting, and blast-radius reduction for the shared ledger cluster (the largest single structural risk).\n- Move from lagging metrics (MTTR) to leading ones (error-budget burn, near-miss rate, drill performance).\n- Make the annual policy review, certification renewal and drill calendar permanent calendar items owned by the Head of Reliability, independent of the audit cycle.\n- Report to the board quarterly on availability, credits and incident trends so the process keeps executive attention after SOC 2 is signed.", "dependencies": ["S24", "S25", "S26"]}], "estimated_complexity": "high", "success_metrics": "- Median time to detect falls from 22 minutes to under 5 minutes within 9 months.\n- Customer-first detection falls from 40% of incidents to under 10% within 6 months and under 5% within 12.\n- Median time to mitigate falls from 3 h 10 min to under 60 minutes within 12 months.\n- An Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes.\n- Status page updated within 15 minutes of SEV1 declaration and 30 minutes of SEV2 in 95% of cases.\n- Monthly alert volume falls from 3,400 to under 500 pages, with actionability above 75%; out-of-hours pages under 2 per person per week.\n- All six legacy alerting tools consolidated into one paging platform, legacy paging paths disabled, by week 16.\n- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.\n- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by their due date.\n- SLA credits fall from $1.3M to under $400k in the first 12 months.\n- Customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known cause under 10%.\n- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar.\n- All 28 teams onboarded by week 24, with 24x7 rotations of 6+ certified responders for every Tier 0/1 team.\n- 30+ certified Incident Commanders and 20+ certified Communications Leads, giving 24x7 primary and secondary command cover.\n- Paid on-call policy approved by HR, Legal and Finance and in payroll before any mandatory night rotation starts.\n- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident-response controls pass with zero exceptions.\n- On-call sentiment score improves quarter over quarter; no increase in attrition among engineers on rotation.\n- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked action items."}19 dense steps that put interim controls live in the first 7 days — one declaration path, staffed interim commanders, and triage of the 53 open historical actions — before any tool consolidation. Uses SEV0–SEV3, domain rotations of ≥6 instead of 28 fragile ones, compensation and fatigue rules (one week in six, paid recovery day, self-declared unfitness) and shadow-mode alerts for 7 days before they can page. Strongest on payments correctness: dual-control preserved during ledger recovery, split-brain/duplication guards, reconciliation required before "resolved".
Milestones are day-anchored (7/14/30/60/90/120/180) with a month-7 mock audit and metric-gaming checks.
The plan produced
1. Establish ownership, authority, and funding
Launch the program within 48 hours under an executive sponsor. Give one program owner authority to standardize incident management across all 28 teams.
- Name the CTO or equivalent as executive sponsor and a Head of Incident Management or Reliability as directly accountable owner.
- Form a working group with Engineering, SRE or Platform, Product, Support, Customer Success, Communications, Security, Legal, Compliance, Risk, HR, Finance, and Internal Audit.
- Approve authority for an incident commander to stop deployments, roll back releases, disable features, shift traffic, invoke continuity plans, and pause payment processing when integrity is at risk.
- Preserve financial controls. The incident commander may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access requirements.
- Fund paging tools, compensation, training, observability work, exercises, and dedicated reliability capacity.
- Reserve engineering capacity for incident remediation. Start with 10% of capacity and adjust through quarterly risk reviews.
- Record the current baselines: 31 customer-impacting incidents, 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M in credits, 3,400 monthly alerts, 85% noise, and 11 of 64 actions closed.
- Maintain a risk register for staffing gaps, shared-ledger concentration, regional failover, alert coverage, third parties, and audit readiness.
2. Install immediate minimum controls (after 1)
Put an interim process in place during the first seven days. Do not wait for tool consolidation, policy perfection, or the SOC 2 audit.
- Publish a one-page interim severity guide and incident declaration procedure.
- Establish one continuously monitored incident declaration path through chat, telephone, and the paging system.
- Create a standard incident channel, conference bridge, incident document, and event naming convention.
- Staff an interim primary and backup incident commander at all times. Compensate this duty retroactively under the final compensation policy.
- Give trained duty personnel access to the status page, paging system, dashboards, support queue, service catalog, and emergency contacts.
- Require an incident commander to be named within 10 minutes for every suspected major incident.
- Direct Support to escalate credible customer reports immediately rather than waiting for engineering confirmation.
- Triage all 53 open historical postmortem actions. Complete, re-plan, or formally risk-accept the items affecting ledger integrity, payment duplication, regional resilience, security, and detection first.
- Hold a daily 15-minute operational review until permanent controls are working.
3. Create the service and dependency catalog (after 1)
Build a reliable ownership map for all production services and customer journeys. This is the basis for paging, escalation, impact assessment, and audit evidence.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, data stores, queues, external processors, banking partners, and customer-facing endpoints.
- Assign each component a single accountable team, primary responder group, secondary escalation group, engineering manager, and product owner.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 according to financial integrity, customer impact, dependency centrality, and contractual obligations.
- Treat the ledger, payment orchestration, authentication, settlement, reconciliation, and critical shared infrastructure as Tier 0 or Tier 1.
- Map every important customer journey to its service, database, cloud-region, and third-party dependencies.
- Record SLOs, RTOs, RPOs, data classification, dashboards, runbooks, deployment controls, feature flags, and failover methods.
- Assign separate but coordinated responders for the ledger application and the shared PostgreSQL platform.
- Document whether each service is active-active, active-passive, or region-bound. Identify dependencies that make nominal regional redundancy ineffective.
- Make missing ownership or missing runbooks a release-blocking risk for Tier 0 and Tier 1 services.
4. Adopt severity and incident lifecycle standards (after 1)
Approve one impact-based severity model for operational, security, data, and third-party incidents. When evidence is incomplete, start at the higher credible severity and downgrade later.
- SEV0, crisis: Use for actual or credible unauthorized, lost, duplicated, or corrupted movement of money; ledger integrity loss; material security compromise; material data exposure; both-region failure; or an event likely to require crisis or regulatory management. Page all roles immediately, engage executives, Security, Legal, Compliance, and Risk, and consider pausing payment activity.
- SEV1, critical: Use for widespread inability to initiate, process, settle, or reconcile payments; a core journey failing without a viable workaround; material regional impact; fast SLA-budget exhaustion; or an imminent integrity risk. Staff all incident roles, notify the executive duty officer, and publish customer communications.
- SEV2, major: Use for a material customer subset, one or more critical customers, significant degradation with a workaround, partial transaction failure, or a likely contractual impact. Assign an incident commander and subject-matter responders; add communications and scribe roles whenever customers are affected.
- SEV3, minor: Use for localized, low-impact degradation with no financial-integrity, security, regulatory, or material contractual risk. The owning team leads the response and keeps an internal record; external communication is not normally required.
- Base severity on actual or credible impact, not the seniority of the reporter, number of alerts, or presumed complexity of the fix.
- Permit any employee to declare an incident. Only the incident commander may lower severity after recording the evidence and rationale.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of customer impact. Define resolution only after stability, backlog processing, transaction recovery, and required ledger reconciliation are complete.
- Start measurement from the earliest reliable indication of impact, including telemetry, customer reports, and partner notifications.
5. Define roles and sustainable 24x7 staffing (after 3, 4)
Separate command from technical remediation. This allows trained commanders to coordinate any incident without asking engineers to debug code they do not own.
- Incident commander: Owns severity, priorities, role assignment, escalation, decision cadence, mitigation strategy, handoffs, and final closure. One person has command at a time.
- Communications lead: Owns internal notices, status-page updates, account-manager briefs, approved customer language, and coordination with Legal or regulators.
- Scribe: Maintains a timestamped timeline of observations, decisions, commands, owners, and status changes. Automation may assist but does not replace human validation for SEV0 and SEV1.
- Subject-matter responders: Diagnose and mitigate only services or domains for which they have accepted ownership, training, access, and runbooks.
- Executive duty officer: Removes organizational obstacles and approves exceptional business decisions. This role does not take command unless a formal transfer occurs.
- Security, Legal, Compliance, Support, Vendor Management, and Business Continuity join according to predefined triggers.
- Create a company-wide incident-command rotation with at least eight certified primary commanders and eight qualified backups. Use weekly rotations with explicit handoffs.
- Create similarly sustainable communications and scribe pools using Engineering Operations, Support, Customer Operations, and Communications personnel.
- Group service responders into approximately 8–12 coherent product or platform domains rather than creating 28 fragile rotations. Each domain rotation should normally contain at least six trained responders.
- Do not place an engineer into another team's responder pool without training, access, runbooks, shadow shifts, and explicit acceptance by both teams.
- Maintain dedicated database-platform and ledger-application escalation coverage for the shared PostgreSQL environment.
- Require distinct people for commander, communications, and primary technical lead during SEV0 and SEV1 incidents.
- Require a verbal and written handoff when any incident role changes. Record the exact time and the new role owner.
6. Implement compensation and fatigue safeguards (after 5)
End unpaid on-call before expanding coverage. Treat availability, interrupted personal time, and overnight recovery as compensable work.
- Pay a fixed stipend for each primary on-call week and a secondary stipend equal to a defined percentage of the primary amount.
- Pay a higher holiday stipend. Apply overtime and call-out rules to non-exempt employees as required by law.
- Give exempt employees a minimum call-out credit or equivalent paid recovery time for material after-hours work.
- Provide a paid recovery day after prolonged overnight work, a SEV0, or a qualifying SEV1. Managers must arrange daytime coverage rather than expecting normal output.
- Have HR, Finance, and employment counsel publish dollar amounts, tax treatment, eligibility, and payroll procedures within 14 days. Apply the policy consistently across teams and locations.
- Target rotations no more frequent than one week in six. Exceptions require a time-limited staffing plan and executive risk acceptance.
- Avoid consecutive primary and secondary weeks. A person must not be primary for two simultaneous domain rotations.
- Track after-hours pages, sleep interruptions, swaps, missed acknowledgements, and reported burnout by rotation.
- Trigger a staffing or alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
- Permit responders to declare themselves temporarily unfit after overnight work without performance penalty.
7. Consolidate incident and paging tooling (after 4)
Create one operational system of record while migrating safely from the six current alerting tools. Consolidation must reduce ambiguity without creating a monitoring gap.
- Select one enterprise paging and escalation platform and one integrated incident record.
- Initially ingest events from all six tools. Deduplicate, correlate, and route them through the new platform before retiring sources.
- Integrate paging with chat, conference bridges, ticket tracking, the service catalog, observability tools, and the customer status page.
- Automatically capture declaration time, acknowledgements, role assignments, severity changes, messages, decisions, mitigated time, and resolved time.
- Use role-based access, multifactor authentication, break-glass controls, immutable audit logs, and periodic access reviews.
- Provide mobile and telephone fallback paths if chat, identity, or the primary paging tool is unavailable.
- Test paging, escalation, status publication, and conference access every week.
- Retire a legacy alert path only after its signals have named owners, successful end-to-end tests, and at least two weeks of verified operation in the new platform.
8. Improve detection and enforce alert quality (after 3, 7)
Shift detection toward customer journeys, payment outcomes, and ledger integrity. Infrastructure metrics alone will not solve the current customer-first detection problem.
- Instrument payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, and reporting with SLOs and business-level success metrics.
- Run external synthetic transactions and API checks from outside the production boundary and from both AWS regions.
- Monitor transaction failure rates, processing latency, queue age, unprocessed volume, reconciliation breaks, unexpected ledger balances, duplicate identifiers, regional asymmetry, and third-party response quality.
- Correlate application telemetry with Kubernetes, AWS, PostgreSQL, network, deployment, and feature-flag events.
- Route high-priority support cases and credible partner notifications into the same incident declaration path within five minutes.
- Define noise as a page that is duplicate, informational, unactionable, non-production, or requires no timely human action.
- Require every paging alert to name an owner, affected service, urgency, customer or SLO risk, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Send non-urgent conditions to a ticket queue rather than a pager.
- Run new alerts in shadow mode for at least seven days unless an emergency risk exception is approved. Test both firing and recovery behavior.
- Review any alert with less than 50% actionability or more than three firings in seven days within two business days.
- Never silently disable a noisy alert. Verify compensating detection, record the decision, and assign a correction owner first.
- Review alert actionability, false positives, missed detection, and page load with every responder group each month.
9. Codify acknowledgement and escalation paths (after 3, 4, 5, 7, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts when a qualifying signal or customer report enters the system.
- For SEV0 and SEV1, page the owning primary immediately; page the secondary after five unacknowledged minutes; page the domain manager and company incident commander at 10 minutes; and engage the executive duty officer by 15 minutes.
- For SEV2, require primary acknowledgement within 10 minutes and incident-command assignment within 15 minutes. Escalate to the secondary and manager when either target is missed.
- For SEV3, require acknowledgement within 30 minutes when immediate production action is needed. Otherwise create a prioritized work item.
- Automatically page the company incident commander for any credible integrity or security concern, cross-team event, customer-visible Tier 0 failure, regional event, or unresolved ownership question.
- If impact remains unknown after 15 minutes, raise severity rather than waiting for certainty.
- Let the incident commander summon dependency owners, cloud support, database support, payment processors, banking partners, and vendors through maintained escalation contacts.
- Test vendor contacts and premium-support entitlements quarterly.
- Route alerts with no valid owner to the central command rotation, then treat the missing ownership record as a control defect.
- Require human acknowledgement. Delivery to a device or chat channel does not count.
- Record every missed acknowledgement, failed escalation, and manual contact workaround for review.
10. Standardize live incident execution (after 4, 5, 7, 9)
Give responders one concise operating procedure for the first minutes through resolution. Prioritize limiting customer and financial harm before proving a root cause.
- Open a dedicated channel, bridge, incident record, and timeline immediately for SEV0 through SEV2.
- Have the incident commander state severity, known impact, current hypothesis, immediate objective, assigned roles, and next update time.
- Freeze unrelated production changes during SEV0 and SEV1 incidents. Record exceptions approved by the incident commander.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Use pre-approved runbooks for region failover, Kubernetes recovery, PostgreSQL failover, credential rotation, queue recovery, and payment suspension.
- Guard against split-brain, replay, duplication, and out-of-order processing during regional or database recovery.
- Require reconciliation and controlled backlog processing before declaring payment or ledger incidents resolved.
- Keep diagnosis and mitigation workstreams separate when enough responders are available.
- State decisions and owners aloud and in the incident record. Avoid unrecorded direct-message command paths.
- Require stability for a severity-specific observation period before closure. Reopen the incident if impact recurs during that period.
- Conduct an explicit operational handback to the owning team, Support, and Customer Success.
11. Standardize internal, customer, and regulatory communications (after 4, 5, 7, 10)
Communicate known impact early without waiting for a root cause. Use approved facts, acknowledge uncertainty, and give the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish an initial customer status within 15 minutes when disclosure is operationally and legally appropriate, then update every 15 minutes.
- For SEV1, notify internal stakeholders within 15 minutes, publish an initial customer status within 15 minutes, and update at least every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send an initial notice within 30 minutes. Update at least every 60 minutes.
- Do not normally publish SEV3 events. Notify specifically affected customers if contracts or material impact require it.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. State affected capabilities, customer symptoms, workarounds, regions, and next update time.
- Do not speculate about root cause, blame, security scope, recovery time, or data integrity.
- Give account managers a single approved briefing and an affected-customer list. Prohibit contradictory or improvised incident explanations.
- Maintain templates for outages, delays, data-integrity investigation, third-party failure, regional failure, security events, and resolution.
- Issue a resolution notice only after operational recovery and required reconciliation. Provide a customer-facing incident summary within five business days for qualifying events.
- Have Legal and Compliance maintain a jurisdiction, regulator, sponsor-bank, network, cyber-insurer, partner, and contract notification matrix.
- Where applicable, explicitly track the current New York cybersecurity-event notification clock, including the 72-hour requirement, without assuming every incident is reportable.
- Have Legal record the reportability decision, decision time, evidence, approver, deadline, and submission confirmation.
- Allow Security or Legal to limit public detail during an active threat, but require the reason and an alternative stakeholder plan to be recorded.
- Coordinate service-credit calculations and contractual notices with Finance and Customer Success from the same incident record.
12. Make postmortems mandatory and actionable (after 4, 7, 10)
Use postmortems to improve systems and controls, not to assign personal blame. Keep performance or misconduct processes separate from the learning review.
- Require a postmortem for every SEV0 and SEV1.
- Require one for a SEV2 that affected customers, incurred credits, breached an SLO or contract, involved financial or data integrity, repeated a prior failure, exposed a control gap, or lasted more than two hours.
- Permit incident command, Security, Compliance, or the service owner to require a review for a near miss.
- Produce a factual draft within three business days and hold the cross-functional review within five business days.
- Use one template covering executive summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, what worked, what failed, and lessons.
- Include why monitoring did or did not detect the event before customers.
- Avoid a single-root-cause assumption. Examine technical, organizational, process, dependency, testing, and incentive factors.
- Give every action one owner, due date, priority, expected risk reduction, verification method, and linked engineering item.
- Classify actions as containment due within 7 days, corrective work due within 30 days, or strategic work normally due within 90 days.
- Require director approval and documented residual-risk acceptance for overdue high-risk actions.
- Verify effectiveness after implementation. Closing a ticket without evidence does not close the action.
- Publish broadly useful reviews internally. Maintain access-restricted versions for security, privacy, personnel, or legally privileged details.
13. Measure performance and review it routinely (after 7, 8, 11, 12)
Use outcome, process, quality, and human-sustainability measures together. Do not reward teams for suppressing declarations or hiding incidents.
- Measure detection time from first impact to first internal signal, declaration time, acknowledgement time, role-staffing time, mitigation time, resolution time, and recurrence.
- Report both median and 90th percentile. Break results down by severity, service tier, customer journey, region, detection source, and owning domain.
- Track customer-first detection, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, and role conflicts.
- Track availability, error-budget consumption, failed-payment volume, delayed value, reconciliation breaks, impacted customers, contractual breaches, and service credits.
- Track alert volume, actionability, duplicates, after-hours pages, missed pages, pages per responder, and tool-source distribution.
- Track required postmortems completed on time, actions completed by due date, action age, verified effectiveness, and repeat contributing factors.
- Track rotation size, on-call frequency, swaps, recovery days, attrition signals, and quarterly responder sentiment.
- Hold a weekly operational review for recent incidents, overdue actions, alert problems, and upcoming risk.
- Hold a monthly executive reliability review covering trends, investment decisions, accepted risks, and SLA exposure.
- Hold a quarterly resilience and control review with Security, Compliance, Risk, Internal Audit, and Product leadership.
- Use team scorecards to direct investment and assistance, not individual performance penalties.
- Reconcile dashboard data against a monthly sample of incident records and customer cases to detect metric gaming or missing incidents.
14. Train and certify participants (after 4, 5, 9, 10, 11, 12)
Train people before assigning full independent duty. Use paid working time for training, shadowing, exercises, and certification.
- Train all employees to recognize impact, declare an incident, and find the incident channel and status page.
- Train engineers and Support on severity, escalation, customer-report handling, evidence preservation, and financial-integrity precautions.
- Certify incident commanders through instruction, tabletop exercises, shadow incidents, and observed command performance.
- Train communications leads in status writing, contractual communications, regulator escalation, and avoiding unsupported claims.
- Train scribes in timestamping, decision capture, evidence hygiene, and separating fact from hypothesis.
- Require subject-matter responders to demonstrate dashboard, runbook, rollback, failover, and access competence for their assigned domain.
- Add training to new-hire onboarding and repeat role-specific certification annually.
- Appoint an incident-management champion in each of the 28 teams to collect feedback and support local adoption.
- Conduct listening sessions focused on pager fairness, cross-team boundaries, tooling friction, and psychological safety.
- Publish that command duty is process coordination, not responsibility for understanding or repairing another team's code.
15. Pilot and expand the on-call model (after 3, 5, 6, 8, 9, 14)
Pilot the model on the highest-risk customer journeys before expanding it. Correct staffing, alert, access, and compensation defects at each stage gate.
- Start with the ledger, payment orchestration, Kubernetes platform, PostgreSQL platform, authentication, settlement, and Support intake.
- Run the command, communications, and domain rotations in parallel with existing paths for two weeks.
- Verify primary and secondary coverage, handoffs, access, runbooks, paging, conference access, status publication, and compensation processing.
- Require at least two shadow shifts before independent primary duty.
- Review every pilot page within one business day for routing accuracy, actionability, responder load, and missing context.
- Expand by customer journey and dependency domain, not by arbitrary team order.
- Provide central first-line triage if useful, but keep technical remediation with the accepted service owner.
- Do not use contractors or a managed service as the sole incident commander or sole owner of payment and ledger remediation.
- Permit temporary shared domain rotations only after service owners document training, access, runbooks, and escalation boundaries.
- Set an executive-reviewed deadline and remediation plan for any production service that cannot provide sustainable 24x7 ownership.
16. Exercise regional, ledger, and communication failures (after 10, 11, 14, 15)
Validate the process under realistic conditions before relying on it. Begin in tabletop and staging environments, then use controlled production tests where risk permits.
- Run a company-wide incident-command tabletop within 30 days of policy approval.
- Exercise loss of one AWS region, Kubernetes control-plane degradation, shared PostgreSQL failure, payment-processor failure, queue backlog, credential compromise, and suspected duplicate payments.
- Exercise a simultaneous operational and security event to test command boundaries and disclosure control.
- Exercise status-page failure and loss of the primary chat or paging provider.
- Exercise overnight staffing, role handoff, executive escalation, account-manager messaging, and a potential regulator-notification decision.
- Validate backups, restore procedures, RPO, RTO, failover prerequisites, and post-recovery reconciliation.
- Do not inject uncontrolled changes into the production ledger. Use replicas, staging, simulations, or tightly governed production tests.
- Record exercise observations as tracked actions under the same ownership and due-date rules as incident actions.
- Run at least one domain exercise per quarter and two cross-company exercises before the SOC 2 audit.
17. Execute a time-boxed enterprise rollout (after 2, 6, 7, 11, 12, 13, 15)
Use fixed implementation waves so the audit deadline does not become the start date. Report progress weekly and escalate missed stage gates as business risks.
- Days 0–7: Establish governance, interim command coverage, one declaration path, provisional severity, and daily operational reviews.
- By day 14: Approve the core policy, role definitions, communications timings, compensation design, and historical-action triage.
- By day 30: Complete Tier 0 ownership, certify the first command roster, begin the on-call pilot, enable standard incident records, and run the first tabletop.
- By day 60: Provide 24x7 coverage for all Tier 0 and Tier 1 customer journeys, integrate the six alert sources, and enforce postmortem tracking.
- By day 90: Migrate critical paging, implement customer-journey detection, complete status and regulatory playbooks, and materially reduce alert noise.
- By day 120: Assign sustainable ownership and escalation for every production service and complete the first controlled regional or continuity exercise.
- By day 180: Complete tool retirement decisions, verify action closure, rerun weak scenarios, and demonstrate improving detection and mitigation trends.
- In month 7: Conduct a mock audit and executive readiness review, leaving at least one month to correct evidence or operating defects.
- Use exception records with owners, expiry dates, compensating controls, and executive approval. Do not allow indefinite verbal exceptions.
18. Build SOC 2 evidence as the process operates (after 1)
Design evidence collection at the start rather than reconstructing it before the audit. Demonstrate both control design and sustained operation.
- Map the incident process to applicable SOC 2 criteria with Compliance and the auditor, including detection, response, communication, change management, access, availability, and corrective action.
- Maintain approved, version-controlled policies, procedures, severity definitions, role descriptions, and exception records.
- Preserve rotation schedules, compensation activation, training attendance, certification, paging tests, access reviews, and exercise results.
- Preserve incident declarations, timestamps, role assignments, communications, decisions, status updates, postmortems, and corrective-action evidence.
- Record regulatory and contractual notification assessments, including decisions that no notification was required.
- Define retention, confidentiality, legal-hold, and access requirements for operational and security records.
- Sample evidence monthly and trace incidents from initial signal through action verification.
- Have Internal Audit or an independent control owner test the process in months 4 and 6.
- Correct control failures through tracked actions rather than editing historical records.
- Conduct the formal mock audit in month 7 using the same evidence populations expected for the external audit.
19. Sustain accountability and continuous improvement (after 13, 17, 18)
Make incident management an operating discipline rather than an audit project. Keep policy, staffing, tools, and investment aligned with changing customer and system risk.
- Assign permanent owners for the incident policy, paging platform, status page, service catalog, training program, and metrics.
- Review severity thresholds, communication timings, compensation, and staffing at least annually and after material incidents.
- Use incident trends to prioritize architectural work on the shared ledger, regional independence, deployment safety, dependency isolation, and graceful degradation.
- Review repeat incidents and repeat contributing factors quarterly. Require executive action when remediation repeatedly loses priority.
- Survey responders quarterly and publish actions addressing fatigue, fairness, psychological safety, and tool friction.
- Recognize effective incident leadership, early declaration, useful postmortems, and preventive work.
- Prohibit retaliation for good-faith incident declaration or escalation.
- Provide the board or risk committee a quarterly summary of severe incidents, SLA exposure, regulatory events, overdue high-risk actions, and resilience investment.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
- Within 7 days, every suspected SEV0–SEV2 has one incident record, one channel, and a named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named incident commander for more than 10 minutes.
- By day 30, 100% of Tier 0 and Tier 1 services have a named owner, primary escalation, secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and subject-matter coverage.
- By day 120, 100% of production services have sustainable ownership and tested escalation paths.
- At least 95% of SEV0 and SEV1 pages are acknowledged within 5 minutes by month 3.
- At least 95% of SEV2 pages are acknowledged within 10 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to 10 minutes by day 90 and 5 minutes by month 6.
- The proportion of incidents first detected by customers falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to below 90 minutes by day 120 and below 60 minutes by month 6.
- At least 95% of qualifying incidents meet their initial customer-communication deadline by month 3.
- At least 95% of published incidents meet their required update cadence by month 3.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without loss of Tier 0 or Tier 1 detection coverage.
- Monthly pages fall from 3,400 to no more than 1,500 by day 90, with alert actionability and missed-detection reviews used as countermeasures against unsafe suppression.
- 100% of new paging alerts satisfy the owner, runbook, dashboard, action, severity, and escalation quality rules by day 60.
- 100% of required SEV0 and SEV1 postmortems are drafted within 3 business days and reviewed within 5 business days by month 2.
- At least 90% of postmortem actions are completed by their approved due dates by month 6.
- All 53 currently open historical actions are triaged within 30 days; all unaccepted high-risk items are completed within 90 days.
- Repeat incidents with the same unaddressed contributing factor decline by at least 50% within 6 months.
- Every primary rotation has at least six trained responders or a documented, time-limited executive exception by day 120.
- No responder is routinely scheduled more frequently than one primary week in six by day 120.
- Two end-to-end cross-company exercises, including regional and ledger scenarios, are completed before the audit, with all critical findings assigned and tracked.
- Monthly availability meets or exceeds the 99.95% contractual target by month 6, with exceptions reviewed at the executive reliability meeting.
- SLA credits decline by at least 50% on an annualized trailing basis by month 8.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
[SYSTEM] You are an expert assistant in complex project planning. Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform. Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible. Take as much time and space as you need to think through all aspects of the problem. After your thorough analysis, answer with the plan in the requested structure. Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text. [HUMAN] Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code". Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit." Generate an initial, detailed, and well-structured plan proposal. The plan has these parts: - "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none). - "estimated_complexity": "low", "medium" or "high". - "success_metrics": clear and measurable success metrics, one per line as a bullet list. Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Establish ownership, authority, and funding", "description": "Launch the program within 48 hours under an executive sponsor. Give one program owner authority to standardize incident management across all 28 teams.\n\n- Name the CTO or equivalent as executive sponsor and a Head of Incident Management or Reliability as directly accountable owner.\n- Form a working group with Engineering, SRE or Platform, Product, Support, Customer Success, Communications, Security, Legal, Compliance, Risk, HR, Finance, and Internal Audit.\n- Approve authority for an incident commander to stop deployments, roll back releases, disable features, shift traffic, invoke continuity plans, and pause payment processing when integrity is at risk.\n- Preserve financial controls. The incident commander may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access requirements.\n- Fund paging tools, compensation, training, observability work, exercises, and dedicated reliability capacity.\n- Reserve engineering capacity for incident remediation. Start with 10% of capacity and adjust through quarterly risk reviews.\n- Record the current baselines: 31 customer-impacting incidents, 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M in credits, 3,400 monthly alerts, 85% noise, and 11 of 64 actions closed.\n- Maintain a risk register for staffing gaps, shared-ledger concentration, regional failover, alert coverage, third parties, and audit readiness.", "dependencies": []}, {"step_id": "S2", "title": "Install immediate minimum controls", "description": "Put an interim process in place during the first seven days. Do not wait for tool consolidation, policy perfection, or the SOC 2 audit.\n\n- Publish a one-page interim severity guide and incident declaration procedure.\n- Establish one continuously monitored incident declaration path through chat, telephone, and the paging system.\n- Create a standard incident channel, conference bridge, incident document, and event naming convention.\n- Staff an interim primary and backup incident commander at all times. Compensate this duty retroactively under the final compensation policy.\n- Give trained duty personnel access to the status page, paging system, dashboards, support queue, service catalog, and emergency contacts.\n- Require an incident commander to be named within 10 minutes for every suspected major incident.\n- Direct Support to escalate credible customer reports immediately rather than waiting for engineering confirmation.\n- Triage all 53 open historical postmortem actions. Complete, re-plan, or formally risk-accept the items affecting ledger integrity, payment duplication, regional resilience, security, and detection first.\n- Hold a daily 15-minute operational review until permanent controls are working.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Create the service and dependency catalog", "description": "Build a reliable ownership map for all production services and customer journeys. This is the basis for paging, escalation, impact assessment, and audit evidence.\n\n- Inventory all 180 services, Kubernetes clusters, AWS accounts, data stores, queues, external processors, banking partners, and customer-facing endpoints.\n- Assign each component a single accountable team, primary responder group, secondary escalation group, engineering manager, and product owner.\n- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 according to financial integrity, customer impact, dependency centrality, and contractual obligations.\n- Treat the ledger, payment orchestration, authentication, settlement, reconciliation, and critical shared infrastructure as Tier 0 or Tier 1.\n- Map every important customer journey to its service, database, cloud-region, and third-party dependencies.\n- Record SLOs, RTOs, RPOs, data classification, dashboards, runbooks, deployment controls, feature flags, and failover methods.\n- Assign separate but coordinated responders for the ledger application and the shared PostgreSQL platform.\n- Document whether each service is active-active, active-passive, or region-bound. Identify dependencies that make nominal regional redundancy ineffective.\n- Make missing ownership or missing runbooks a release-blocking risk for Tier 0 and Tier 1 services.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Adopt severity and incident lifecycle standards", "description": "Approve one impact-based severity model for operational, security, data, and third-party incidents. When evidence is incomplete, start at the higher credible severity and downgrade later.\n\n- SEV0, crisis: Use for actual or credible unauthorized, lost, duplicated, or corrupted movement of money; ledger integrity loss; material security compromise; material data exposure; both-region failure; or an event likely to require crisis or regulatory management. Page all roles immediately, engage executives, Security, Legal, Compliance, and Risk, and consider pausing payment activity.\n- SEV1, critical: Use for widespread inability to initiate, process, settle, or reconcile payments; a core journey failing without a viable workaround; material regional impact; fast SLA-budget exhaustion; or an imminent integrity risk. Staff all incident roles, notify the executive duty officer, and publish customer communications.\n- SEV2, major: Use for a material customer subset, one or more critical customers, significant degradation with a workaround, partial transaction failure, or a likely contractual impact. Assign an incident commander and subject-matter responders; add communications and scribe roles whenever customers are affected.\n- SEV3, minor: Use for localized, low-impact degradation with no financial-integrity, security, regulatory, or material contractual risk. The owning team leads the response and keeps an internal record; external communication is not normally required.\n- Base severity on actual or credible impact, not the seniority of the reporter, number of alerts, or presumed complexity of the fix.\n- Permit any employee to declare an incident. Only the incident commander may lower severity after recording the evidence and rationale.\n- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.\n- Define mitigation as the end of customer impact. Define resolution only after stability, backlog processing, transaction recovery, and required ledger reconciliation are complete.\n- Start measurement from the earliest reliable indication of impact, including telemetry, customer reports, and partner notifications.", "dependencies": ["S1"]}, {"step_id": "S5", "title": "Define roles and sustainable 24x7 staffing", "description": "Separate command from technical remediation. This allows trained commanders to coordinate any incident without asking engineers to debug code they do not own.\n\n- Incident commander: Owns severity, priorities, role assignment, escalation, decision cadence, mitigation strategy, handoffs, and final closure. One person has command at a time.\n- Communications lead: Owns internal notices, status-page updates, account-manager briefs, approved customer language, and coordination with Legal or regulators.\n- Scribe: Maintains a timestamped timeline of observations, decisions, commands, owners, and status changes. Automation may assist but does not replace human validation for SEV0 and SEV1.\n- Subject-matter responders: Diagnose and mitigate only services or domains for which they have accepted ownership, training, access, and runbooks.\n- Executive duty officer: Removes organizational obstacles and approves exceptional business decisions. This role does not take command unless a formal transfer occurs.\n- Security, Legal, Compliance, Support, Vendor Management, and Business Continuity join according to predefined triggers.\n- Create a company-wide incident-command rotation with at least eight certified primary commanders and eight qualified backups. Use weekly rotations with explicit handoffs.\n- Create similarly sustainable communications and scribe pools using Engineering Operations, Support, Customer Operations, and Communications personnel.\n- Group service responders into approximately 8–12 coherent product or platform domains rather than creating 28 fragile rotations. Each domain rotation should normally contain at least six trained responders.\n- Do not place an engineer into another team's responder pool without training, access, runbooks, shadow shifts, and explicit acceptance by both teams.\n- Maintain dedicated database-platform and ledger-application escalation coverage for the shared PostgreSQL environment.\n- Require distinct people for commander, communications, and primary technical lead during SEV0 and SEV1 incidents.\n- Require a verbal and written handoff when any incident role changes. Record the exact time and the new role owner.", "dependencies": ["S3", "S4"]}, {"step_id": "S6", "title": "Implement compensation and fatigue safeguards", "description": "End unpaid on-call before expanding coverage. Treat availability, interrupted personal time, and overnight recovery as compensable work.\n\n- Pay a fixed stipend for each primary on-call week and a secondary stipend equal to a defined percentage of the primary amount.\n- Pay a higher holiday stipend. Apply overtime and call-out rules to non-exempt employees as required by law.\n- Give exempt employees a minimum call-out credit or equivalent paid recovery time for material after-hours work.\n- Provide a paid recovery day after prolonged overnight work, a SEV0, or a qualifying SEV1. Managers must arrange daytime coverage rather than expecting normal output.\n- Have HR, Finance, and employment counsel publish dollar amounts, tax treatment, eligibility, and payroll procedures within 14 days. Apply the policy consistently across teams and locations.\n- Target rotations no more frequent than one week in six. Exceptions require a time-limited staffing plan and executive risk acceptance.\n- Avoid consecutive primary and secondary weeks. A person must not be primary for two simultaneous domain rotations.\n- Track after-hours pages, sleep interruptions, swaps, missed acknowledgements, and reported burnout by rotation.\n- Trigger a staffing or alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.\n- Permit responders to declare themselves temporarily unfit after overnight work without performance penalty.", "dependencies": ["S5"]}, {"step_id": "S7", "title": "Consolidate incident and paging tooling", "description": "Create one operational system of record while migrating safely from the six current alerting tools. Consolidation must reduce ambiguity without creating a monitoring gap.\n\n- Select one enterprise paging and escalation platform and one integrated incident record.\n- Initially ingest events from all six tools. Deduplicate, correlate, and route them through the new platform before retiring sources.\n- Integrate paging with chat, conference bridges, ticket tracking, the service catalog, observability tools, and the customer status page.\n- Automatically capture declaration time, acknowledgements, role assignments, severity changes, messages, decisions, mitigated time, and resolved time.\n- Use role-based access, multifactor authentication, break-glass controls, immutable audit logs, and periodic access reviews.\n- Provide mobile and telephone fallback paths if chat, identity, or the primary paging tool is unavailable.\n- Test paging, escalation, status publication, and conference access every week.\n- Retire a legacy alert path only after its signals have named owners, successful end-to-end tests, and at least two weeks of verified operation in the new platform.", "dependencies": ["S4"]}, {"step_id": "S8", "title": "Improve detection and enforce alert quality", "description": "Shift detection toward customer journeys, payment outcomes, and ledger integrity. Infrastructure metrics alone will not solve the current customer-first detection problem.\n\n- Instrument payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, and reporting with SLOs and business-level success metrics.\n- Run external synthetic transactions and API checks from outside the production boundary and from both AWS regions.\n- Monitor transaction failure rates, processing latency, queue age, unprocessed volume, reconciliation breaks, unexpected ledger balances, duplicate identifiers, regional asymmetry, and third-party response quality.\n- Correlate application telemetry with Kubernetes, AWS, PostgreSQL, network, deployment, and feature-flag events.\n- Route high-priority support cases and credible partner notifications into the same incident declaration path within five minutes.\n- Define noise as a page that is duplicate, informational, unactionable, non-production, or requires no timely human action.\n- Require every paging alert to name an owner, affected service, urgency, customer or SLO risk, dashboard, runbook, expected action, deduplication key, and escalation policy.\n- Send non-urgent conditions to a ticket queue rather than a pager.\n- Run new alerts in shadow mode for at least seven days unless an emergency risk exception is approved. Test both firing and recovery behavior.\n- Review any alert with less than 50% actionability or more than three firings in seven days within two business days.\n- Never silently disable a noisy alert. Verify compensating detection, record the decision, and assign a correction owner first.\n- Review alert actionability, false positives, missed detection, and page load with every responder group each month.", "dependencies": ["S3", "S7"]}, {"step_id": "S9", "title": "Codify acknowledgement and escalation paths", "description": "Make escalation automatic, time-bound, and independent of personal contacts. The clock starts when a qualifying signal or customer report enters the system.\n\n- For SEV0 and SEV1, page the owning primary immediately; page the secondary after five unacknowledged minutes; page the domain manager and company incident commander at 10 minutes; and engage the executive duty officer by 15 minutes.\n- For SEV2, require primary acknowledgement within 10 minutes and incident-command assignment within 15 minutes. Escalate to the secondary and manager when either target is missed.\n- For SEV3, require acknowledgement within 30 minutes when immediate production action is needed. Otherwise create a prioritized work item.\n- Automatically page the company incident commander for any credible integrity or security concern, cross-team event, customer-visible Tier 0 failure, regional event, or unresolved ownership question.\n- If impact remains unknown after 15 minutes, raise severity rather than waiting for certainty.\n- Let the incident commander summon dependency owners, cloud support, database support, payment processors, banking partners, and vendors through maintained escalation contacts.\n- Test vendor contacts and premium-support entitlements quarterly.\n- Route alerts with no valid owner to the central command rotation, then treat the missing ownership record as a control defect.\n- Require human acknowledgement. Delivery to a device or chat channel does not count.\n- Record every missed acknowledgement, failed escalation, and manual contact workaround for review.", "dependencies": ["S3", "S4", "S5", "S7", "S8"]}, {"step_id": "S10", "title": "Standardize live incident execution", "description": "Give responders one concise operating procedure for the first minutes through resolution. Prioritize limiting customer and financial harm before proving a root cause.\n\n- Open a dedicated channel, bridge, incident record, and timeline immediately for SEV0 through SEV2.\n- Have the incident commander state severity, known impact, current hypothesis, immediate objective, assigned roles, and next update time.\n- Freeze unrelated production changes during SEV0 and SEV1 incidents. Record exceptions approved by the incident commander.\n- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.\n- Use pre-approved runbooks for region failover, Kubernetes recovery, PostgreSQL failover, credential rotation, queue recovery, and payment suspension.\n- Guard against split-brain, replay, duplication, and out-of-order processing during regional or database recovery.\n- Require reconciliation and controlled backlog processing before declaring payment or ledger incidents resolved.\n- Keep diagnosis and mitigation workstreams separate when enough responders are available.\n- State decisions and owners aloud and in the incident record. Avoid unrecorded direct-message command paths.\n- Require stability for a severity-specific observation period before closure. Reopen the incident if impact recurs during that period.\n- Conduct an explicit operational handback to the owning team, Support, and Customer Success.", "dependencies": ["S4", "S5", "S7", "S9"]}, {"step_id": "S11", "title": "Standardize internal, customer, and regulatory communications", "description": "Communicate known impact early without waiting for a root cause. Use approved facts, acknowledge uncertainty, and give the next update time.\n\n- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish an initial customer status within 15 minutes when disclosure is operationally and legally appropriate, then update every 15 minutes.\n- For SEV1, notify internal stakeholders within 15 minutes, publish an initial customer status within 15 minutes, and update at least every 30 minutes.\n- For customer-visible SEV2, brief Support and account managers and publish or directly send an initial notice within 30 minutes. Update at least every 60 minutes.\n- Do not normally publish SEV3 events. Notify specifically affected customers if contracts or material impact require it.\n- Use status states such as Investigating, Identified, Monitoring, and Resolved. State affected capabilities, customer symptoms, workarounds, regions, and next update time.\n- Do not speculate about root cause, blame, security scope, recovery time, or data integrity.\n- Give account managers a single approved briefing and an affected-customer list. Prohibit contradictory or improvised incident explanations.\n- Maintain templates for outages, delays, data-integrity investigation, third-party failure, regional failure, security events, and resolution.\n- Issue a resolution notice only after operational recovery and required reconciliation. Provide a customer-facing incident summary within five business days for qualifying events.\n- Have Legal and Compliance maintain a jurisdiction, regulator, sponsor-bank, network, cyber-insurer, partner, and contract notification matrix.\n- Where applicable, explicitly track the current New York cybersecurity-event notification clock, including the 72-hour requirement, without assuming every incident is reportable.\n- Have Legal record the reportability decision, decision time, evidence, approver, deadline, and submission confirmation.\n- Allow Security or Legal to limit public detail during an active threat, but require the reason and an alternative stakeholder plan to be recorded.\n- Coordinate service-credit calculations and contractual notices with Finance and Customer Success from the same incident record.", "dependencies": ["S4", "S5", "S7", "S10"]}, {"step_id": "S12", "title": "Make postmortems mandatory and actionable", "description": "Use postmortems to improve systems and controls, not to assign personal blame. Keep performance or misconduct processes separate from the learning review.\n\n- Require a postmortem for every SEV0 and SEV1.\n- Require one for a SEV2 that affected customers, incurred credits, breached an SLO or contract, involved financial or data integrity, repeated a prior failure, exposed a control gap, or lasted more than two hours.\n- Permit incident command, Security, Compliance, or the service owner to require a review for a near miss.\n- Produce a factual draft within three business days and hold the cross-functional review within five business days.\n- Use one template covering executive summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, what worked, what failed, and lessons.\n- Include why monitoring did or did not detect the event before customers.\n- Avoid a single-root-cause assumption. Examine technical, organizational, process, dependency, testing, and incentive factors.\n- Give every action one owner, due date, priority, expected risk reduction, verification method, and linked engineering item.\n- Classify actions as containment due within 7 days, corrective work due within 30 days, or strategic work normally due within 90 days.\n- Require director approval and documented residual-risk acceptance for overdue high-risk actions.\n- Verify effectiveness after implementation. Closing a ticket without evidence does not close the action.\n- Publish broadly useful reviews internally. Maintain access-restricted versions for security, privacy, personnel, or legally privileged details.", "dependencies": ["S4", "S7", "S10"]}, {"step_id": "S13", "title": "Measure performance and review it routinely", "description": "Use outcome, process, quality, and human-sustainability measures together. Do not reward teams for suppressing declarations or hiding incidents.\n\n- Measure detection time from first impact to first internal signal, declaration time, acknowledgement time, role-staffing time, mitigation time, resolution time, and recurrence.\n- Report both median and 90th percentile. Break results down by severity, service tier, customer journey, region, detection source, and owning domain.\n- Track customer-first detection, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, and role conflicts.\n- Track availability, error-budget consumption, failed-payment volume, delayed value, reconciliation breaks, impacted customers, contractual breaches, and service credits.\n- Track alert volume, actionability, duplicates, after-hours pages, missed pages, pages per responder, and tool-source distribution.\n- Track required postmortems completed on time, actions completed by due date, action age, verified effectiveness, and repeat contributing factors.\n- Track rotation size, on-call frequency, swaps, recovery days, attrition signals, and quarterly responder sentiment.\n- Hold a weekly operational review for recent incidents, overdue actions, alert problems, and upcoming risk.\n- Hold a monthly executive reliability review covering trends, investment decisions, accepted risks, and SLA exposure.\n- Hold a quarterly resilience and control review with Security, Compliance, Risk, Internal Audit, and Product leadership.\n- Use team scorecards to direct investment and assistance, not individual performance penalties.\n- Reconcile dashboard data against a monthly sample of incident records and customer cases to detect metric gaming or missing incidents.", "dependencies": ["S7", "S8", "S11", "S12"]}, {"step_id": "S14", "title": "Train and certify participants", "description": "Train people before assigning full independent duty. Use paid working time for training, shadowing, exercises, and certification.\n\n- Train all employees to recognize impact, declare an incident, and find the incident channel and status page.\n- Train engineers and Support on severity, escalation, customer-report handling, evidence preservation, and financial-integrity precautions.\n- Certify incident commanders through instruction, tabletop exercises, shadow incidents, and observed command performance.\n- Train communications leads in status writing, contractual communications, regulator escalation, and avoiding unsupported claims.\n- Train scribes in timestamping, decision capture, evidence hygiene, and separating fact from hypothesis.\n- Require subject-matter responders to demonstrate dashboard, runbook, rollback, failover, and access competence for their assigned domain.\n- Add training to new-hire onboarding and repeat role-specific certification annually.\n- Appoint an incident-management champion in each of the 28 teams to collect feedback and support local adoption.\n- Conduct listening sessions focused on pager fairness, cross-team boundaries, tooling friction, and psychological safety.\n- Publish that command duty is process coordination, not responsibility for understanding or repairing another team's code.", "dependencies": ["S4", "S5", "S9", "S10", "S11", "S12"]}, {"step_id": "S15", "title": "Pilot and expand the on-call model", "description": "Pilot the model on the highest-risk customer journeys before expanding it. Correct staffing, alert, access, and compensation defects at each stage gate.\n\n- Start with the ledger, payment orchestration, Kubernetes platform, PostgreSQL platform, authentication, settlement, and Support intake.\n- Run the command, communications, and domain rotations in parallel with existing paths for two weeks.\n- Verify primary and secondary coverage, handoffs, access, runbooks, paging, conference access, status publication, and compensation processing.\n- Require at least two shadow shifts before independent primary duty.\n- Review every pilot page within one business day for routing accuracy, actionability, responder load, and missing context.\n- Expand by customer journey and dependency domain, not by arbitrary team order.\n- Provide central first-line triage if useful, but keep technical remediation with the accepted service owner.\n- Do not use contractors or a managed service as the sole incident commander or sole owner of payment and ledger remediation.\n- Permit temporary shared domain rotations only after service owners document training, access, runbooks, and escalation boundaries.\n- Set an executive-reviewed deadline and remediation plan for any production service that cannot provide sustainable 24x7 ownership.", "dependencies": ["S3", "S5", "S6", "S8", "S9", "S14"]}, {"step_id": "S16", "title": "Exercise regional, ledger, and communication failures", "description": "Validate the process under realistic conditions before relying on it. Begin in tabletop and staging environments, then use controlled production tests where risk permits.\n\n- Run a company-wide incident-command tabletop within 30 days of policy approval.\n- Exercise loss of one AWS region, Kubernetes control-plane degradation, shared PostgreSQL failure, payment-processor failure, queue backlog, credential compromise, and suspected duplicate payments.\n- Exercise a simultaneous operational and security event to test command boundaries and disclosure control.\n- Exercise status-page failure and loss of the primary chat or paging provider.\n- Exercise overnight staffing, role handoff, executive escalation, account-manager messaging, and a potential regulator-notification decision.\n- Validate backups, restore procedures, RPO, RTO, failover prerequisites, and post-recovery reconciliation.\n- Do not inject uncontrolled changes into the production ledger. Use replicas, staging, simulations, or tightly governed production tests.\n- Record exercise observations as tracked actions under the same ownership and due-date rules as incident actions.\n- Run at least one domain exercise per quarter and two cross-company exercises before the SOC 2 audit.", "dependencies": ["S10", "S11", "S14", "S15"]}, {"step_id": "S17", "title": "Execute a time-boxed enterprise rollout", "description": "Use fixed implementation waves so the audit deadline does not become the start date. Report progress weekly and escalate missed stage gates as business risks.\n\n- Days 0–7: Establish governance, interim command coverage, one declaration path, provisional severity, and daily operational reviews.\n- By day 14: Approve the core policy, role definitions, communications timings, compensation design, and historical-action triage.\n- By day 30: Complete Tier 0 ownership, certify the first command roster, begin the on-call pilot, enable standard incident records, and run the first tabletop.\n- By day 60: Provide 24x7 coverage for all Tier 0 and Tier 1 customer journeys, integrate the six alert sources, and enforce postmortem tracking.\n- By day 90: Migrate critical paging, implement customer-journey detection, complete status and regulatory playbooks, and materially reduce alert noise.\n- By day 120: Assign sustainable ownership and escalation for every production service and complete the first controlled regional or continuity exercise.\n- By day 180: Complete tool retirement decisions, verify action closure, rerun weak scenarios, and demonstrate improving detection and mitigation trends.\n- In month 7: Conduct a mock audit and executive readiness review, leaving at least one month to correct evidence or operating defects.\n- Use exception records with owners, expiry dates, compensating controls, and executive approval. Do not allow indefinite verbal exceptions.", "dependencies": ["S2", "S6", "S7", "S11", "S12", "S13", "S15"]}, {"step_id": "S18", "title": "Build SOC 2 evidence as the process operates", "description": "Design evidence collection at the start rather than reconstructing it before the audit. Demonstrate both control design and sustained operation.\n\n- Map the incident process to applicable SOC 2 criteria with Compliance and the auditor, including detection, response, communication, change management, access, availability, and corrective action.\n- Maintain approved, version-controlled policies, procedures, severity definitions, role descriptions, and exception records.\n- Preserve rotation schedules, compensation activation, training attendance, certification, paging tests, access reviews, and exercise results.\n- Preserve incident declarations, timestamps, role assignments, communications, decisions, status updates, postmortems, and corrective-action evidence.\n- Record regulatory and contractual notification assessments, including decisions that no notification was required.\n- Define retention, confidentiality, legal-hold, and access requirements for operational and security records.\n- Sample evidence monthly and trace incidents from initial signal through action verification.\n- Have Internal Audit or an independent control owner test the process in months 4 and 6.\n- Correct control failures through tracked actions rather than editing historical records.\n- Conduct the formal mock audit in month 7 using the same evidence populations expected for the external audit.", "dependencies": ["S1"]}, {"step_id": "S19", "title": "Sustain accountability and continuous improvement", "description": "Make incident management an operating discipline rather than an audit project. Keep policy, staffing, tools, and investment aligned with changing customer and system risk.\n\n- Assign permanent owners for the incident policy, paging platform, status page, service catalog, training program, and metrics.\n- Review severity thresholds, communication timings, compensation, and staffing at least annually and after material incidents.\n- Use incident trends to prioritize architectural work on the shared ledger, regional independence, deployment safety, dependency isolation, and graceful degradation.\n- Review repeat incidents and repeat contributing factors quarterly. Require executive action when remediation repeatedly loses priority.\n- Survey responders quarterly and publish actions addressing fatigue, fairness, psychological safety, and tool friction.\n- Recognize effective incident leadership, early declaration, useful postmortems, and preventive work.\n- Prohibit retaliation for good-faith incident declaration or escalation.\n- Provide the board or risk committee a quarterly summary of severe incidents, SLA exposure, regulatory events, overdue high-risk actions, and resilience investment.\n- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.", "dependencies": ["S13", "S17", "S18"]}], "estimated_complexity": "high", "success_metrics": "- Within 7 days, every suspected SEV0–SEV2 has one incident record, one channel, and a named incident commander.\n- After day 30, no SEV0 or SEV1 remains without a named incident commander for more than 10 minutes.\n- By day 30, 100% of Tier 0 and Tier 1 services have a named owner, primary escalation, secondary escalation, dashboard, and runbook.\n- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and subject-matter coverage.\n- By day 120, 100% of production services have sustainable ownership and tested escalation paths.\n- At least 95% of SEV0 and SEV1 pages are acknowledged within 5 minutes by month 3.\n- At least 95% of SEV2 pages are acknowledged within 10 minutes by month 3.\n- Median customer-impact detection time falls from 22 minutes to 10 minutes by day 90 and 5 minutes by month 6.\n- The proportion of incidents first detected by customers falls from 40% to below 20% by day 90 and below 10% by month 6.\n- Median mitigation time falls from 190 minutes to below 90 minutes by day 120 and below 60 minutes by month 6.\n- At least 95% of qualifying incidents meet their initial customer-communication deadline by month 3.\n- At least 95% of published incidents meet their required update cadence by month 3.\n- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without loss of Tier 0 or Tier 1 detection coverage.\n- Monthly pages fall from 3,400 to no more than 1,500 by day 90, with alert actionability and missed-detection reviews used as countermeasures against unsafe suppression.\n- 100% of new paging alerts satisfy the owner, runbook, dashboard, action, severity, and escalation quality rules by day 60.\n- 100% of required SEV0 and SEV1 postmortems are drafted within 3 business days and reviewed within 5 business days by month 2.\n- At least 90% of postmortem actions are completed by their approved due dates by month 6.\n- All 53 currently open historical actions are triaged within 30 days; all unaccepted high-risk items are completed within 90 days.\n- Repeat incidents with the same unaddressed contributing factor decline by at least 50% within 6 months.\n- Every primary rotation has at least six trained responders or a documented, time-limited executive exception by day 120.\n- No responder is routinely scheduled more frequently than one primary week in six by day 120.\n- Two end-to-end cross-company exercises, including regional and ledger scenarios, are completed before the audit, with all critical findings assigned and tracked.\n- Monthly availability meets or exceeds the 99.95% contractual target by month 6, with exceptions reviewed at the executive reliability meeting.\n- SLA credits decline by at least 50% on an annualized trailing basis by month 8.\n- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence."}13 steps run by a new Incident Management Office and a 28-manager Council, starting with a baseline plus peer benchmarking against Stripe/Adyen. Distinctive is the three-tier staffing where a 6-person SRE Tier 2 absorbs Kubernetes/PostgreSQL pages and a Ledger Duty Officer joins any ledger-touching incident, explicitly to defuse the cross-team pager objection. Compensation is fully costed ($500/$250 weekly, $75–150 per page, $300 SEV1 bonus, ~$420K/yr).
Weak spots: a per-team budget of 100 alerts/month contradicts the <600/month target, "56 certified ICs and all 260 engineers trained in 14 weeks" is optimistic, and step 12 funnels all eleven prior steps into one rollout.
The plan produced
1. Assess Current State and Baseline Metrics
Build an evidence-based picture of the current incident management reality before designing anything new.
Collect and catalog the last 12 months of incident data: all 31 customer-impacting incidents, 3,400 monthly alerts, on-call coverage gaps across the 28 teams, and the 11 of 64 closed postmortem action items. Interview one lead from each of the 28 teams to surface pain points, political concerns (the 'carrying a pager for other teams' pushback), and tool sprawl.
Deliverables to produce:
- Alert inventory: which of the six alert tools feed which teams, alert volume per team, noise rate per tool, and overlap between tools.
- Incident timeline analysis: median detection-to-notification-to-mitigation-to-resolution times, who detected first (internal vs. customer), who mitigated, and where handoff gaps occurred.
- On-call coverage map: which 12 teams have on-call, which 16 do not, rotation length, compensation status, and escalation paths (or lack thereof).
- Postmortem audit: format variance, action-item tracking gaps, and the two incidents with no clear owner for over an hour.
- Tooling and integration audit: Kubernetes observability stack, six alerting tools, status-page provider, communication channels (Slack, email, phone), and any existing runbooks.
- Compliance gap analysis: SOC 2 Type II CC7.3/CC7.4 requirements vs. current practice, with a risk register for the eight-month window.
- Peer benchmarking: incident management practices at 3–4 comparable B2B fintech platforms (e.g., Plaid, Stripe, Adyen) for severity scales, on-call comp, and MTTR targets.
2. Secure Executive Sponsorship and Form the IM Governance Body (after 1)
Anchor the program with visible, top-down authority so that 28 teams adopt changes they did not individually request.
The CEO email about 'outages we hear about from clients' is a ready-made mandate. Convert it into a formal sponsorship structure.
- Appoint an executive sponsor (CTO or VP Engineering) who owns the program end-to-end and reports to the CEO monthly.
- Create an Incident Management Office (IMO): one dedicated senior incident management lead, one tooling/platform engineer, and one part-time data analyst.
- Establish an Incident Management Council (IMC): one engineering manager from each of the 28 teams, plus the VP of Customer Success, a compliance lead, and a security lead. The IMC meets bi-weekly during rollout, monthly thereafter.
- Draft and circulate an executive mandate memo that states: incident response is a shared operational obligation, not a per-team favor; participation in on-call rotations is a condition of employment for production-facing roles; and the program is not optional pending the SOC 2 audit.
- Allocate a dedicated budget line for on-call compensation, tooling consolidation, status-page licensing, training, and external facilitation.
3. Define Severity Levels and Automatic Triggers (after 1)
Replace the current ad-hoc triage with a five-level severity taxonomy that every engineer, support agent, and account manager can apply in under 60 seconds.
- SEV1 – Critical: Ledger data corruption or loss, complete payment processing halt, confirmed data breach affecting customer PII or funds, or regulatory reporting breach. Triggers: automatic all-hands page to on-call, CTO + CEO paged within 10 minutes, dedicated bridge call within 5 minutes, status-page update within 15 minutes, regulator notification assessment within 1 hour, customer comms within 30 minutes.
- SEV2 – High: Payment processing degraded >30% throughput or >5% error rate, single-region failover failure, ledger read-only mode, or any condition likely to breach the 99.95% SLA within the current window. Triggers: primary + secondary on-call paged, incident commander assigned within 10 minutes, bridge call within 15 minutes, status-page update within 30 minutes, VP Engineering notified within 20 minutes.
- SEV3 – Medium: Non-critical service degradation affecting <30% of customers, single non-ledger microservice outage with working fallback, or elevated latency above SLA threshold but below halt. Triggers: primary on-call paged, IM notified within 30 minutes, status-page update within 1 hour if customer-visible, daily standup update.
- SEV4 – Low: Degraded internal tooling, minor UX bug with workaround, non-customer-facing alert. Triggers: next-business-day response, ticket created, no page unless on-call agrees.
- SEV5 – Informational / Noise: Cosmetic issues, planned-maintenance notifications, alert misfires logged for tuning. Triggers: no page, logged for weekly alert-quality review.
Define escalation rules: any SEV3 unresolved after 4 hours auto-escalates to SEV2; any SEV2 unresolved after 2 hours auto-escalates to SEV1. Severity can be downgraded only by the incident commander with IMC notification.
Publish the taxonomy as a one-page decision tree, a Slack slash-command (
/sev), and an integration into the alerting tool so that every alert carries a suggested severity.4. Define Incident Roles and Staffing Model (after 3)
Codify four mandatory roles for every SEV1/SEV2 incident and optional roles for SEV3, then solve the 24×7 staffing problem across 28 teams.
Roles
- Incident Commander (IC): owns the incident end-to-end, declares severity, assigns tasks, authorizes mitigations, decides when to escalate or stand down. Never writes code during the incident.
- Communications Lead (CL): owns status-page updates, internal Slack channels, account-manager briefings, and regulator notifications. Separate from the IC so the IC can focus on mitigation.
- Scribe / Timeline Keeper: logs every decision, action, and timestamp in the incident channel and the incident-management tool. Produces the raw timeline for the postmortem.
- Subject-Matter Responders (SMRs): 1–3 engineers from the owning team(s) who diagnose and fix. For the shared PostgreSQL ledger, a dedicated DBA responder is always required.
24×7 Staffing via a Three-Tier Follow-the-Sun Model
- Tier 1 – Front-line on-call: Primary + secondary responder per team, paged first. Covers the team's own services.
- Tier 2 – Platform / SRE on-call: A dedicated 6-person SRE rotation covering cross-cutting infrastructure: Kubernetes, the shared PostgreSQL ledger, networking, and the two AWS regions. This tier directly addresses the 'carrying a pager for other teams' concern by absorbing infrastructure incidents.
- Tier 3 – IMC escalation: Engineering managers and the IMO on-call for multi-team or SEV1 incidents. Provides the IC and CL when no team-level IC is available.
Follow-the-Sun: If any engineering hub exists in a second timezone, use it for overnight Tier 1 coverage. If not, partner with a managed on-call service for overnight first-response triage (severity declaration + paging the correct team), reducing 3 a.m. pages for NY-based engineers.
IC and CL pools: Nominate at least 2 ICs and 1 CL per team (56 ICs, 28 CLs minimum). ICs are trained and certified before they rotate. For SEV1 incidents, the IC must be a certified IC from the IMC pool, not just 'whoever is around'.
Ledger-specific rule: Because the PostgreSQL ledger is shared, a Ledger Duty Officer from the SRE Tier 2 is always on the bridge for any incident touching ledger services, regardless of which team owns the failing microservice.
5. Design On-Call Rotations, Compensation, and Alert-Quality Rules (after 4)
Make on-call sustainable, fairly compensated, and free of alert noise so engineers stop resisting participation.
Rotation Design
- 7-day rotations, one primary + one secondary per team per week. No engineer is on-call more than one week in four.
- Minimum 48-hour rest between rotations. No on-call during approved PTO.
- All 28 teams participate. Teams without current on-call get a 90-day ramp with a shadow rotation before going live.
- Tier 2 SRE rotation: 6 engineers, one week on / five weeks off, with a dedicated backup.
Compensation Package
- Base on-call stipend: $500 per week of primary on-call, $250 for secondary, paid regardless of whether pages fire.
- Page pay: $75 per acknowledged page outside business hours; $150 if the page leads to active incident work.
- Time-off-in-lieu (TOIL): Any engineer who works >4 hours overnight (00:00–06:00 local) gets a full TOIL day. >8 hours in a single incident gets 1.5 TOIL days.
- SEV1 bonus: $300 flat bonus for every engineer who actively works a SEV1 incident, paid within the next pay cycle.
- Annual on-call cap: No engineer exceeds 13 weeks of on-call per year. Exceeding the cap triggers a mandatory team-staffing review.
- Budget estimate: ~$420K/year for stipends and page pay across 28 teams; present this to the CFO as a fraction of the $1.3M annual SLA credit cost.
Alert-Quality Rules (the '85% noise' problem)
- Every alert must carry: owning team, suggested severity, runbook link, and a 30-day noise score.
- Alert budget: each team gets a maximum of 100 actionable alerts per month. Exceeding the budget triggers a mandatory alert-tuning session with the IMO.
- Noise threshold: any alert that fires >10 times in 7 days with no human action is auto-flagged for suppression or tuning within 14 days.
- Alert review cadence: weekly 30-minute alert-quality review per team; monthly cross-team alert review in the IMC.
- Sunset rule: alerts with no runbook are demoted to SEV5 after 30 days and suppressed after 60 days unless a runbook is written.
- Target: reduce monthly alert volume from 3,400 to <600 actionable alerts within 6 months.
6. Build Detection, Escalation, and Communication Paths (after 3, 4)
Eliminate the 22-minute median detection gap and the 40% customer-first-detection rate with layered monitoring and a single escalation spine.
Detection Layers
- Synthetic transactions: run a payment end-to-end through the full stack (API → service → ledger → confirmation) every 60 seconds from both AWS regions. Alert if latency >2× baseline or any step fails. This catches what per-service metrics miss.
- Customer-traffic anomaly detection: monitor API error rates, payment success rates, and latency percentiles per customer cohort. Alert on >2σ deviation.
- SLO-based alerting: define SLIs for the 99.95% SLA (availability, latency p99, ledger consistency). Alert when error budget burn rate exceeds threshold, before the SLA actually breaches.
- Infrastructure health: Kubernetes node/pod health, PostgreSQL replication lag, disk I/O, and cross-region latency.
- Support-ticket spike detection: if >5 customers open tickets about the same symptom within 10 minutes, auto-create a SEV3 candidate.
Escalation Path
- Alert fires → PagerDuty routes to Tier 1 primary → 5-min no-ack → Tier 1 secondary → 10-min no-ack → Tier 2 SRE → 15-min no-ack → IMC on-call manager → 20-min no-ack → VP Engineering auto-page.
- Any SEV1 declaration auto-pages the CTO, opens a dedicated Slack channel + Zoom bridge, and notifies the CL.
- No incident goes unowned for >15 minutes. If no IC is assigned by minute 15, the IMC on-call manager assumes IC role by default.
Internal Communications
- Dedicated Slack channels:
#inc-sev1,#inc-sev2,#inc-sev3(auto-created per incident), plus#inc-updatesfor broadcast. - IC posts a structured update every 15 minutes (SEV1), 30 minutes (SEV2), 1 hour (SEV3) into the incident channel.
- CL posts a summary to
#inc-updatesand notifies relevant engineering managers.
Customer Communications
- Status page: auto-updated via API. SEV1: first update within 15 minutes, then every 30 minutes until resolved. SEV2: first update within 30 minutes, then every hour. SEV3: within 1 hour if customer-visible.
- Account managers: CL briefs AMs via a dedicated Slack channel within 30 minutes (SEV1) or 1 hour (SEV2). AMs contact their top-20 revenue accounts directly.
- Customer email/SMS: for SEV1 and SEV2, automated notification to all 2,100 customers via the status-page subscription system within 30 minutes.
- Regulator notification: Legal/Compliance assesses within 1 hour whether NY DFS, FinCEN, or card-network notification is required. If yes, file within the regulatory deadline (typically 72 hours for NY DFS cybersecurity events). Log the decision and filing in the incident record.
Post-resolution: CL publishes a 'resolved' update within 30 minutes of mitigation. For SEV1/SEV2, a preliminary customer-facing RCA summary is published within 5 business days.
7. Standardize Postmortems with Tracking and Accountability (after 3, 6)
Fix the '11 of 64 action items closed' problem with a mandatory, uniform, blameless postmortem process backed by engineering-manager accountability.
When Mandatory
- All SEV1 and SEV2 incidents: postmortem required within 5 business days.
- SEV3 incidents: postmortem required if >50 customers affected, if the incident lasted >4 hours, or if it was customer-detected.
- SEV4/SEV5: optional, but any recurring SEV4 (≥3 times in 30 days) triggers a mandatory review.
Format (single template, enforced by tooling)
- Incident summary (severity, duration, customers affected, revenue impact, SLA credit exposure).
- Timeline (auto-generated from scribe notes + alert timestamps).
- Detection analysis: how was it detected, why did it take X minutes, could it have been faster.
- Root cause analysis using 5-Whys or fault-tree, not blame.
- Contributing factors (process, tooling, staffing, knowledge gaps).
- Impact quantification: customers affected, transactions failed, SLA credits triggered.
- Action items: each with a named owner, due date, priority, and ticket in Jira.
- Lessons learned and what went well.
Blameless Review Meeting
- Held within 5 business days, facilitated by the IMO or a trained facilitator (never the IC of that incident).
- All responders, the CL, relevant engineering managers, and an IMC representative attend.
- Ground rules: focus on system and process failures, not individual mistakes. The facilitator enforces this.
- Meeting recorded; notes published to the engineering-wide wiki within 48 hours.
Action-Item Tracking and Accountability
- Every action item is created as a Jira ticket with a due date and a named owner.
- Engineering managers are accountable: action-item completion is a standing agenda item in the bi-weekly IMC meeting. Any action item >7 days overdue is escalated to the VP Engineering.
- Completion gate: no team may close its postmortem until 100% of its action items have Jira tickets. Postmortem is 'closed' only when all tickets are resolved.
- Quarterly audit: the IMO audits action-item completion rates and reports to the IMC and the executive sponsor. Target: >90% completion within 30 days of the postmortem.
- Link postmortem quality and action-item completion to team health scores and engineering-manager performance reviews.
8. Define KPIs, Dashboards, and Governance Reviews (after 7)
Create a measurable feedback loop so leadership can see whether the program is working and where to intervene.
Primary KPIs (tracked weekly, reported monthly)
- MTTD (Median Time to Detect): target <5 minutes (from 22).
- MTTA (Median Time to Acknowledge): target <5 minutes.
- MTTM (Median Time to Mitigate): target <60 minutes for SEV1 (from 3 h 10 min), <4 hours for SEV2.
- Customer-first detection rate: target <10% (from 40%).
- SLA compliance: maintain 99.95%; track monthly SLA credit payouts, target <$100K/quarter (from ~$325K/quarter).
- Alert signal-to-noise ratio: target >80% actionable (from ~15%).
- Alert volume: target <600/month (from 3,400).
- Postmortem completion rate: 100% for SEV1/SEV2 within 5 business days.
- Action-item completion rate: >90% within 30 days (from ~17%).
- On-call health: pages per engineer per week (target <5), TOIL usage, on-call satisfaction survey score.
- Unowned incident duration: target 0 incidents with >15 minutes without an IC.
Dashboards
- Real-time operational dashboard (Grafana): current incidents, active alerts, on-call roster, SLA error-budget burn.
- Weekly leadership dashboard (auto-generated): KPI trends, open action items, alert-noise report, on-call load distribution.
- Quarterly IMC scorecard per team.
Review Cadence
- Weekly: IMO publishes KPI snapshot to
#inc-updates. - Bi-weekly IMC: review open incidents, overdue action items, alert-quality exceptions, and on-call load.
- Monthly executive review: CTO presents KPI trends, SLA credit cost, and risk register to the CEO.
- Quarterly incident-management review: deep-dive into trends, training gaps, tooling needs, and process improvements. Output fed into the next quarter's roadmap.
9. Consolidate Tooling and Build the Incident Management Platform (after 2, 3)
Replace six alerting tools and ad-hoc status-page updates with a single, integrated incident management stack.
Target Tool Architecture
- Single alerting and on-call platform (e.g., PagerDuty or Opsgenie): ingest all alerts, apply severity routing, manage on-call schedules, handle escalations, and send pages. Retire the other five tools within 6 months.
- Observability consolidation: standardize on one APM/metrics stack (e.g., Datadog or Grafana Cloud) for all 180 Kubernetes services across both AWS regions. Ensure the shared PostgreSQL ledger has dedicated dashboards.
- Status page: a dedicated, branded status page (e.g., Statuspage.io or Instatus) with API integration for auto-updates. Subscribe all 2,100 customers.
- Incident coordination tool: integrate incident-management workflows into Slack (auto-create channels, invite responders, post templates) and a dedicated incident record system (e.g., Jira Service Management, incident.io, or Rootly) for timelines, postmortems, and action-item tracking.
- Runbook repository: a central wiki (Confluence or Notion) with mandatory runbooks for every alert. No alert goes live without a linked runbook.
Implementation Tasks
- Migrate all 28 teams' alert rules into the single platform in three waves (highest-volume teams first).
- Build the severity-based routing rules and escalation policies per S3 and S6.
- Automate status-page updates triggered by severity declaration.
- Build the synthetic-transaction monitor and SLO-based alerting per S6.
- Integrate Jira for automatic action-item ticket creation from postmortems.
- Decommission legacy tools only after all teams have completed training on the new stack.
- Budget: allocate $150K–$250K/year for licensing, plus engineering time for migration.
10. Prepare for the SOC 2 Type II Audit (after 7, 8, 9)
Ensure the incident management process produces the evidence the auditor will need, well before the audit window opens in eight months.
SOC 2 Requirements to Address (CC7.3, CC7.4, CC7.5)
- Documented incident response procedures (the severity taxonomy, role definitions, communication templates).
- Evidence of incident detection, response, and recovery for every SEV1/SEV2 incident during the audit period.
- Postmortem records with action-item tracking.
- On-call schedules, training records, and escalation evidence.
- Status-page update logs and customer notification records.
- Regulator notification logs (if any).
Preparation Tasks
- The IMO maintains a SOC 2 evidence folder: every incident record, postmortem, action-item ticket, status-page update, and training completion certificate is stored and indexed.
- Conduct a mock SOC 2 audit at month 5: an internal or external auditor reviews the incident management process end-to-end and identifies gaps.
- Remediate mock-audit findings before month 7.
- Ensure the incident management tool retains all records for at least 12 months (the SOC 2 Type II observation window).
- Document the chain of custody for incident records: who accessed, modified, or closed each record.
- Prepare a narrative document describing the incident management process, roles, and controls for the auditor.
- Coordinate with the compliance lead to align incident management evidence with the broader SOC 2 scope (access controls, change management, etc.).
11. Design and Deliver Training, Runbooks, and Change Management (after 4, 5, 9)
Equip all 260 engineers, 28 team leads, account managers, and support staff with the knowledge and muscle memory to execute the new process.
Training Tracks
- All 260 engineers (2-hour session): severity taxonomy, how to acknowledge a page, how to join an incident bridge, how to hand off to an IC, and how to write a postmortem contribution. Delivered in team-level sessions over 4 weeks.
- IC pool (56+ engineers) (8-hour certification): incident command techniques, severity declaration, escalation decision-making, bridge facilitation, and blameless postmortem facilitation. Includes two tabletop exercises. Certification valid for 12 months, renewed annually.
- CL pool (28+ staff) (4-hour session): status-page writing, customer communication templates, regulator notification triggers, and AM briefing protocol.
- Account managers and support staff (1-hour session): how to read the status page, how to escalate a customer report into an incident, and what information to collect.
- SRE Tier 2 (16-hour onboarding): Kubernetes and PostgreSQL ledger deep-dive, cross-region failover runbooks, and escalation authority.
Runbooks
- Every alert must have a runbook before it is routed to on-call. The IMO provides a runbook template and audits compliance weekly.
- Priority runbooks to write first: shared PostgreSQL ledger failover, Kubernetes cluster degradation, payment-processing pipeline failure, cross-region failover, and ledger data-integrity check.
- Runbooks are peer-reviewed and version-controlled.
Change Management for Adoption
- Address the 'carrying a pager for other teams' concern directly: publish an FAQ explaining the three-tier model, the SRE Tier 2 absorbing cross-team infrastructure, the compensation package, and the TOIL policy.
- Run office hours weekly for the first 8 weeks where any engineer can ask questions or raise concerns.
- Identify team champions: one engineer per team who volunteers as an early adopter and peer mentor.
- Publish a weekly 'incident management newsletter' during rollout: what changed, what improved, KPI trends, and success stories.
- Make on-call participation a documented expectation in job descriptions and performance reviews for production-facing roles.
12. Execute Phased Rollout, Tabletop Exercises, and Continuous Improvement (after 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11)
Introduce the process in three waves so teams are not overwhelmed, then validate with exercises and iterate continuously.
Phase 1 – Weeks 1–6: Foundation
- Publish the severity taxonomy, role definitions, and communication protocols (S3, S4, S6).
- Launch the single alerting platform for the 12 teams already on-call; begin migration for the other 16.
- Activate the SRE Tier 2 rotation for the shared PostgreSQL ledger and cross-cutting infrastructure.
- Deploy the status page and test the API integration.
- Begin IC and CL training (first cohort of 20 ICs, 10 CLs).
- Publish the on-call compensation package; HR integrates stipends into payroll.
- Write the top 10 priority runbooks.
Phase 2 – Weeks 7–14: Expansion
- All 28 teams live on the single alerting platform; legacy tools in read-only mode.
- All 28 teams on the on-call rotation schedule (the 16 new teams in shadow mode for the first 4 weeks).
- Second and third IC/CL training cohorts completed.
- First tabletop exercise: simulate a SEV1 ledger corruption scenario with all roles, test the escalation path, status-page updates, and AM briefings. Debrief and fix gaps.
- Postmortem template and Jira integration live; all new incidents use the standard process.
- Alert-tuning sprint: each team reduces its alert volume by 50%.
Phase 3 – Weeks 15–24: Optimization
- All teams fully live; legacy alerting tools decommissioned.
- Second tabletop exercise: simulate a SEV1 cross-region failure with regulator notification.
- First quarterly IMC review with full KPI dashboard.
- Mock SOC 2 audit (month 5) and remediation.
- Retrospective on the rollout: survey all 260 engineers for feedback, adjust compensation or rotation rules if needed.
- Establish the continuous improvement cadence: quarterly process review, annual severity-taxonomy review, and annual on-call compensation benchmarking.
Ongoing Governance
- The IMC owns the process document and approves changes.
- The IMO tracks all KPIs and reports to the CTO monthly.
- Any process change requires IMC approval and a 2-week notice period before enforcement.
- Annual external benchmarking against peer B2B payments platforms.
13. Establish Ongoing Governance, Annual Review, and Audit Readiness Cycle (after 12)
Embed incident management as a permanent organizational capability, not a one-time project.
- Annual process review: the IMC reviews the severity taxonomy, role definitions, on-call structure, and compensation against industry benchmarks and internal KPIs. Update as needed.
- Bi-annual tabletop exercises: one SEV1 infrastructure scenario, one SEV1 data-breach/regulator scenario. Rotate the IC and CL assignments so everyone gets practice.
- Quarterly alert-quality audit: the IMO reviews alert volumes, noise rates, and runbook coverage across all 28 teams.
- On-call health survey: quarterly anonymous survey measuring burnout, fairness, and compensation satisfaction. Results reviewed by the IMC.
- SOC 2 readiness cycle: begin evidence collection immediately after each audit ends. The IMO maintains a rolling evidence folder. Mock audit at month 5 of every 12-month audit cycle.
- Postmortem maturity tracking: track the action-item completion rate monthly. If it drops below 80%, the VP Engineering intervenes.
- Incident management maturity model: adopt a 5-level maturity model (ad-hoc → defined → managed → optimized → predictive). Assess annually. Target: Level 3 within 12 months, Level 4 within 24 months.
- Budget review: annually review on-call compensation, tooling costs, and training budget against the reduction in SLA credits and incident frequency.
- MTTD reduced from 22 minutes to <5 minutes within 6 months of full rollout.
- Customer-first detection rate reduced from 40% to <10% within 6 months.
- MTTM for SEV1 incidents reduced from 3 h 10 min to <60 minutes within 9 months.
- Monthly SLA credit payouts reduced from ~$325K to <$100K per quarter within 12 months.
- Alert volume reduced from 3,400/month to <600 actionable alerts/month within 6 months; signal-to-noise ratio >80%.
- Postmortem completion rate: 100% of SEV1/SEV2 incidents have a blameless postmortem within 5 business days.
- Postmortem action-item completion rate >90% within 30 days of the postmortem (up from ~17%).
- Zero incidents with >15 minutes of unowned command (down from 2 incidents with >1 hour).
- 100% on-call coverage: all 28 teams staffed with primary + secondary on-call 24×7 within 14 weeks.
- On-call compensation adopted: 100% of on-call engineers receiving stipends and page pay; on-call satisfaction score ≥4/5 in quarterly survey.
- Status-page first-update within 15 minutes for SEV1 and 30 minutes for SEV2, 100% compliance.
- SOC 2 Type II audit passed at month 8 with zero incident-response findings.
- All 260 engineers trained; 56+ certified ICs and 28+ certified CLs active within 14 weeks.
- Six alerting tools consolidated to one within 6 months; legacy tools decommissioned.
- Regulator notification process tested: at least one tabletop exercise includes a NY DFS / FinCEN notification drill, and the legal/compliance playbook is documented and approved.
- Quarterly IMC reviews held consistently with published KPI dashboards and action-item tracking.
- On-call participation resistance resolved: <10% of engineers report 'unwilling to participate' in the 6-month pulse survey (baseline to be measured in S1).
[SYSTEM] You are an expert assistant in complex project planning. Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform. Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible. Take as much time and space as you need to think through all aspects of the problem. After your thorough analysis, answer with the plan in the requested structure. Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text. [HUMAN] Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code". Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit." Generate an initial, detailed, and well-structured plan proposal. The plan has these parts: - "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none). - "estimated_complexity": "low", "medium" or "high". - "success_metrics": clear and measurable success metrics, one per line as a bullet list. Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Assess Current State and Baseline Metrics", "description": "Build an evidence-based picture of the current incident management reality before designing anything new.\n\nCollect and catalog the last 12 months of incident data: all 31 customer-impacting incidents, 3,400 monthly alerts, on-call coverage gaps across the 28 teams, and the 11 of 64 closed postmortem action items. Interview one lead from each of the 28 teams to surface pain points, political concerns (the 'carrying a pager for other teams' pushback), and tool sprawl.\n\nDeliverables to produce:\n- **Alert inventory**: which of the six alert tools feed which teams, alert volume per team, noise rate per tool, and overlap between tools.\n- **Incident timeline analysis**: median detection-to-notification-to-mitigation-to-resolution times, who detected first (internal vs. customer), who mitigated, and where handoff gaps occurred.\n- **On-call coverage map**: which 12 teams have on-call, which 16 do not, rotation length, compensation status, and escalation paths (or lack thereof).\n- **Postmortem audit**: format variance, action-item tracking gaps, and the two incidents with no clear owner for over an hour.\n- **Tooling and integration audit**: Kubernetes observability stack, six alerting tools, status-page provider, communication channels (Slack, email, phone), and any existing runbooks.\n- **Compliance gap analysis**: SOC 2 Type II CC7.3/CC7.4 requirements vs. current practice, with a risk register for the eight-month window.\n- **Peer benchmarking**: incident management practices at 3–4 comparable B2B fintech platforms (e.g., Plaid, Stripe, Adyen) for severity scales, on-call comp, and MTTR targets.", "dependencies": []}, {"step_id": "S2", "title": "Secure Executive Sponsorship and Form the IM Governance Body", "description": "Anchor the program with visible, top-down authority so that 28 teams adopt changes they did not individually request.\n\nThe CEO email about 'outages we hear about from clients' is a ready-made mandate. Convert it into a formal sponsorship structure.\n\n- Appoint an **executive sponsor** (CTO or VP Engineering) who owns the program end-to-end and reports to the CEO monthly.\n- Create an **Incident Management Office (IMO)**: one dedicated senior incident management lead, one tooling/platform engineer, and one part-time data analyst.\n- Establish an **Incident Management Council (IMC)**: one engineering manager from each of the 28 teams, plus the VP of Customer Success, a compliance lead, and a security lead. The IMC meets bi-weekly during rollout, monthly thereafter.\n- Draft and circulate an **executive mandate memo** that states: incident response is a shared operational obligation, not a per-team favor; participation in on-call rotations is a condition of employment for production-facing roles; and the program is not optional pending the SOC 2 audit.\n- Allocate a dedicated budget line for on-call compensation, tooling consolidation, status-page licensing, training, and external facilitation.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Define Severity Levels and Automatic Triggers", "description": "Replace the current ad-hoc triage with a five-level severity taxonomy that every engineer, support agent, and account manager can apply in under 60 seconds.\n\n- **SEV1 – Critical**: Ledger data corruption or loss, complete payment processing halt, confirmed data breach affecting customer PII or funds, or regulatory reporting breach. Triggers: automatic all-hands page to on-call, CTO + CEO paged within 10 minutes, dedicated bridge call within 5 minutes, status-page update within 15 minutes, regulator notification assessment within 1 hour, customer comms within 30 minutes.\n- **SEV2 – High**: Payment processing degraded >30% throughput or >5% error rate, single-region failover failure, ledger read-only mode, or any condition likely to breach the 99.95% SLA within the current window. Triggers: primary + secondary on-call paged, incident commander assigned within 10 minutes, bridge call within 15 minutes, status-page update within 30 minutes, VP Engineering notified within 20 minutes.\n- **SEV3 – Medium**: Non-critical service degradation affecting <30% of customers, single non-ledger microservice outage with working fallback, or elevated latency above SLA threshold but below halt. Triggers: primary on-call paged, IM notified within 30 minutes, status-page update within 1 hour if customer-visible, daily standup update.\n- **SEV4 – Low**: Degraded internal tooling, minor UX bug with workaround, non-customer-facing alert. Triggers: next-business-day response, ticket created, no page unless on-call agrees.\n- **SEV5 – Informational / Noise**: Cosmetic issues, planned-maintenance notifications, alert misfires logged for tuning. Triggers: no page, logged for weekly alert-quality review.\n\nDefine **escalation rules**: any SEV3 unresolved after 4 hours auto-escalates to SEV2; any SEV2 unresolved after 2 hours auto-escalates to SEV1. Severity can be **downgraded** only by the incident commander with IMC notification.\n\nPublish the taxonomy as a one-page decision tree, a Slack slash-command (`/sev`), and an integration into the alerting tool so that every alert carries a suggested severity.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Define Incident Roles and Staffing Model", "description": "Codify four mandatory roles for every SEV1/SEV2 incident and optional roles for SEV3, then solve the 24×7 staffing problem across 28 teams.\n\n**Roles**\n- **Incident Commander (IC)**: owns the incident end-to-end, declares severity, assigns tasks, authorizes mitigations, decides when to escalate or stand down. Never writes code during the incident.\n- **Communications Lead (CL)**: owns status-page updates, internal Slack channels, account-manager briefings, and regulator notifications. Separate from the IC so the IC can focus on mitigation.\n- **Scribe / Timeline Keeper**: logs every decision, action, and timestamp in the incident channel and the incident-management tool. Produces the raw timeline for the postmortem.\n- **Subject-Matter Responders (SMRs)**: 1–3 engineers from the owning team(s) who diagnose and fix. For the shared PostgreSQL ledger, a dedicated DBA responder is always required.\n\n**24×7 Staffing via a Three-Tier Follow-the-Sun Model**\n- **Tier 1 – Front-line on-call**: Primary + secondary responder per team, paged first. Covers the team's own services.\n- **Tier 2 – Platform / SRE on-call**: A dedicated 6-person SRE rotation covering cross-cutting infrastructure: Kubernetes, the shared PostgreSQL ledger, networking, and the two AWS regions. This tier directly addresses the 'carrying a pager for other teams' concern by absorbing infrastructure incidents.\n- **Tier 3 – IMC escalation**: Engineering managers and the IMO on-call for multi-team or SEV1 incidents. Provides the IC and CL when no team-level IC is available.\n\n**Follow-the-Sun**: If any engineering hub exists in a second timezone, use it for overnight Tier 1 coverage. If not, partner with a managed on-call service for overnight first-response triage (severity declaration + paging the correct team), reducing 3 a.m. pages for NY-based engineers.\n\n**IC and CL pools**: Nominate at least 2 ICs and 1 CL per team (56 ICs, 28 CLs minimum). ICs are trained and certified before they rotate. For SEV1 incidents, the IC must be a certified IC from the IMC pool, not just 'whoever is around'.\n\n**Ledger-specific rule**: Because the PostgreSQL ledger is shared, a **Ledger Duty Officer** from the SRE Tier 2 is always on the bridge for any incident touching ledger services, regardless of which team owns the failing microservice.", "dependencies": ["S3"]}, {"step_id": "S5", "title": "Design On-Call Rotations, Compensation, and Alert-Quality Rules", "description": "Make on-call sustainable, fairly compensated, and free of alert noise so engineers stop resisting participation.\n\n**Rotation Design**\n- 7-day rotations, one primary + one secondary per team per week. No engineer is on-call more than one week in four.\n- Minimum 48-hour rest between rotations. No on-call during approved PTO.\n- All 28 teams participate. Teams without current on-call get a 90-day ramp with a shadow rotation before going live.\n- Tier 2 SRE rotation: 6 engineers, one week on / five weeks off, with a dedicated backup.\n\n**Compensation Package**\n- **Base on-call stipend**: $500 per week of primary on-call, $250 for secondary, paid regardless of whether pages fire.\n- **Page pay**: $75 per acknowledged page outside business hours; $150 if the page leads to active incident work.\n- **Time-off-in-lieu (TOIL)**: Any engineer who works >4 hours overnight (00:00–06:00 local) gets a full TOIL day. >8 hours in a single incident gets 1.5 TOIL days.\n- **SEV1 bonus**: $300 flat bonus for every engineer who actively works a SEV1 incident, paid within the next pay cycle.\n- **Annual on-call cap**: No engineer exceeds 13 weeks of on-call per year. Exceeding the cap triggers a mandatory team-staffing review.\n- Budget estimate: ~$420K/year for stipends and page pay across 28 teams; present this to the CFO as a fraction of the $1.3M annual SLA credit cost.\n\n**Alert-Quality Rules (the '85% noise' problem)**\n- Every alert must carry: owning team, suggested severity, runbook link, and a 30-day noise score.\n- **Alert budget**: each team gets a maximum of 100 actionable alerts per month. Exceeding the budget triggers a mandatory alert-tuning session with the IMO.\n- **Noise threshold**: any alert that fires >10 times in 7 days with no human action is auto-flagged for suppression or tuning within 14 days.\n- **Alert review cadence**: weekly 30-minute alert-quality review per team; monthly cross-team alert review in the IMC.\n- **Sunset rule**: alerts with no runbook are demoted to SEV5 after 30 days and suppressed after 60 days unless a runbook is written.\n- Target: reduce monthly alert volume from 3,400 to <600 actionable alerts within 6 months.", "dependencies": ["S4"]}, {"step_id": "S6", "title": "Build Detection, Escalation, and Communication Paths", "description": "Eliminate the 22-minute median detection gap and the 40% customer-first-detection rate with layered monitoring and a single escalation spine.\n\n**Detection Layers**\n- **Synthetic transactions**: run a payment end-to-end through the full stack (API → service → ledger → confirmation) every 60 seconds from both AWS regions. Alert if latency >2× baseline or any step fails. This catches what per-service metrics miss.\n- **Customer-traffic anomaly detection**: monitor API error rates, payment success rates, and latency percentiles per customer cohort. Alert on >2σ deviation.\n- **SLO-based alerting**: define SLIs for the 99.95% SLA (availability, latency p99, ledger consistency). Alert when error budget burn rate exceeds threshold, before the SLA actually breaches.\n- **Infrastructure health**: Kubernetes node/pod health, PostgreSQL replication lag, disk I/O, and cross-region latency.\n- **Support-ticket spike detection**: if >5 customers open tickets about the same symptom within 10 minutes, auto-create a SEV3 candidate.\n\n**Escalation Path**\n- Alert fires → PagerDuty routes to Tier 1 primary → 5-min no-ack → Tier 1 secondary → 10-min no-ack → Tier 2 SRE → 15-min no-ack → IMC on-call manager → 20-min no-ack → VP Engineering auto-page.\n- Any SEV1 declaration auto-pages the CTO, opens a dedicated Slack channel + Zoom bridge, and notifies the CL.\n- **No incident goes unowned for >15 minutes.** If no IC is assigned by minute 15, the IMC on-call manager assumes IC role by default.\n\n**Internal Communications**\n- Dedicated Slack channels: `#inc-sev1`, `#inc-sev2`, `#inc-sev3` (auto-created per incident), plus `#inc-updates` for broadcast.\n- IC posts a structured update every 15 minutes (SEV1), 30 minutes (SEV2), 1 hour (SEV3) into the incident channel.\n- CL posts a summary to `#inc-updates` and notifies relevant engineering managers.\n\n**Customer Communications**\n- **Status page**: auto-updated via API. SEV1: first update within 15 minutes, then every 30 minutes until resolved. SEV2: first update within 30 minutes, then every hour. SEV3: within 1 hour if customer-visible.\n- **Account managers**: CL briefs AMs via a dedicated Slack channel within 30 minutes (SEV1) or 1 hour (SEV2). AMs contact their top-20 revenue accounts directly.\n- **Customer email/SMS**: for SEV1 and SEV2, automated notification to all 2,100 customers via the status-page subscription system within 30 minutes.\n- **Regulator notification**: Legal/Compliance assesses within 1 hour whether NY DFS, FinCEN, or card-network notification is required. If yes, file within the regulatory deadline (typically 72 hours for NY DFS cybersecurity events). Log the decision and filing in the incident record.\n\n**Post-resolution**: CL publishes a 'resolved' update within 30 minutes of mitigation. For SEV1/SEV2, a preliminary customer-facing RCA summary is published within 5 business days.", "dependencies": ["S3", "S4"]}, {"step_id": "S7", "title": "Standardize Postmortems with Tracking and Accountability", "description": "Fix the '11 of 64 action items closed' problem with a mandatory, uniform, blameless postmortem process backed by engineering-manager accountability.\n\n**When Mandatory**\n- All SEV1 and SEV2 incidents: postmortem required within 5 business days.\n- SEV3 incidents: postmortem required if >50 customers affected, if the incident lasted >4 hours, or if it was customer-detected.\n- SEV4/SEV5: optional, but any recurring SEV4 (≥3 times in 30 days) triggers a mandatory review.\n\n**Format (single template, enforced by tooling)**\n- Incident summary (severity, duration, customers affected, revenue impact, SLA credit exposure).\n- Timeline (auto-generated from scribe notes + alert timestamps).\n- Detection analysis: how was it detected, why did it take X minutes, could it have been faster.\n- Root cause analysis using 5-Whys or fault-tree, not blame.\n- Contributing factors (process, tooling, staffing, knowledge gaps).\n- Impact quantification: customers affected, transactions failed, SLA credits triggered.\n- Action items: each with a **named owner**, **due date**, **priority**, and **ticket in Jira**.\n- Lessons learned and what went well.\n\n**Blameless Review Meeting**\n- Held within 5 business days, facilitated by the IMO or a trained facilitator (never the IC of that incident).\n- All responders, the CL, relevant engineering managers, and an IMC representative attend.\n- Ground rules: focus on system and process failures, not individual mistakes. The facilitator enforces this.\n- Meeting recorded; notes published to the engineering-wide wiki within 48 hours.\n\n**Action-Item Tracking and Accountability**\n- Every action item is created as a Jira ticket with a due date and a named owner.\n- **Engineering managers are accountable**: action-item completion is a standing agenda item in the bi-weekly IMC meeting. Any action item >7 days overdue is escalated to the VP Engineering.\n- **Completion gate**: no team may close its postmortem until 100% of its action items have Jira tickets. Postmortem is 'closed' only when all tickets are resolved.\n- **Quarterly audit**: the IMO audits action-item completion rates and reports to the IMC and the executive sponsor. Target: >90% completion within 30 days of the postmortem.\n- Link postmortem quality and action-item completion to team health scores and engineering-manager performance reviews.", "dependencies": ["S3", "S6"]}, {"step_id": "S8", "title": "Define KPIs, Dashboards, and Governance Reviews", "description": "Create a measurable feedback loop so leadership can see whether the program is working and where to intervene.\n\n**Primary KPIs (tracked weekly, reported monthly)**\n- **MTTD** (Median Time to Detect): target <5 minutes (from 22).\n- **MTTA** (Median Time to Acknowledge): target <5 minutes.\n- **MTTM** (Median Time to Mitigate): target <60 minutes for SEV1 (from 3 h 10 min), <4 hours for SEV2.\n- **Customer-first detection rate**: target <10% (from 40%).\n- **SLA compliance**: maintain 99.95%; track monthly SLA credit payouts, target <$100K/quarter (from ~$325K/quarter).\n- **Alert signal-to-noise ratio**: target >80% actionable (from ~15%).\n- **Alert volume**: target <600/month (from 3,400).\n- **Postmortem completion rate**: 100% for SEV1/SEV2 within 5 business days.\n- **Action-item completion rate**: >90% within 30 days (from ~17%).\n- **On-call health**: pages per engineer per week (target <5), TOIL usage, on-call satisfaction survey score.\n- **Unowned incident duration**: target 0 incidents with >15 minutes without an IC.\n\n**Dashboards**\n- Real-time operational dashboard (Grafana): current incidents, active alerts, on-call roster, SLA error-budget burn.\n- Weekly leadership dashboard (auto-generated): KPI trends, open action items, alert-noise report, on-call load distribution.\n- Quarterly IMC scorecard per team.\n\n**Review Cadence**\n- **Weekly**: IMO publishes KPI snapshot to `#inc-updates`.\n- **Bi-weekly IMC**: review open incidents, overdue action items, alert-quality exceptions, and on-call load.\n- **Monthly executive review**: CTO presents KPI trends, SLA credit cost, and risk register to the CEO.\n- **Quarterly incident-management review**: deep-dive into trends, training gaps, tooling needs, and process improvements. Output fed into the next quarter's roadmap.", "dependencies": ["S7"]}, {"step_id": "S9", "title": "Consolidate Tooling and Build the Incident Management Platform", "description": "Replace six alerting tools and ad-hoc status-page updates with a single, integrated incident management stack.\n\n**Target Tool Architecture**\n- **Single alerting and on-call platform** (e.g., PagerDuty or Opsgenie): ingest all alerts, apply severity routing, manage on-call schedules, handle escalations, and send pages. Retire the other five tools within 6 months.\n- **Observability consolidation**: standardize on one APM/metrics stack (e.g., Datadog or Grafana Cloud) for all 180 Kubernetes services across both AWS regions. Ensure the shared PostgreSQL ledger has dedicated dashboards.\n- **Status page**: a dedicated, branded status page (e.g., Statuspage.io or Instatus) with API integration for auto-updates. Subscribe all 2,100 customers.\n- **Incident coordination tool**: integrate incident-management workflows into Slack (auto-create channels, invite responders, post templates) and a dedicated incident record system (e.g., Jira Service Management, incident.io, or Rootly) for timelines, postmortems, and action-item tracking.\n- **Runbook repository**: a central wiki (Confluence or Notion) with mandatory runbooks for every alert. No alert goes live without a linked runbook.\n\n**Implementation Tasks**\n- Migrate all 28 teams' alert rules into the single platform in three waves (highest-volume teams first).\n- Build the severity-based routing rules and escalation policies per S3 and S6.\n- Automate status-page updates triggered by severity declaration.\n- Build the synthetic-transaction monitor and SLO-based alerting per S6.\n- Integrate Jira for automatic action-item ticket creation from postmortems.\n- Decommission legacy tools only after all teams have completed training on the new stack.\n- Budget: allocate $150K–$250K/year for licensing, plus engineering time for migration.", "dependencies": ["S2", "S3"]}, {"step_id": "S10", "title": "Prepare for the SOC 2 Type II Audit", "description": "Ensure the incident management process produces the evidence the auditor will need, well before the audit window opens in eight months.\n\n**SOC 2 Requirements to Address (CC7.3, CC7.4, CC7.5)**\n- Documented incident response procedures (the severity taxonomy, role definitions, communication templates).\n- Evidence of incident detection, response, and recovery for every SEV1/SEV2 incident during the audit period.\n- Postmortem records with action-item tracking.\n- On-call schedules, training records, and escalation evidence.\n- Status-page update logs and customer notification records.\n- Regulator notification logs (if any).\n\n**Preparation Tasks**\n- The IMO maintains a **SOC 2 evidence folder**: every incident record, postmortem, action-item ticket, status-page update, and training completion certificate is stored and indexed.\n- Conduct a **mock SOC 2 audit** at month 5: an internal or external auditor reviews the incident management process end-to-end and identifies gaps.\n- Remediate mock-audit findings before month 7.\n- Ensure the incident management tool retains all records for at least 12 months (the SOC 2 Type II observation window).\n- Document the **chain of custody** for incident records: who accessed, modified, or closed each record.\n- Prepare a **narrative document** describing the incident management process, roles, and controls for the auditor.\n- Coordinate with the compliance lead to align incident management evidence with the broader SOC 2 scope (access controls, change management, etc.).", "dependencies": ["S7", "S8", "S9"]}, {"step_id": "S11", "title": "Design and Deliver Training, Runbooks, and Change Management", "description": "Equip all 260 engineers, 28 team leads, account managers, and support staff with the knowledge and muscle memory to execute the new process.\n\n**Training Tracks**\n- **All 260 engineers** (2-hour session): severity taxonomy, how to acknowledge a page, how to join an incident bridge, how to hand off to an IC, and how to write a postmortem contribution. Delivered in team-level sessions over 4 weeks.\n- **IC pool (56+ engineers)** (8-hour certification): incident command techniques, severity declaration, escalation decision-making, bridge facilitation, and blameless postmortem facilitation. Includes two tabletop exercises. Certification valid for 12 months, renewed annually.\n- **CL pool (28+ staff)** (4-hour session): status-page writing, customer communication templates, regulator notification triggers, and AM briefing protocol.\n- **Account managers and support staff** (1-hour session): how to read the status page, how to escalate a customer report into an incident, and what information to collect.\n- **SRE Tier 2** (16-hour onboarding): Kubernetes and PostgreSQL ledger deep-dive, cross-region failover runbooks, and escalation authority.\n\n**Runbooks**\n- Every alert must have a runbook before it is routed to on-call. The IMO provides a runbook template and audits compliance weekly.\n- Priority runbooks to write first: shared PostgreSQL ledger failover, Kubernetes cluster degradation, payment-processing pipeline failure, cross-region failover, and ledger data-integrity check.\n- Runbooks are peer-reviewed and version-controlled.\n\n**Change Management for Adoption**\n- Address the 'carrying a pager for other teams' concern directly: publish an FAQ explaining the three-tier model, the SRE Tier 2 absorbing cross-team infrastructure, the compensation package, and the TOIL policy.\n- Run **office hours** weekly for the first 8 weeks where any engineer can ask questions or raise concerns.\n- Identify **team champions**: one engineer per team who volunteers as an early adopter and peer mentor.\n- Publish a **weekly 'incident management newsletter'** during rollout: what changed, what improved, KPI trends, and success stories.\n- Make on-call participation a documented expectation in job descriptions and performance reviews for production-facing roles.", "dependencies": ["S4", "S5", "S9"]}, {"step_id": "S12", "title": "Execute Phased Rollout, Tabletop Exercises, and Continuous Improvement", "description": "Introduce the process in three waves so teams are not overwhelmed, then validate with exercises and iterate continuously.\n\n**Phase 1 – Weeks 1–6: Foundation**\n- Publish the severity taxonomy, role definitions, and communication protocols (S3, S4, S6).\n- Launch the single alerting platform for the 12 teams already on-call; begin migration for the other 16.\n- Activate the SRE Tier 2 rotation for the shared PostgreSQL ledger and cross-cutting infrastructure.\n- Deploy the status page and test the API integration.\n- Begin IC and CL training (first cohort of 20 ICs, 10 CLs).\n- Publish the on-call compensation package; HR integrates stipends into payroll.\n- Write the top 10 priority runbooks.\n\n**Phase 2 – Weeks 7–14: Expansion**\n- All 28 teams live on the single alerting platform; legacy tools in read-only mode.\n- All 28 teams on the on-call rotation schedule (the 16 new teams in shadow mode for the first 4 weeks).\n- Second and third IC/CL training cohorts completed.\n- First **tabletop exercise**: simulate a SEV1 ledger corruption scenario with all roles, test the escalation path, status-page updates, and AM briefings. Debrief and fix gaps.\n- Postmortem template and Jira integration live; all new incidents use the standard process.\n- Alert-tuning sprint: each team reduces its alert volume by 50%.\n\n**Phase 3 – Weeks 15–24: Optimization**\n- All teams fully live; legacy alerting tools decommissioned.\n- Second **tabletop exercise**: simulate a SEV1 cross-region failure with regulator notification.\n- First quarterly IMC review with full KPI dashboard.\n- Mock SOC 2 audit (month 5) and remediation.\n- Retrospective on the rollout: survey all 260 engineers for feedback, adjust compensation or rotation rules if needed.\n- Establish the **continuous improvement cadence**: quarterly process review, annual severity-taxonomy review, and annual on-call compensation benchmarking.\n\n**Ongoing Governance**\n- The IMC owns the process document and approves changes.\n- The IMO tracks all KPIs and reports to the CTO monthly.\n- Any process change requires IMC approval and a 2-week notice period before enforcement.\n- Annual external benchmarking against peer B2B payments platforms.", "dependencies": ["S1", "S2", "S3", "S4", "S5", "S6", "S7", "S8", "S9", "S10", "S11"]}, {"step_id": "S13", "title": "Establish Ongoing Governance, Annual Review, and Audit Readiness Cycle", "description": "Embed incident management as a permanent organizational capability, not a one-time project.\n\n- **Annual process review**: the IMC reviews the severity taxonomy, role definitions, on-call structure, and compensation against industry benchmarks and internal KPIs. Update as needed.\n- **Bi-annual tabletop exercises**: one SEV1 infrastructure scenario, one SEV1 data-breach/regulator scenario. Rotate the IC and CL assignments so everyone gets practice.\n- **Quarterly alert-quality audit**: the IMO reviews alert volumes, noise rates, and runbook coverage across all 28 teams.\n- **On-call health survey**: quarterly anonymous survey measuring burnout, fairness, and compensation satisfaction. Results reviewed by the IMC.\n- **SOC 2 readiness cycle**: begin evidence collection immediately after each audit ends. The IMO maintains a rolling evidence folder. Mock audit at month 5 of every 12-month audit cycle.\n- **Postmortem maturity tracking**: track the action-item completion rate monthly. If it drops below 80%, the VP Engineering intervenes.\n- **Incident management maturity model**: adopt a 5-level maturity model (ad-hoc → defined → managed → optimized → predictive). Assess annually. Target: Level 3 within 12 months, Level 4 within 24 months.\n- **Budget review**: annually review on-call compensation, tooling costs, and training budget against the reduction in SLA credits and incident frequency.", "dependencies": ["S12"]}], "estimated_complexity": "high", "success_metrics": "- **MTTD reduced from 22 minutes to <5 minutes** within 6 months of full rollout.\n- **Customer-first detection rate reduced from 40% to <10%** within 6 months.\n- **MTTM for SEV1 incidents reduced from 3 h 10 min to <60 minutes** within 9 months.\n- **Monthly SLA credit payouts reduced from ~$325K to <$100K per quarter** within 12 months.\n- **Alert volume reduced from 3,400/month to <600 actionable alerts/month** within 6 months; signal-to-noise ratio >80%.\n- **Postmortem completion rate: 100% of SEV1/SEV2 incidents** have a blameless postmortem within 5 business days.\n- **Postmortem action-item completion rate >90% within 30 days** of the postmortem (up from ~17%).\n- **Zero incidents with >15 minutes of unowned command** (down from 2 incidents with >1 hour).\n- **100% on-call coverage**: all 28 teams staffed with primary + secondary on-call 24×7 within 14 weeks.\n- **On-call compensation adopted**: 100% of on-call engineers receiving stipends and page pay; on-call satisfaction score ≥4/5 in quarterly survey.\n- **Status-page first-update within 15 minutes for SEV1 and 30 minutes for SEV2**, 100% compliance.\n- **SOC 2 Type II audit passed** at month 8 with zero incident-response findings.\n- **All 260 engineers trained**; 56+ certified ICs and 28+ certified CLs active within 14 weeks.\n- **Six alerting tools consolidated to one** within 6 months; legacy tools decommissioned.\n- **Regulator notification process tested**: at least one tabletop exercise includes a NY DFS / FinCEN notification drill, and the legal/compliance playbook is documented and approved.\n- **Quarterly IMC reviews held consistently** with published KPI dashboards and action-item tracking.\n- **On-call participation resistance resolved**: <10% of engineers report 'unwilling to participate' in the 6-month pulse survey (baseline to be measured in S1)."}30 short, opinionated steps that time-box design to four weeks and aim to be live in ~10 weeks so five-plus months of SOC 2 evidence accumulate. Organises coverage as Layer A (24 trained company ICs, 24x7), Layer B (critical path 24x7), Layer C (business hours plus an EM escalation list), with mechanical rotation rules and a page-load SLO of p50 ≤4 per night shift. Maps contractual and regulatory notification clocks (S4) before setting severity triggers, runs a forced noise burn-down with a weekly leaderboard, and gates wave expansion on a red on-call pulse survey.
Closes with a mock auditor interview and a deliberate wording freeze for the observation window.
The plan produced
1. Secure executive mandate and budget
Get a written CEO/CTO mandate that incident command is a company process, not a team hobby.
The mandate must state that paid on-call is required for production ownership. "No pager for other teams' code" is solved by named ownership, not by refusing coverage.
- Approve budget for tooling, stipends, training, and a dedicated program lead for six months.
- Name an executive sponsor (CTO or VP Engineering) who will chair the weekly incident review.
- Tie the clock to SOC 2 Type II: the process must be live in about 10 weeks so ~6 months of evidence remain.
- Commit that the CEO will hear about outages from this process, not from customers.
2. Form the working group and decision rights (after 1)
Stand up a small group that can decide. Do not form a 28-team committee.
Core seats: SRE/platform lead, payments/ledger engineering manager, support lead, legal/compliance, HR, one rotating team EM, and a program manager.
- Meet twice a week for 10 weeks, then weekly.
- RACI: the group proposes; the sponsor decides in 48 hours; teams implement.
- Publish one Slack channel and one source-of-truth doc on day one.
- Time-box design to four weeks. Ship v1 rather than wait for consensus.
3. Inventory services, owners, and on-call gaps (after 1)
Build a living catalog of all ~180 services: owning team, criticality, current on-call, alert sources, and runbook link.
Walk the last 31 customer-impacting incidents and the two events where nobody was in charge. Record who detected, who led, time to mitigate, and which alerts fired.
- Tag each service as critical-path, customer-visible, or internal.
- List the 16 teams with no on-call and every orphan service with no owner.
- Map all six alerting tools and the 3,400 monthly alerts onto services.
- Flag the shared PostgreSQL ledger and two-region failover as named special cases.
4. Map regulatory and contractual notification duties (after 1)
Legal and compliance list every duty an incident can trigger. Do not invent clocks that violate a contract.
Cover SOC 2 CC7, customer MSA/SLA credit terms, money-transmitter rules, NYDFS 23 NYCRR 500 if applicable, PCI if in scope, and breach clocks.
- Extract notification timings from the largest customer contracts (status page, named AM, written notice).
- Define when legal, regulators, insurers, or the board must be told.
- Feed these clocks into severity triggers and the communications playbook.
5. Approve paid on-call and incident pay (after 1, 4)
Unpaid on-call is why 16 teams refuse the pager and why nights are uncovered. Fix the money before asking for coverage.
HR, legal, and finance design a New York–compliant package: weekly stipend for primary and secondary, extra stipend for company IC and comms, and after-hours incident pay or comp time.
- Treat exempt vs non-exempt staff explicitly under NY wage-hour rules.
- Put stipend in the next pay cycle after policy publish, not "later."
- Cap consecutive night weeks. Fund hiring if a team cannot rotate fairly (minimum six people for 24x7 primary plus secondary).
- Publish the package before any new rotation starts. This is the main answer to pager pushback.
6. Ratify four severity levels and their triggers (after 3, 4)
Adopt a business-impact scale. Engineers do not invent severity in the moment.
SEV-1: material payments failure; ledger down or inconsistent; security or customer-data incident; both regions impaired; or many customers already in SLA-credit territory.
SEV-2: degraded payments or a contracted feature down for multiple customers; SLA at risk.
SEV-3: narrow or single-customer impact with a workaround; no fleet-wide SLA risk.
SEV-4: no customer impact; ticket only.
- SEV-1 pages IC, comms, scribe, owning SMEs, and an exec; war room in 5 minutes; status page in 10; AM outreach in 20.
- SEV-2 pages IC and owning SMEs; comms may be the IC; status page in 15 minutes; updates every 30 minutes.
- SEV-3 pages the owning team only; customer notice only if that customer is affected.
- Anyone may declare. Only the IC may downgrade. When unsure, start high.
7. Define incident roles and their authority (after 6)
Four roles. Separate coordination from debugging so "who is in charge" cannot stall for an hour again.
- Incident Commander: owns severity, the room, the clock, and the next action. Does not write code. May page anyone, freeze deploys, and invoke failover. Staffed from a company-wide trained pool, not from the failing team.
- Communications Lead: status page, customers, AMs, execs, regulators. Speaks only from IC-approved facts.
- Scribe: timeline in the incident tool. Required for SEV-1 and SEV-2.
- SME responders: the owning team's on-call. They mitigate. They do not run the room.
Publish a one-page authority card. The IC stays in charge even if a VP joins.
8. Design 24x7 coverage without 28 night rotations (after 3, 5, 7)
Do not put 28 teams on 24x7. That is what engineers are rejecting.
Use three layers so people page for their own code, plus a trained commander.
- Layer A — company IC and SEV-1 comms: 24x7; about 24 trained people; week-long primary and secondary.
- Layer B — critical-path team on-call (ledger, payments processing, auth, API edge, platform/Kubernetes, data stores): 24x7 primary plus secondary.
- Layer C — all other teams: business-hours on-call; after hours the IC pages the team EM, who has a written escalation list.
Platform on-call is the safety net for unknown-owner pages, never the permanent owner. Every service must have a named team within 60 days or be scheduled to shut off.
9. Set rotation, handoff, and load rules (after 8)
Write mechanical rules so rotations are fair and load is visible.
Primary week, then secondary week, then at least two weeks off. No one holds primary on two rotations at once.
- Handoff is a 30-minute overlap covering open incidents, silenced alerts, and upcoming changes.
- Page-load SLO: p50 ≤ 4 pages per 12-hour night shift; p95 ≤ 10. A breach opens an alert-quality action.
- Require a shadow week before a first IC shift or a first critical-path rotation.
- Swaps live in the paging tool. Managers own coverage gaps, not the last person on the roster.
10. Write detection and escalation paths (after 6, 9)
Customers currently detect 40% of incidents and median time to detect is 22 minutes. That is the first failure mode.
Detection path: synthetic full-payment probes in both regions, SLO burn-rate alerts, support-to-incident intake, and one customer callback path that can create a SEV.
- A page must be acked in 5 minutes or it auto-escalates to secondary, then IC, then the EM, then the VP.
- Support may declare SEV-2 or higher without engineering permission.
- If ownership is unclear for 10 minutes, the IC keeps the incident and assigns a temporary owner. Never wait.
- An exec bridge auto-opens for every SEV-1 at T+15 minutes.
11. Write internal, customer, and regulator communications (after 4, 6, 7)
Stop "whoever is around" from writing the status page. Comms follow the clock, not convenience.
Timings from declaration:
- Internal war room: immediate. Exec summary for SEV-1/2 at 15 minutes, then every 30 minutes.
- Public status page: SEV-1 in 10 minutes, SEV-2 in 15. Updates at least every 30 minutes until resolve. Templates only. No speculation.
- Account managers get an affected-customer list and a script at T+20 minutes for SEV-1/2.
- Resolve notice and credit assessment within one business day.
- Comms pages legal on SEV-1 security, ledger integrity, or any outage that will breach contractual notice. Legal owns outbound regulatory letters; the IC owns facts.
12. Standardize blameless postmortems and action tracking (after 6)
A written postmortem is mandatory for every SEV-1 and SEV-2 within 5 business days. SEV-3 if the IC or EM requests it.
Use one template: timeline, customer impact (volume, duration, credits), detection gap, what went well, what did not, process-focused five whys, and numbered actions with owner and due date.
- Review is blameless and scheduled. The IC attends. The exec sponsor reads every SEV-1.
- Actions live in one tracker, not in the doc. No action without an owner and a date. Default due date 14 days; 30 days max unless architecture work with a milestone.
- Close rate is a published metric. The old 11-of-64 pattern is a process failure.
13. Set alert quality rules that make paging acceptable (after 3, 6)
3,400 alerts a month and 85% noise is why on-call feels like punishment. Pages are a product with a quality bar.
A page (not a ticket) must map to a customer-facing SLO or a hard dependency of one. It must have an owner team, a runbook link, and a default severity. It must be actionable at 3 a.m. by the person who is paged.
- Ban parallel paging from six tools. One paging policy: symptom-based; burn-rate preferred over raw thresholds.
- Every team gets a monthly noise budget. Exceeding it is a sprint task, not heroics.
- A human may silence a flapping alert only with a linked ticket.
14. Publish Incident Management Policy v1 (after 5, 6, 7, 8, 9, 10, 11, 12, 13)
Collapse the design into a short policy people will open during an outage.
Ten pages or fewer, plus one-page cards for severity, roles, and comms timings. Host it where the incident tool can link it.
- Include the compensation summary and the rule: you are not on-call for other teams' services.
- Version it. v1 is mandatory from the pilot start date.
- Legal, HR, and the exec sponsor sign. Announce in all-hands, not only in Slack.
15. Implement a single incident command tool (after 7, 10, 11)
Put one tool in the path that creates the room, pages roles from severity, records the timeline, and prompts status-page updates.
Requirements: Slack (or equivalent) incident bot, severity in one click, role assignment, stakeholder groups, and timeline export for postmortems and auditors.
- Integrate with the pager so IC, comms, and SME pages are automatic.
- Retain artifacts at least one year for SOC 2.
- Ad-hoc Zoom/Slack threads are no longer the system of record.
16. Consolidate six alerting tools onto one pager (after 9, 13)
Pick one paging product. Connect existing monitors to it. Migrate pages first, tickets second.
- Inventory every page-producing rule. Delete or downgrade the noisy majority in S23.
- Route by service label → owning team schedule → the escalation policy from S10.
- IC and comms schedules live in the same product.
- Set a hard date after which pages outside the chosen tool are not valid on-call obligations.
17. Operationalize the status page and AM path (after 11, 15)
Put the status page behind the Comms Lead role. Use templates for investigating, identified, mitigating, and resolved.
Subscribe AMs and customers whose contracts require it. Generate the affected-customer list from the incident (tenant, region, payment method).
- Dry-run a SEV-2 update before the pilot goes live.
- Record every public update in the incident timeline for the audit.
- Define partial vs full-outage wording so impact cannot be understated.
18. Stand up one postmortem repo and action board (after 12)
Create the template, the filing location, and one Jira/Linear board with states: open, in progress, blocked, done, and won't-do (with exec reason).
Wire the incident tool so a SEV-1/2 automatically opens a draft postmortem and action tickets.
- Each week the program manager reports actions past due to the exec sponsor.
- If it is not in the board, it does not exist.
19. Assign owners and write critical-path runbooks (after 3, 9)
Close the ownership gaps that cause "pager for other people's code."
Every production service gets a team in the catalog. Unowned services get an owner in 30 days or a decommission date.
- Write runbooks for ledger Postgres, regional failover, payments API, auth, and the Kubernetes control plane: symptoms, dashboards, mitigate vs escalate, customer impact.
- Link runbooks from alerts. If there is no runbook, the alert cannot page at night unless the EM accepts the gap in writing.
20. Detect payments failures before customers do (after 3, 13)
Build or finish end-to-end synthetics: create payment, ledger write, webhook, both regions, both critical payment methods.
Alert on SLO burn, not on a single 500. Page SEV-2 or SEV-1 from these probes. This is the fastest lever on the 22-minute MTTD and the 40% customer-detected rate.
- Add ledger lag, replication, disk, and failover-readiness as first-class pages to the ledger team.
- Every postmortem asks first: why did a customer see this first?
21. Train the first cadre of ICs, comms leads, and scribes (after 7, 14, 15)
Train about 24 ICs and 12 comms leads before the pilot. Classroom plus a recorded shadow of a simulated SEV-1.
Curriculum: severity, authority card, tool, comms timings, when to call legal, how to run a room of 20, how to hand off at 2 a.m.
- Certification: pass a tabletop. No certificate, no rotation.
- Recertify yearly and after any SEV-1 where process failed.
- Managers of ICs protect calendar time. This is part of the job.
22. Pilot on the payments critical path for six weeks (after 5, 14, 15, 16, 17, 18, 19, 21)
Go live with policy, tool, paid rotations, and IC coverage for ledger, payments, API, platform, and support intake.
Keep old paths as backup for one week, then cut over. Real incidents use the new process only.
- Staff the program lead in every SEV-2+ as coach, not as secret IC.
- Collect friction daily. Fix tooling and wording in 48 hours.
- Expansion gate: named IC in under 5 minutes, first status update on time, no unpaid pages, postmortem filed.
23. Cut alert noise with a forced burn-down (after 13, 16)
Give every team a numbered list of their noisiest alerts. Move noise from 85% to under 15%, and monthly pages from 3,400 toward 500.
Each sprint, critical-path teams must delete, debounce, or convert to ticket a fixed quota. Platform provides burn-rate and grouping libraries.
- Publish a weekly noise leaderboard. Shame systems, not people.
- After eight weeks, any page without a runbook or with >30% false pages in 14 days is auto-downgraded until fixed.
24. Roll remaining teams onto the model by risk (after 22)
After the pilot gate, add teams in waves of four to five every two weeks. Highest customer-impact first.
Layer C teams get business-hours schedules and the EM night list. Do not surprise anyone with a pager.
- Each wave: ownership confirmed, alerts routed, runbooks for paging alerts, paid rotation in HR, one tabletop.
- Finish all 28 teams at least five months before the SOC 2 report date so the observation window covers the company.
- Orphan services still unowned at wave end are escalated to the sponsor for shutdown or reassignment.
25. Run tabletops and multi-region game days (after 21, 22)
Schedule a monthly tabletop: SEV-1 ledger, SEV-1 region loss, SEV-2 degraded payments, customer-detected incident, and a "who is in charge" chaos drill.
Quarterly game day: fail a region or a ledger replica in staging or a controlled production drill.
- Include support, AMs, legal, and an exec. Process fails if only engineers show up.
- Capture actions on the same board as real postmortems.
- Use results as SOC 2 evidence that IR is tested.
26. Resolve ownership fights and pager culture (after 5, 8, 22)
Treat pushback as design input, not defiance. Repeat the contract in office hours: you carry a pager for your services; IC is coordination; nights are paid; noise is a defect.
- EMs who cannot staff a fair rotation get headcount or have services reassigned. Do not run two-person 24x7.
- Publicly close the two historical "nobody in charge" incidents with what would be different now.
- Pulse-survey on-call at 60 and 120 days. If load or fairness is red, stop expansion until fixed.
27. Launch metrics, weekly review, and error budgets (after 15, 22)
The CEO email exists because there was no operating rhythm. Stand up a dashboard and use it.
Track incident count by SEV, MTTD, MTTA, MTTM, customer-detected %, time-to-IC, time-to-status, pages per shift, noise %, SLA credits, postmortem SLA, and action close rate.
- Weekly 30-minute review chaired by the exec sponsor. Last week's incidents, open actions, alert SLO breaches. Mandatory for EMs of teams that had SEV-2+.
- Error budgets on critical SLOs: burn too fast and the team stops features and pays down reliability.
- Monthly report to the CEO so customer-first news does not recur.
28. Collect SOC 2 Type II evidence as a side effect (after 4, 14, 22, 27)
Do not create a parallel "audit process." The real process is the evidence.
From pilot start, retain policy versions, on-call schedules, incident records, status-page history, postmortems, action tickets, training attendance, tabletop notes, and access reviews of the incident and pager tools.
- Map each artifact to SOC 2 CC7.2–CC7.5. Compliance owns the matrix; engineering owns completeness.
- Internal audit or a consultant runs readiness reviews at week 12 and week 20, not the week before the auditor.
29. Inspect and adapt the process after 90 days (after 22, 27)
After 90 days live, revise v1 using data, not opinions.
Look at MTTD/MTTM, uncovered shifts, IC burnout, missed status updates, action close rate, and survey results.
- Change severity definitions if teams inflate or deflate them.
- Adjust Layer B vs C membership from real pages.
- Cut steps nobody follows. Add only what the last 90 days proved missing.
- Freeze v2 as the SOC 2 described process unless a defect is found.
30. Rehearse the audit and freeze the observation window (after 24, 28, 29)
Six to eight weeks before the auditor, run a mock interview with an IC, a random engineer, support, and compliance.
Show one SEV-1 and one SEV-2 end-to-end: detect, page, command, comms, postmortem, closed actions.
- Fix evidence gaps immediately.
- Freeze process wording for the rest of the observation window; log exceptions.
- Brief the CEO and customer success with the metrics so the "outages we hear about from clients" story is retired.
- Median time to detect customer-impacting incidents ≤ 5 minutes within 6 months of go-live.
- Share of SEV-1/SEV-2 incidents first detected by customers ≤ 5% (from 40%).
- Median time to mitigate SEV-1/SEV-2 ≤ 45 minutes (from 3 h 10 min).
- Named Incident Commander assigned within 5 minutes for ≥ 95% of SEV-1/SEV-2.
- First status-page update within policy time for ≥ 95% of SEV-1/SEV-2.
- SLA credits down ≥ 80% versus the trailing $1.3M within 12 months.
- Paging volume ≤ 500 per month and noise ≤ 15% (from 3,400 and 85%).
- 100% of production services have a named owning team and a paging policy.
- Postmortems filed within 5 business days for 100% of SEV-1/SEV-2; action-item close rate ≥ 80% within 30 days.
- 24×7 IC and critical-path coverage with zero unfilled shifts per quarter.
- Paid on-call live for every rotation before that rotation pages humans.
- SOC 2 Type II incident-response controls evidenced for ≥ 5 months before the auditor's report.
- On-call pulse: ≥ 70% of engineers agree rotations are fair and limited to their services.
[SYSTEM] You are an expert assistant in complex project planning. Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform. Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible. Take as much time and space as you need to think through all aspects of the problem. After your thorough analysis, answer with the plan in the requested structure. Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text. [HUMAN] Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code". Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit." Generate an initial, detailed, and well-structured plan proposal. The plan has these parts: - "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none). - "estimated_complexity": "low", "medium" or "high". - "success_metrics": clear and measurable success metrics, one per line as a bullet list. Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Secure executive mandate and budget", "description": "Get a written CEO/CTO mandate that incident command is a company process, not a team hobby.\n\nThe mandate must state that **paid on-call** is required for production ownership. \"No pager for other teams' code\" is solved by named ownership, not by refusing coverage.\n\n- Approve budget for tooling, stipends, training, and a dedicated program lead for six months.\n- Name an executive sponsor (CTO or VP Engineering) who will chair the weekly incident review.\n- Tie the clock to SOC 2 Type II: the process must be live in about 10 weeks so ~6 months of evidence remain.\n- Commit that the CEO will hear about outages from this process, not from customers.", "dependencies": []}, {"step_id": "S2", "title": "Form the working group and decision rights", "description": "Stand up a small group that can decide. Do not form a 28-team committee.\n\n**Core seats:** SRE/platform lead, payments/ledger engineering manager, support lead, legal/compliance, HR, one rotating team EM, and a program manager.\n\n- Meet twice a week for 10 weeks, then weekly.\n- RACI: the group proposes; the sponsor decides in 48 hours; teams implement.\n- Publish one Slack channel and one source-of-truth doc on day one.\n- Time-box design to four weeks. Ship v1 rather than wait for consensus.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Inventory services, owners, and on-call gaps", "description": "Build a living catalog of all ~180 services: owning team, criticality, current on-call, alert sources, and runbook link.\n\nWalk the last 31 customer-impacting incidents and the two events where **nobody was in charge**. Record who detected, who led, time to mitigate, and which alerts fired.\n\n- Tag each service as critical-path, customer-visible, or internal.\n- List the 16 teams with no on-call and every orphan service with no owner.\n- Map all six alerting tools and the 3,400 monthly alerts onto services.\n- Flag the shared PostgreSQL ledger and two-region failover as named special cases.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Map regulatory and contractual notification duties", "description": "Legal and compliance list every duty an incident can trigger. Do not invent clocks that violate a contract.\n\nCover SOC 2 CC7, customer MSA/SLA credit terms, money-transmitter rules, NYDFS 23 NYCRR 500 if applicable, PCI if in scope, and breach clocks.\n\n- Extract **notification timings** from the largest customer contracts (status page, named AM, written notice).\n- Define when legal, regulators, insurers, or the board must be told.\n- Feed these clocks into severity triggers and the communications playbook.", "dependencies": ["S1"]}, {"step_id": "S5", "title": "Approve paid on-call and incident pay", "description": "Unpaid on-call is why 16 teams refuse the pager and why nights are uncovered. Fix the money before asking for coverage.\n\nHR, legal, and finance design a New York–compliant package: weekly stipend for primary and secondary, extra stipend for company IC and comms, and **after-hours incident pay** or comp time.\n\n- Treat exempt vs non-exempt staff explicitly under NY wage-hour rules.\n- Put stipend in the next pay cycle after policy publish, not \"later.\"\n- Cap consecutive night weeks. Fund hiring if a team cannot rotate fairly (minimum six people for 24x7 primary plus secondary).\n- Publish the package before any new rotation starts. This is the main answer to pager pushback.", "dependencies": ["S1", "S4"]}, {"step_id": "S6", "title": "Ratify four severity levels and their triggers", "description": "Adopt a business-impact scale. Engineers do not invent severity in the moment.\n\n**SEV-1:** material payments failure; ledger down or inconsistent; security or customer-data incident; both regions impaired; or many customers already in SLA-credit territory.\n\n**SEV-2:** degraded payments or a contracted feature down for multiple customers; SLA at risk.\n\n**SEV-3:** narrow or single-customer impact with a workaround; no fleet-wide SLA risk.\n\n**SEV-4:** no customer impact; ticket only.\n\n- SEV-1 pages IC, comms, scribe, owning SMEs, and an exec; war room in 5 minutes; status page in 10; AM outreach in 20.\n- SEV-2 pages IC and owning SMEs; comms may be the IC; status page in 15 minutes; updates every 30 minutes.\n- SEV-3 pages the owning team only; customer notice only if that customer is affected.\n- Anyone may declare. Only the IC may downgrade. When unsure, start high.", "dependencies": ["S3", "S4"]}, {"step_id": "S7", "title": "Define incident roles and their authority", "description": "Four roles. Separate coordination from debugging so \"who is in charge\" cannot stall for an hour again.\n\n- **Incident Commander:** owns severity, the room, the clock, and the next action. Does not write code. May page anyone, freeze deploys, and invoke failover. Staffed from a company-wide trained pool, not from the failing team.\n- **Communications Lead:** status page, customers, AMs, execs, regulators. Speaks only from IC-approved facts.\n- **Scribe:** timeline in the incident tool. Required for SEV-1 and SEV-2.\n- **SME responders:** the owning team's on-call. They mitigate. They do not run the room.\n\nPublish a one-page authority card. The IC stays in charge even if a VP joins.", "dependencies": ["S6"]}, {"step_id": "S8", "title": "Design 24x7 coverage without 28 night rotations", "description": "Do not put 28 teams on 24x7. That is what engineers are rejecting.\n\nUse **three layers** so people page for their own code, plus a trained commander.\n\n- Layer A — company IC and SEV-1 comms: 24x7; about 24 trained people; week-long primary and secondary.\n- Layer B — critical-path team on-call (ledger, payments processing, auth, API edge, platform/Kubernetes, data stores): 24x7 primary plus secondary.\n- Layer C — all other teams: business-hours on-call; after hours the IC pages the team EM, who has a written escalation list.\n\nPlatform on-call is the safety net for unknown-owner pages, never the permanent owner. Every service must have a named team within 60 days or be scheduled to shut off.", "dependencies": ["S3", "S5", "S7"]}, {"step_id": "S9", "title": "Set rotation, handoff, and load rules", "description": "Write mechanical rules so rotations are fair and load is visible.\n\nPrimary week, then secondary week, then at least two weeks off. No one holds primary on two rotations at once.\n\n- Handoff is a 30-minute overlap covering open incidents, silenced alerts, and upcoming changes.\n- Page-load SLO: p50 ≤ 4 pages per 12-hour night shift; p95 ≤ 10. A breach opens an alert-quality action.\n- Require a shadow week before a first IC shift or a first critical-path rotation.\n- Swaps live in the paging tool. Managers own coverage gaps, not the last person on the roster.", "dependencies": ["S8"]}, {"step_id": "S10", "title": "Write detection and escalation paths", "description": "Customers currently detect 40% of incidents and median time to detect is 22 minutes. That is the first failure mode.\n\nDetection path: synthetic full-payment probes in both regions, SLO burn-rate alerts, support-to-incident intake, and one customer callback path that can create a SEV.\n\n- A page must be acked in **5 minutes** or it auto-escalates to secondary, then IC, then the EM, then the VP.\n- Support may declare SEV-2 or higher without engineering permission.\n- If ownership is unclear for 10 minutes, the IC keeps the incident and assigns a temporary owner. Never wait.\n- An exec bridge auto-opens for every SEV-1 at T+15 minutes.", "dependencies": ["S6", "S9"]}, {"step_id": "S11", "title": "Write internal, customer, and regulator communications", "description": "Stop \"whoever is around\" from writing the status page. Comms follow the clock, not convenience.\n\nTimings from declaration:\n\n- Internal war room: immediate. Exec summary for SEV-1/2 at 15 minutes, then every 30 minutes.\n- **Public status page:** SEV-1 in 10 minutes, SEV-2 in 15. Updates at least every 30 minutes until resolve. Templates only. No speculation.\n- Account managers get an affected-customer list and a script at T+20 minutes for SEV-1/2.\n- Resolve notice and credit assessment within one business day.\n- Comms pages legal on SEV-1 security, ledger integrity, or any outage that will breach contractual notice. Legal owns outbound regulatory letters; the IC owns facts.", "dependencies": ["S4", "S6", "S7"]}, {"step_id": "S12", "title": "Standardize blameless postmortems and action tracking", "description": "A written postmortem is mandatory for every SEV-1 and SEV-2 within 5 business days. SEV-3 if the IC or EM requests it.\n\nUse one template: timeline, customer impact (volume, duration, credits), detection gap, what went well, what did not, process-focused five whys, and numbered actions with owner and due date.\n\n- Review is **blameless** and scheduled. The IC attends. The exec sponsor reads every SEV-1.\n- Actions live in one tracker, not in the doc. No action without an owner and a date. Default due date 14 days; 30 days max unless architecture work with a milestone.\n- Close rate is a published metric. The old 11-of-64 pattern is a process failure.", "dependencies": ["S6"]}, {"step_id": "S13", "title": "Set alert quality rules that make paging acceptable", "description": "3,400 alerts a month and 85% noise is why on-call feels like punishment. Pages are a product with a quality bar.\n\nA page (not a ticket) must map to a customer-facing SLO or a hard dependency of one. It must have an owner team, a runbook link, and a default severity. It must be actionable at 3 a.m. by the person who is paged.\n\n- Ban parallel paging from six tools. One paging policy: symptom-based; burn-rate preferred over raw thresholds.\n- Every team gets a monthly noise budget. Exceeding it is a sprint task, not heroics.\n- A human may silence a flapping alert only with a linked ticket.", "dependencies": ["S3", "S6"]}, {"step_id": "S14", "title": "Publish Incident Management Policy v1", "description": "Collapse the design into a short policy people will open during an outage.\n\nTen pages or fewer, plus one-page cards for severity, roles, and comms timings. Host it where the incident tool can link it.\n\n- Include the compensation summary and the rule: you are **not on-call for other teams' services**.\n- Version it. v1 is mandatory from the pilot start date.\n- Legal, HR, and the exec sponsor sign. Announce in all-hands, not only in Slack.", "dependencies": ["S5", "S6", "S7", "S8", "S9", "S10", "S11", "S12", "S13"]}, {"step_id": "S15", "title": "Implement a single incident command tool", "description": "Put one tool in the path that creates the room, pages roles from severity, records the timeline, and prompts status-page updates.\n\nRequirements: Slack (or equivalent) incident bot, severity in one click, role assignment, stakeholder groups, and timeline export for postmortems and auditors.\n\n- Integrate with the pager so IC, comms, and SME pages are automatic.\n- Retain artifacts at least one year for SOC 2.\n- Ad-hoc Zoom/Slack threads are no longer the system of record.", "dependencies": ["S7", "S10", "S11"]}, {"step_id": "S16", "title": "Consolidate six alerting tools onto one pager", "description": "Pick one paging product. Connect existing monitors to it. Migrate **pages** first, tickets second.\n\n- Inventory every page-producing rule. Delete or downgrade the noisy majority in S23.\n- Route by service label → owning team schedule → the escalation policy from S10.\n- IC and comms schedules live in the same product.\n- Set a hard date after which pages outside the chosen tool are not valid on-call obligations.", "dependencies": ["S9", "S13"]}, {"step_id": "S17", "title": "Operationalize the status page and AM path", "description": "Put the status page behind the Comms Lead role. Use templates for investigating, identified, mitigating, and resolved.\n\nSubscribe AMs and customers whose contracts require it. Generate the affected-customer list from the incident (tenant, region, payment method).\n\n- Dry-run a SEV-2 update before the pilot goes live.\n- Record every public update in the incident timeline for the audit.\n- Define partial vs full-outage wording so impact cannot be understated.", "dependencies": ["S11", "S15"]}, {"step_id": "S18", "title": "Stand up one postmortem repo and action board", "description": "Create the template, the filing location, and one Jira/Linear board with states: open, in progress, blocked, done, and won't-do (with exec reason).\n\nWire the incident tool so a SEV-1/2 automatically opens a draft postmortem and action tickets.\n\n- Each week the program manager reports actions past due to the exec sponsor.\n- If it is not in the board, it does not exist.", "dependencies": ["S12"]}, {"step_id": "S19", "title": "Assign owners and write critical-path runbooks", "description": "Close the ownership gaps that cause \"pager for other people's code.\"\n\nEvery production service gets a team in the catalog. Unowned services get an owner in 30 days or a decommission date.\n\n- Write runbooks for ledger Postgres, regional failover, payments API, auth, and the Kubernetes control plane: symptoms, dashboards, mitigate vs escalate, customer impact.\n- Link runbooks from alerts. If there is no runbook, the alert cannot page at night unless the EM accepts the gap in writing.", "dependencies": ["S3", "S9"]}, {"step_id": "S20", "title": "Detect payments failures before customers do", "description": "Build or finish end-to-end synthetics: create payment, ledger write, webhook, both regions, both critical payment methods.\n\nAlert on SLO burn, not on a single 500. Page SEV-2 or SEV-1 from these probes. This is the fastest lever on the 22-minute MTTD and the 40% customer-detected rate.\n\n- Add ledger lag, replication, disk, and failover-readiness as first-class pages to the ledger team.\n- Every postmortem asks first: **why did a customer see this first?**", "dependencies": ["S3", "S13"]}, {"step_id": "S21", "title": "Train the first cadre of ICs, comms leads, and scribes", "description": "Train about 24 ICs and 12 comms leads before the pilot. Classroom plus a recorded shadow of a simulated SEV-1.\n\nCurriculum: severity, authority card, tool, comms timings, when to call legal, how to run a room of 20, how to hand off at 2 a.m.\n\n- Certification: pass a tabletop. No certificate, no rotation.\n- Recertify yearly and after any SEV-1 where process failed.\n- Managers of ICs protect calendar time. This is part of the job.", "dependencies": ["S7", "S14", "S15"]}, {"step_id": "S22", "title": "Pilot on the payments critical path for six weeks", "description": "Go live with policy, tool, paid rotations, and IC coverage for ledger, payments, API, platform, and support intake.\n\nKeep old paths as backup for one week, then cut over. Real incidents use the new process only.\n\n- Staff the program lead in every SEV-2+ as coach, not as secret IC.\n- Collect friction daily. Fix tooling and wording in 48 hours.\n- Expansion gate: named IC in under 5 minutes, first status update on time, no unpaid pages, postmortem filed.", "dependencies": ["S5", "S14", "S15", "S16", "S17", "S18", "S19", "S21"]}, {"step_id": "S23", "title": "Cut alert noise with a forced burn-down", "description": "Give every team a numbered list of their noisiest alerts. Move noise from 85% to **under 15%**, and monthly pages from 3,400 toward 500.\n\nEach sprint, critical-path teams must delete, debounce, or convert to ticket a fixed quota. Platform provides burn-rate and grouping libraries.\n\n- Publish a weekly noise leaderboard. Shame systems, not people.\n- After eight weeks, any page without a runbook or with >30% false pages in 14 days is auto-downgraded until fixed.", "dependencies": ["S13", "S16"]}, {"step_id": "S24", "title": "Roll remaining teams onto the model by risk", "description": "After the pilot gate, add teams in waves of four to five every two weeks. Highest customer-impact first.\n\nLayer C teams get business-hours schedules and the EM night list. Do not surprise anyone with a pager.\n\n- Each wave: ownership confirmed, alerts routed, runbooks for paging alerts, paid rotation in HR, one tabletop.\n- Finish all 28 teams at least five months before the SOC 2 report date so the observation window covers the company.\n- Orphan services still unowned at wave end are escalated to the sponsor for shutdown or reassignment.", "dependencies": ["S22"]}, {"step_id": "S25", "title": "Run tabletops and multi-region game days", "description": "Schedule a monthly tabletop: SEV-1 ledger, SEV-1 region loss, SEV-2 degraded payments, customer-detected incident, and a \"who is in charge\" chaos drill.\n\nQuarterly game day: fail a region or a ledger replica in staging or a controlled production drill.\n\n- Include support, AMs, legal, and an exec. Process fails if only engineers show up.\n- Capture actions on the same board as real postmortems.\n- Use results as SOC 2 evidence that IR is tested.", "dependencies": ["S21", "S22"]}, {"step_id": "S26", "title": "Resolve ownership fights and pager culture", "description": "Treat pushback as design input, not defiance. Repeat the contract in office hours: you carry a pager for **your** services; IC is coordination; nights are paid; noise is a defect.\n\n- EMs who cannot staff a fair rotation get headcount or have services reassigned. Do not run two-person 24x7.\n- Publicly close the two historical \"nobody in charge\" incidents with what would be different now.\n- Pulse-survey on-call at 60 and 120 days. If load or fairness is red, stop expansion until fixed.", "dependencies": ["S5", "S8", "S22"]}, {"step_id": "S27", "title": "Launch metrics, weekly review, and error budgets", "description": "The CEO email exists because there was no operating rhythm. Stand up a dashboard and use it.\n\nTrack incident count by SEV, MTTD, MTTA, MTTM, customer-detected %, time-to-IC, time-to-status, pages per shift, noise %, SLA credits, postmortem SLA, and action close rate.\n\n- **Weekly 30-minute review** chaired by the exec sponsor. Last week's incidents, open actions, alert SLO breaches. Mandatory for EMs of teams that had SEV-2+.\n- Error budgets on critical SLOs: burn too fast and the team stops features and pays down reliability.\n- Monthly report to the CEO so customer-first news does not recur.", "dependencies": ["S15", "S22"]}, {"step_id": "S28", "title": "Collect SOC 2 Type II evidence as a side effect", "description": "Do not create a parallel \"audit process.\" The real process is the evidence.\n\nFrom pilot start, retain policy versions, on-call schedules, incident records, status-page history, postmortems, action tickets, training attendance, tabletop notes, and access reviews of the incident and pager tools.\n\n- Map each artifact to SOC 2 CC7.2–CC7.5. Compliance owns the matrix; engineering owns completeness.\n- Internal audit or a consultant runs readiness reviews at week 12 and week 20, not the week before the auditor.", "dependencies": ["S4", "S14", "S22", "S27"]}, {"step_id": "S29", "title": "Inspect and adapt the process after 90 days", "description": "After 90 days live, revise v1 using data, not opinions.\n\nLook at MTTD/MTTM, uncovered shifts, IC burnout, missed status updates, action close rate, and survey results.\n\n- Change severity definitions if teams inflate or deflate them.\n- Adjust Layer B vs C membership from real pages.\n- Cut steps nobody follows. Add only what the last 90 days proved missing.\n- Freeze v2 as the SOC 2 described process unless a defect is found.", "dependencies": ["S22", "S27"]}, {"step_id": "S30", "title": "Rehearse the audit and freeze the observation window", "description": "Six to eight weeks before the auditor, run a mock interview with an IC, a random engineer, support, and compliance.\n\nShow one SEV-1 and one SEV-2 end-to-end: detect, page, command, comms, postmortem, closed actions.\n\n- Fix evidence gaps immediately.\n- Freeze process wording for the rest of the observation window; log exceptions.\n- Brief the CEO and customer success with the metrics so the \"outages we hear about from clients\" story is retired.", "dependencies": ["S24", "S28", "S29"]}], "estimated_complexity": "high", "success_metrics": "- Median time to detect customer-impacting incidents ≤ 5 minutes within 6 months of go-live.\n- Share of SEV-1/SEV-2 incidents first detected by customers ≤ 5% (from 40%).\n- Median time to mitigate SEV-1/SEV-2 ≤ 45 minutes (from 3 h 10 min).\n- Named Incident Commander assigned within 5 minutes for ≥ 95% of SEV-1/SEV-2.\n- First status-page update within policy time for ≥ 95% of SEV-1/SEV-2.\n- SLA credits down ≥ 80% versus the trailing $1.3M within 12 months.\n- Paging volume ≤ 500 per month and noise ≤ 15% (from 3,400 and 85%).\n- 100% of production services have a named owning team and a paging policy.\n- Postmortems filed within 5 business days for 100% of SEV-1/SEV-2; action-item close rate ≥ 80% within 30 days.\n- 24×7 IC and critical-path coverage with zero unfilled shifts per quarter.\n- Paid on-call live for every rotation before that rotation pages humans.\n- SOC 2 Type II incident-response controls evidenced for ≥ 5 months before the auditor's report.\n- On-call pulse: ≥ 70% of engineers agree rotations are fair and limited to their services."}14 conventional steps covering the full brief: SEV1–SEV4, RACI roles, service on-call plus a central IC guild, stipends and TOIL, one paging platform, comms timings, mandatory postmortems, metrics and a month-6 readiness assessment. Correct in outline but the thinnest on specifics — no dates until months 6–8, no wave plan, no page budget, no risk register, and no targeted answer to the "other teams' code" objection beyond a survey. Two realism problems: a 5-minute SEV1 status-page post, and a 24x7 model built on "follow-the-sun between the two AWS regions" for a single New York organisation.
The pilot lands at step 11 with training and full rollout after it, leaving little room to iterate before the audit.
The plan produced
1. Baseline current incident response and align stakeholders
Collect data from the last 12 months of incidents, all six alert tools, on-call practices, and team interviews. Identify gaps against the target incident management process and secure executive sponsorship.
- Gather incident timeline, detection source, mitigation time, customer impact, SLA credit, and postmortem status for all 31 incidents.
- Survey 28 teams on on-call burden, alert quality, and operational pain.
- Map current tools, escalation paths, and communication workflows.
- Create baseline metrics and a stakeholder map with executive sponsor and audit owner.
2. Define severity levels and response triggers (after 1)
Define a four-level severity scale with objective business-impact criteria so any engineer can classify an incident consistently.
- SEV1: widespread transaction processing outage, data breach, security incident, or severe SLA breach; triggers full incident command, executive notification, and a 5-minute status page update.
- SEV2: major feature outage, significant degradation without workaround, or customer financial risk; triggers incident commander, full communications role, and status page updates.
- SEV3: partial impairment with workaround or limited customer impact; triggers on-call response, internal communication, and optional status page update.
- SEV4: minor or internal issue, no customer impact; handled during business hours through ticketing.
- Include an escalation matrix showing who can declare, downgrade, and invoke regulatory or legal involvement.
3. Define incident roles and decision authority (after 1, 2)
Define incident roles, responsibilities, and decision authority using RACI to remove ambiguity about who is in charge.
- Incident Commander: owns the incident, declares severity, and coordinates resolution.
- Communications Lead: owns internal and external messaging, status page updates, and account manager notifications.
- Scribe: maintains timeline, incident log, and postmortem notes.
- Subject-matter responders: diagnose and fix the incident; may come from multiple teams.
- Executive sponsor: optional for SEV1; customer liaison: handles account managers.
- Define decision rights for severity declaration, escalation, rollback, customer communications, and incident closure.
4. Design 24x7 staffing model across 28 teams (after 3)
Design 24x7 coverage across 28 teams without overloading engineers. Use service-based on-call plus a central incident command pool.
- Each service or domain team assigns primary and secondary on-call for its own services.
- Create central incident commander, communications, and scribe rotations staffed from a trained incident response guild across all teams; use follow-the-sun between the two AWS regions and time zones.
- Define escalation layers: service on-call to team lead or manager to service owner to executive.
- Define handoff times, shadow shifts, and load balancing; target at most one week of on-call per engineer per month.
- Bridge the current 12-team paid on-call to 28-team paid coverage; no team remains uncovered.
5. Define on-call rotations, compensation, and alert quality rules (after 4)
Define sustainable rotations, pay, and rules that eliminate noisy pages.
- Rotations: weekly or biweekly, at least one primary and one secondary, with 12-hour shifts where possible or 24-hour for low-volume services.
- Compensation: monthly on-call stipend for all on-call engineers, additional incident response bonus for after-hours work, and time off in lieu; align with market rates.
- Alert quality rules: every page must be actionable, have a runbook link, specify a service owner, include severity, and be based on SLO burn or known failure signals; no dashboard-only alerts.
- Noise budget: reject or downgrade non-actionable alerts; all pages must go to on-call only after suppression and deduplication.
- Weekly alert review removes the top noisy alerts.
6. Design detection, escalation, and alert routing (after 2, 4, 5)
Define how incidents are detected, routed, and escalated so nothing waits on a human to notice.
- Consolidate the six alert tools into one alerting and paging platform with routing by service, severity, and tags.
- Detection sources: infrastructure metrics, application synthetic transactions, log-based anomalies, business transaction SLI monitoring, and customer-reported issues through support or account managers.
- Routing: alert is paged to service on-call within 30 seconds; primary must acknowledge within 5 minutes; if no ack, page secondary then on-call manager.
- Escalation timeouts: unresolved SEV1 escalates to service owner at 15 minutes and to leadership at 30 minutes; any engineer can escalate to the incident commander.
- Define customer-reported incident intake and classification in the same tool.
7. Define internal and external communication protocols (after 2, 3)
Define communication channels, templates, and timing for internal, customer, and regulator audiences.
- Internal: dedicated incident Slack channel, internal status page mirror, and war room bridge for SEV1; incident commander and communications lead own these channels.
- Status page: SEV1 post within 5 minutes, updates every 30 minutes or on material change, resolution within 60 minutes of mitigation; SEV2 post within 15 minutes, updates hourly; SEV3 optional.
- Account managers: SEV1 and SEV2 notify account managers within 15 minutes with an approved customer-facing description and expected impact.
- Regulators: legal or compliance determines notification for data breaches, security incidents, funds availability issues, or regulatory reportable events; criteria and timing follow legal and regulatory requirements; communications lead coordinates.
- Use pre-approved message templates and an approval chain; no ad-hoc wording.
8. Define postmortem policy and action tracking (after 3)
Define mandatory blameless postmortems and action tracking.
- Mandatory for all SEV1 and SEV2 incidents, and any SEV3 that breaches SLA or is customer-detected.
- Format: impact, timeline, root causes, contributing factors, detection and response gaps, what worked well, and action items.
- Blameless: focus on system and process causes, not individual blame; use trained facilitators.
- Ownership: each action has an owner, due date, and tracking ID in a single backlog.
- Review postmortems at the weekly incident review; track action closure; expect 100% completion.
- Complete postmortems within 5 business days for SEV1 and SEV2 incidents.
9. Define metrics, dashboards, and review cadence (after 2, 3, 8)
Define metrics and review cadence to measure process health.
- Metrics: MTTD, MTTM, customer detected percentage, alert noise percentage, on-call response time, on-call load, SLA credits paid, and postmortem action completion.
- Dashboards: real-time operational dashboard for on-call engineers and management.
- Weekly incident review: review all SEV1 and SEV2 incidents, action items, and noisy alerts.
- Monthly trends with leadership; quarterly review against SLOs and audit controls.
- Success thresholds: MTTD under 5 minutes, MTTM under 60 minutes for SEV1, customer detected under 15%, and alert noise under 10%.
10. Configure incident tooling and integrations (after 5, 6, 7, 8, 9)
Implement and integrate the tools that automate the defined process.
- Aggregate alerts from the existing six tools into PagerDuty, Opsgenie, or a similar platform.
- Configure on-call schedules, escalation policies, and paging targeted at service owners.
- Integrate status page API for automated or one-click updates.
- Add Slack commands to declare incidents, start war rooms, assign roles, and post status updates.
- Integrate runbook and service catalog access; create postmortem templates in Jira or Notion with action item tracking.
- Ensure audit trails and role assignments are logged for SOC 2.
11. Pilot with 2-3 volunteer teams and iterate (after 10)
Run a controlled pilot before full rollout to validate and refine the process.
- Select 2-3 volunteer teams with representative services and on-call patterns.
- Run the new severity, roles, on-call, alerting, and communication process for 2 weeks.
- Track metrics and gather feedback from on-call engineers, incident commanders, and communications leads.
- Iterate severity thresholds, alert rules, templates, and runbooks based on findings.
- Exit criteria: no SEV1 without a declared incident commander, alert noise below target, and positive on-call survey results.
12. Train and certify all 28 teams (after 11)
Prepare all 28 teams so the process is not a paper exercise.
- Create role-specific training: all engineers, incident commanders, communications leads, scribes, subject-matter responders, and managers.
- Deliver interactive workshops and online modules.
- Run at least two game-day drills per quarter simulating SEV1 and SEV2 incidents, testing escalation, communication, and postmortem.
- Certify incident commanders and communications leads; only certified staff can hold those roles in rotations.
- Include scenario-specific training for account managers and legal or regulatory handling.
13. Full rollout and enforce process (after 12)
Move every team onto the new process and retire old habits.
- Migrate all 28 teams into consolidated on-call schedules and alert routing.
- Retire old alerting tools or methods that bypass the process.
- Activate status page templates and postmortem action tracking globally.
- Monitor the first 30 days with a daily operational review and quick fixes.
- Communicate rollout status to leadership and all engineers; provide office hours and support.
14. SOC 2 readiness and continuous improvement (after 13)
Map the process to SOC 2 and keep improving after go-live.
- Map each process component, including severity, roles, detection, communications, postmortems, and metrics, to SOC 2 control requirements.
- Produce runbooks, evidence of training, incident logs, postmortems, and action item tracking.
- Conduct an internal readiness assessment at month 6 and remediate gaps before the audit at month 8.
- Maintain weekly and monthly metrics reviews to drive continuous improvement.
- Refresh on-call compensation and alert quality rules based on operational data.
- Customer-detected incidents decrease from 40% to less than 15% within six months.
- Median time to detect (MTTD) is under 5 minutes for SEV1 and SEV2 incidents.
- Median time to mitigate (MTTM) is under 60 minutes for SEV1 and under 2 hours for SEV2.
- Alert noise decreases from 85% to below 10% within three months.
- 100% of SEV1 and SEV2 incidents have a completed blameless postmortem within 5 business days.
- 100% of postmortem action items are tracked with an owner and due date; 90% are completed on time.
- 24x7 on-call coverage achieved across all 28 teams with no unpaid on-call.
- 99% of pages are acknowledged within 5 minutes.
- SLA credits paid reduce by at least 50% over the next 12 months.
- SOC 2 readiness: all incident response controls are documented, tested, and evidence is produced by month 7.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
Generate an initial, detailed, and well-structured plan proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
Answer with a single JSON object, and nothing else, that conforms to this JSON schema:
{"description": "The structure every proposing agent must answer with, in every round.", "properties": {"steps": {"items": {"description": "One step of a proposed plan. Dependencies are step_ids of the same proposal.", "properties": {"step_id": {"title": "Step Id", "type": "string"}, "title": {"title": "Title", "type": "string"}, "description": {"title": "Description", "type": "string"}, "dependencies": {"items": {"type": "string"}, "title": "Dependencies", "type": "array"}}, "required": ["step_id", "title", "description", "dependencies"], "title": "Step", "type": "object", "additionalProperties": false}, "title": "Steps", "type": "array"}, "estimated_complexity": {"title": "Estimated Complexity", "type": "string"}, "success_metrics": {"title": "Success Metrics", "type": "string"}}, "required": ["steps", "estimated_complexity", "success_metrics"], "title": "ProposalContent", "type": "object", "additionalProperties": false}
{"steps": [{"step_id": "S1", "title": "Baseline current incident response and align stakeholders", "description": "Collect data from the last 12 months of incidents, all six alert tools, on-call practices, and team interviews. Identify gaps against the target incident management process and secure executive sponsorship.\n\n- Gather incident timeline, detection source, mitigation time, customer impact, SLA credit, and postmortem status for all 31 incidents.\n- Survey 28 teams on on-call burden, alert quality, and operational pain.\n- Map current tools, escalation paths, and communication workflows.\n- Create baseline metrics and a stakeholder map with executive sponsor and audit owner.", "dependencies": []}, {"step_id": "S2", "title": "Define severity levels and response triggers", "description": "Define a four-level severity scale with objective business-impact criteria so any engineer can classify an incident consistently.\n\n- SEV1: widespread transaction processing outage, data breach, security incident, or severe SLA breach; triggers full incident command, executive notification, and a 5-minute status page update.\n- SEV2: major feature outage, significant degradation without workaround, or customer financial risk; triggers incident commander, full communications role, and status page updates.\n- SEV3: partial impairment with workaround or limited customer impact; triggers on-call response, internal communication, and optional status page update.\n- SEV4: minor or internal issue, no customer impact; handled during business hours through ticketing.\n- Include an escalation matrix showing who can declare, downgrade, and invoke regulatory or legal involvement.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Define incident roles and decision authority", "description": "Define incident roles, responsibilities, and decision authority using RACI to remove ambiguity about who is in charge.\n\n- Incident Commander: owns the incident, declares severity, and coordinates resolution.\n- Communications Lead: owns internal and external messaging, status page updates, and account manager notifications.\n- Scribe: maintains timeline, incident log, and postmortem notes.\n- Subject-matter responders: diagnose and fix the incident; may come from multiple teams.\n- Executive sponsor: optional for SEV1; customer liaison: handles account managers.\n- Define decision rights for severity declaration, escalation, rollback, customer communications, and incident closure.", "dependencies": ["S1", "S2"]}, {"step_id": "S4", "title": "Design 24x7 staffing model across 28 teams", "description": "Design 24x7 coverage across 28 teams without overloading engineers. Use service-based on-call plus a central incident command pool.\n\n- Each service or domain team assigns primary and secondary on-call for its own services.\n- Create central incident commander, communications, and scribe rotations staffed from a trained incident response guild across all teams; use follow-the-sun between the two AWS regions and time zones.\n- Define escalation layers: service on-call to team lead or manager to service owner to executive.\n- Define handoff times, shadow shifts, and load balancing; target at most one week of on-call per engineer per month.\n- Bridge the current 12-team paid on-call to 28-team paid coverage; no team remains uncovered.", "dependencies": ["S3"]}, {"step_id": "S5", "title": "Define on-call rotations, compensation, and alert quality rules", "description": "Define sustainable rotations, pay, and rules that eliminate noisy pages.\n\n- Rotations: weekly or biweekly, at least one primary and one secondary, with 12-hour shifts where possible or 24-hour for low-volume services.\n- Compensation: monthly on-call stipend for all on-call engineers, additional incident response bonus for after-hours work, and time off in lieu; align with market rates.\n- Alert quality rules: every page must be actionable, have a runbook link, specify a service owner, include severity, and be based on SLO burn or known failure signals; no dashboard-only alerts.\n- Noise budget: reject or downgrade non-actionable alerts; all pages must go to on-call only after suppression and deduplication.\n- Weekly alert review removes the top noisy alerts.", "dependencies": ["S4"]}, {"step_id": "S6", "title": "Design detection, escalation, and alert routing", "description": "Define how incidents are detected, routed, and escalated so nothing waits on a human to notice.\n\n- Consolidate the six alert tools into one alerting and paging platform with routing by service, severity, and tags.\n- Detection sources: infrastructure metrics, application synthetic transactions, log-based anomalies, business transaction SLI monitoring, and customer-reported issues through support or account managers.\n- Routing: alert is paged to service on-call within 30 seconds; primary must acknowledge within 5 minutes; if no ack, page secondary then on-call manager.\n- Escalation timeouts: unresolved SEV1 escalates to service owner at 15 minutes and to leadership at 30 minutes; any engineer can escalate to the incident commander.\n- Define customer-reported incident intake and classification in the same tool.", "dependencies": ["S2", "S4", "S5"]}, {"step_id": "S7", "title": "Define internal and external communication protocols", "description": "Define communication channels, templates, and timing for internal, customer, and regulator audiences.\n\n- Internal: dedicated incident Slack channel, internal status page mirror, and war room bridge for SEV1; incident commander and communications lead own these channels.\n- Status page: SEV1 post within 5 minutes, updates every 30 minutes or on material change, resolution within 60 minutes of mitigation; SEV2 post within 15 minutes, updates hourly; SEV3 optional.\n- Account managers: SEV1 and SEV2 notify account managers within 15 minutes with an approved customer-facing description and expected impact.\n- Regulators: legal or compliance determines notification for data breaches, security incidents, funds availability issues, or regulatory reportable events; criteria and timing follow legal and regulatory requirements; communications lead coordinates.\n- Use pre-approved message templates and an approval chain; no ad-hoc wording.", "dependencies": ["S2", "S3"]}, {"step_id": "S8", "title": "Define postmortem policy and action tracking", "description": "Define mandatory blameless postmortems and action tracking.\n\n- Mandatory for all SEV1 and SEV2 incidents, and any SEV3 that breaches SLA or is customer-detected.\n- Format: impact, timeline, root causes, contributing factors, detection and response gaps, what worked well, and action items.\n- Blameless: focus on system and process causes, not individual blame; use trained facilitators.\n- Ownership: each action has an owner, due date, and tracking ID in a single backlog.\n- Review postmortems at the weekly incident review; track action closure; expect 100% completion.\n- Complete postmortems within 5 business days for SEV1 and SEV2 incidents.", "dependencies": ["S3"]}, {"step_id": "S9", "title": "Define metrics, dashboards, and review cadence", "description": "Define metrics and review cadence to measure process health.\n\n- Metrics: MTTD, MTTM, customer detected percentage, alert noise percentage, on-call response time, on-call load, SLA credits paid, and postmortem action completion.\n- Dashboards: real-time operational dashboard for on-call engineers and management.\n- Weekly incident review: review all SEV1 and SEV2 incidents, action items, and noisy alerts.\n- Monthly trends with leadership; quarterly review against SLOs and audit controls.\n- Success thresholds: MTTD under 5 minutes, MTTM under 60 minutes for SEV1, customer detected under 15%, and alert noise under 10%.", "dependencies": ["S2", "S3", "S8"]}, {"step_id": "S10", "title": "Configure incident tooling and integrations", "description": "Implement and integrate the tools that automate the defined process.\n\n- Aggregate alerts from the existing six tools into PagerDuty, Opsgenie, or a similar platform.\n- Configure on-call schedules, escalation policies, and paging targeted at service owners.\n- Integrate status page API for automated or one-click updates.\n- Add Slack commands to declare incidents, start war rooms, assign roles, and post status updates.\n- Integrate runbook and service catalog access; create postmortem templates in Jira or Notion with action item tracking.\n- Ensure audit trails and role assignments are logged for SOC 2.", "dependencies": ["S5", "S6", "S7", "S8", "S9"]}, {"step_id": "S11", "title": "Pilot with 2-3 volunteer teams and iterate", "description": "Run a controlled pilot before full rollout to validate and refine the process.\n\n- Select 2-3 volunteer teams with representative services and on-call patterns.\n- Run the new severity, roles, on-call, alerting, and communication process for 2 weeks.\n- Track metrics and gather feedback from on-call engineers, incident commanders, and communications leads.\n- Iterate severity thresholds, alert rules, templates, and runbooks based on findings.\n- Exit criteria: no SEV1 without a declared incident commander, alert noise below target, and positive on-call survey results.", "dependencies": ["S10"]}, {"step_id": "S12", "title": "Train and certify all 28 teams", "description": "Prepare all 28 teams so the process is not a paper exercise.\n\n- Create role-specific training: all engineers, incident commanders, communications leads, scribes, subject-matter responders, and managers.\n- Deliver interactive workshops and online modules.\n- Run at least two game-day drills per quarter simulating SEV1 and SEV2 incidents, testing escalation, communication, and postmortem.\n- Certify incident commanders and communications leads; only certified staff can hold those roles in rotations.\n- Include scenario-specific training for account managers and legal or regulatory handling.", "dependencies": ["S11"]}, {"step_id": "S13", "title": "Full rollout and enforce process", "description": "Move every team onto the new process and retire old habits.\n\n- Migrate all 28 teams into consolidated on-call schedules and alert routing.\n- Retire old alerting tools or methods that bypass the process.\n- Activate status page templates and postmortem action tracking globally.\n- Monitor the first 30 days with a daily operational review and quick fixes.\n- Communicate rollout status to leadership and all engineers; provide office hours and support.", "dependencies": ["S12"]}, {"step_id": "S14", "title": "SOC 2 readiness and continuous improvement", "description": "Map the process to SOC 2 and keep improving after go-live.\n\n- Map each process component, including severity, roles, detection, communications, postmortems, and metrics, to SOC 2 control requirements.\n- Produce runbooks, evidence of training, incident logs, postmortems, and action item tracking.\n- Conduct an internal readiness assessment at month 6 and remediate gaps before the audit at month 8.\n- Maintain weekly and monthly metrics reviews to drive continuous improvement.\n- Refresh on-call compensation and alert quality rules based on operational data.", "dependencies": ["S13"]}], "estimated_complexity": "high", "success_metrics": "- Customer-detected incidents decrease from 40% to less than 15% within six months.\n- Median time to detect (MTTD) is under 5 minutes for SEV1 and SEV2 incidents.\n- Median time to mitigate (MTTM) is under 60 minutes for SEV1 and under 2 hours for SEV2.\n- Alert noise decreases from 85% to below 10% within three months.\n- 100% of SEV1 and SEV2 incidents have a completed blameless postmortem within 5 business days.\n- 100% of postmortem action items are tracked with an owner and due date; 90% are completed on time.\n- 24x7 on-call coverage achieved across all 28 teams with no unpaid on-call.\n- 99% of pages are acknowledged within 5 minutes.\n- SLA credits paid reduce by at least 50% over the next 12 months.\n- SOC 2 readiness: all incident response controls are documented, tested, and evidence is produced by month 7."}Round 1 — refinement 1 of 3
All five plans converged on a common architecture — 7-day interim command, forensic baseline, service catalog with tiers, 4–5 level severity, IC/Comms/Scribe/SME roles, a central command pool plus 8–12 domain rotations, paid on-call before mandatory nights, page budgets, money-path detection, one pager, 3/5/10-day postmortems with tracked actions, critical-path pilot, gated waves, month-6/7 audit rehearsal. The real remaining spread is in severity thresholds, compensation concreteness, and how much program scaffolding (policy v1, risk register, 90-day revision) each keeps.
What still separates them
- Severity design. P2 (S5) alone gives quantitative declaration guardrails (>10% payment failures for 5 min = SEV1; 1–10% = SEV2) plus a SEV0 crisis tier and no time-based auto-escalation; P5 (S5) uses five levels with aggressive clocks (SEV2 unresolved >1h becomes SEV1); P1/P3/P4 (S6) use SEV1–SEV4 with ledger-touch and 2-hour auto-escalation.
- Compensation and load caps. P1 (S9) and P5 (S9) publish dollar figures (~$800–1,200/primary week, $150/night page); P4 (S9) budgets $600k–1M with structure only; P2 (S8) and P3 (S7) defer amounts to HR in 14 days. Caps diverge: one primary week in six (P1, P2, P4) versus one in four (P5 S8; P3 S7 says "four or six depending on staffing").
- Program scaffolding. P1 keeps a standalone Policy v1 with exception register (S24), risk register (S31) and a 90-day inspect-and-adapt (S32); P3 keeps a risk register (S26). P4 dropped its own Policy v1 and 90-day revision; P2 and P5 have no contingency/risk step at all.
- Interim protection and density. P1 (S2), P2 (S2), P4 (S2) and P5 (S1) stand up interim command in 48h–7 days; P3 only lists "interim controls week 1" in S1 with no step behind it. On density, P4 is tightest at 24 steps, P1 heaviest at 32 with overlapping alert work (S10 standard vs S28 burn-down) and an 11-dependency join at S24.
Who took what from whom
- P2's 7-day interim minimum controls, retroactive pay for interim duty, triage of the 53 open actions and the daily 15-minute ops review (R0 S2) were taken by P1 (S2), P4 (S2) and P5 (S1); P3 left it as a single timeline bullet.
- P4's three-layer coverage and the "deal" fairness contract (R0 S8, S26) were taken by P1 (S8, S29), P3 (S8, S24) and P5 (S7–S8); P2 reached the same result via 8–12 responder domains (S7) without the layer naming.
- P1's forensic baseline, page budget, SLA-credit workflow, Incident Review Board and gated waves (R0 S2, S10, S17, S18, S25) were adopted almost wholesale by P3 (S2, S9, S14, S16, S22) and selectively by P4 (S3, S10, S16, S17, S22) and P2 (S3, S11, S15, S18, S21).
- P2's safety rails travelled widely: shadow-mode alerts and "never silently disable", two weeks of verified operation before retiring a legacy path, human acknowledgement required, mitigation-vs-resolution semantics and ledger dual-control (R0 S1, S4, S7, S8, S9) now appear in P1 (S6, S7, S10, S12), P4 (S7, S10, S12, S14), P3 (S9, S11) and P5 (S13).
- Nobody took P3's outsourced overnight triage vendor or its per-page pay table (R0 S4, S5) — P3 dropped both itself — and nobody kept P5's 5-minute SEV1 status-page deadline (R0 S7); all five settled on 15 minutes for the top availability tier.
The calls of this round
Influences: who took what from whom
| Round 1 ↓ · round 0 → | Proposal 1 | Proposal 2 | Proposal 3 | Proposal 4 | Proposal 5 | New steps |
|---|---|---|---|---|---|---|
| Proposal 1 |
kept18 | same titles1 analyst sees+4 / −1 | same titles0 analyst sees+0 / −1 | same titles3 analyst sees+2 / −0 | same titles0 analyst sees+0 / −1 | new10 |
| Proposal 2 |
same titles3 analyst sees+4 / −2 | kept7 | same titles0 analyst sees+0 / −1 | same titles3 analyst sees+1 / −0 | same titles0 analyst sees+0 / −0 | new11 |
| Proposal 3 |
same titles17 analyst sees+3 / −0 | same titles0 analyst sees+1 / −2 | kept1 | same titles0 analyst sees+1 / −0 | same titles0 analyst sees+0 / −1 | new9 |
| Proposal 4 |
same titles11 analyst sees+2 / −1 | same titles1 analyst sees+3 / −1 | same titles0 analyst sees+0 / −1 | kept4 | same titles0 analyst sees+0 / −0 | new8 |
| Proposal 5 |
same titles9 analyst sees+2 / −0 | same titles5 analyst sees+3 / −2 | same titles1 analyst sees+0 / −0 | same titles1 analyst sees+0 / −1 | kept4 | new6 |
Closed its biggest R0 gap — six weeks of design with no protection — by adding a 7-day interim command bridge (S2). Replaced "14–18 teams at 24x7" with a Layer A/B/C model over 10–12 domains (S8), added live-execution doctrine (S14), Policy v1 (S24), a noise burn-down campaign (S28) and staged day-90/month-6 metric targets. Now the most complete plan, but at 32 steps it is also the heaviest.
- S2 interim bridge: named commander within 10 min, Support escalates without engineering sign-off, 53 legacy actions triaged, daily 15-min review.
- S6 adds a lifecycle (Detected→Reviewed) with mitigation = customer impact ends, resolution = backlog processed and ledger reconciled, plus detection clock starting at first reliable impact signal.
- S8 consolidates 28 teams into 10–12 domain rotations with a ~30-person command corps, directly answering the pager objection instead of asserting it away.
- S14 payments-specific execution guards (split-brain, replay, duplication, controlled backlog drain) that R0 lacked.
- Metrics now staged (MTTD <10 min day 90, <5 min month 6; pages 3,400→1,500 by day 90→<500 by month 6) rather than single endpoints.
- S26 anti-gaming: monthly reconciliation of incident counts against support tickets, credits and status-page history.
- S30 adds independent internal control testing in months 4 and 6 before the month-7 mock audit.
- 32 steps with overlap: S10 (alert standard) and S28 (burn-down campaign) cover the same ground, as do S14 playbooks and S21 runbooks.
- S24 Policy v1 depends on 11 steps and gates the pilot (S25) — a fragile critical path late in the schedule.
- Internal inconsistency: S27 claims teams finish "at least five months before the audit report date" while the S1 timeline ends rollout at week 24 against a month-8 audit (~2.5 months).
- Metric list has grown to 21 items; no prioritisation of which three the CEO sees.
- Proposal 2 : Interim minimum controls within seven days, compensated retroactively, with Support empowered to escalate.
- Proposal 2 : Impact-based lifecycle with mitigation and resolution defined separately, resolution requiring reconciliation.
- Proposal 2 : Preserve dual control and privileged-access rules during incidents; group teams into 8–12 coherent domains.
- Proposal 2 : Shadow mode for new alerts, never silently disable, and reconcile dashboards against customer cases to catch gaming.
- Proposal 4 : Publish a short Policy v1 with laminated cards, run noise reduction as a quota campaign with a leaderboard, and revise the process at 90 days on data.
- Proposal 4 : Hard date after which pages outside the chosen tool create no on-call obligation.
- Proposal 2 : A distinct SEV0 crisis tier above SEV1.
- Proposal 3 : All 28 teams staffed with 24x7 primary and secondary on-call.
- Proposal 5 : Status page within 5 minutes of SEV1 declaration.
+ Executive mandate, single owner, funding and non-negotiables+ Seven-day interim command bridge+ 24x7 coverage model: central command corps, local expertise+ Detection uplift on the money path+ Consolidate to one pager, one incident record, one status page+ Live incident execution doctrine+ Action ownership, capacity reservation and enforcement+ Runbooks, readiness bar and ledger blast-radius reduction+ Exercise programme: tabletops, game days and unannounced drills+ Publish Incident Management Policy v1+ Pilot on the payments critical path+ Alert noise burn-down campaign+ SOC 2 evidence by design, internal testing and mock audit+ Ninety-day inspect-and-adapt, then year-two sustainabilityProgram charter, executive mandate and funding24x7 incident command coverage modelTeam on-call structure, rotations and routing rulesDetection uplift: SLOs, synthetic journeys and ledger assuranceTool consolidation and incident platform implementationAction item ownership and tracking systemRunbooks, major-incident playbooks and the on-call readiness barSimulation programme: game days, drills and wheel of misfortunePilot with wave 0 teamsSOC 2 control mapping, evidence automation and internal dry-run auditContinuous improvement, maturity roadmap and post-audit sustainability
The plan produced
1. Executive mandate, single owner, funding and non-negotiables new
Convert the CEO email into a chartered program with one accountable owner and authority over all 28 teams. Incident response becomes a company operating process, not a per-team choice.
- Name an executive sponsor (CTO) and one accountable owner (Head of Reliability / Incident Management) with a small permanent office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group with Engineering, SRE/Platform, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours. It is not a 28-person committee.
- Fix the non-negotiables now: one severity scale, one paging tool, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock deliberately earlier than the audit: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready week 28.
- Approve budget against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises.
2. Seven-day interim command bridge (after 1) new
Do not let design work leave the company unprotected for six weeks. Put a crude but real process in place within seven days and improve it later.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster from engineering managers and senior SREs, primary plus backup, 24x7. Pay it retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Tell Support to escalate credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items: complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (after 1)
Rebuild the facts before designing anything. This becomes both the design input and the "before" picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start / detect / declare / commander-assigned / mitigate / resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. This list drives the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome. Identify the top 50 rules producing most of the noise and every rule with no owner or runbook.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the on-call deal (after 1)
Engineer pushback is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS and Sales in two weeks. Test what the objection actually is: unpaid work, lost sleep, unfamiliar code, missing runbooks, or fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey (fairness, trust in alerts, willingness, burnout) to re-measure at 60, 120 and 365 days.
5. Service catalog, ownership and money-path tiering (after 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that is the single source of truth for paging, impact and audit.
- One accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not by technology: Tier 0 (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions and third parties, including sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Every orphan service gets an owner in 30 days or a decommission date approved by the sponsor. Tier 0 without an owner is an executive escalation.
6. Severity scale, declaration rules and incident lifecycle (after 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- SEV1 (crisis): money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- SEV2 (critical): material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- SEV3 (major, contained): narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring; only the commander may downgrade, with the evidence recorded.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer impact ends) → Monitoring → Resolved (backlog processed and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Incident roles, authority and handover discipline (after 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- Incident Commander: owns severity, priorities, roles, cadence and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates for SEV1/SEV2.
- Subject-Matter Responders: engineers of the owning team; they mitigate, they do not run the room.
- Executive Duty Officer (SEV1): removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance and Vendor Management join on defined triggers.
- Rules: command claimed within 5 minutes, stated in channel ("I am IC"), distinct people for command, comms and technical lead at SEV1/SEV2, and every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation or privileged-access rules.
8. 24x7 coverage model: central command corps, local expertise (after 5, 7) new
Do not create 28 night rotations. Centralise coordination in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- Layer A — Incident Command corps: ~30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of ~18 from Support, CS and engineering management; a scribe pool used as the training entry point.
- Layer B — critical-path domain rotations: consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- Layer C — everyone else: business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing, with cost and timeline: a US-only paid night rotation now, a Lisbon or APAC follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
9. On-call compensation, labour compliance and fatigue safeguards (after 4, 8)
Unpaid on-call in New York is both a retention problem and a legal exposure. Pay for it before asking anyone to sign up, and publish the numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 for business-hours rotations, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover rather than expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, tax treatment and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a ~15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
10. Alert quality standard, page budget and noise burn-down (after 3, 5)
3,400 alerts at 85% noise is the reason detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human, and burn the backlog down deliberately rather than by mass silencing.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements; return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
- Target 3,400 → under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
11. Detection uplift on the money path (after 5, 10) new
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls; five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. Consolidate to one pager, one incident record, one status page (after 6, 8, 10) new
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrating without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack or the identity provider is unavailable. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Escalation ladder and the five-minute command rule (after 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live incident execution doctrine (after 7, 12, 13) from P2 step 10
Give responders one short operating procedure for the first minutes through closure. Priority is limiting customer and financial harm, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; controlled backlog drain; reconciliation completed before any payment or ledger incident is declared resolved.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and Customer Success, and reopening if impact recurs.
15. Internal communications protocol (after 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (after 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook and RSS; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack; the long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity or blame; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (after 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded even when the answer is "not reportable", with evidence, decision-maker and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review; Legal owns outbound regulatory text, the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
18. SLA credit and financial impact workflow (after 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one, and use the delta as the standing business case for on-call pay and reliability capacity.
19. Blameless postmortem standard and Incident Review Board (after 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first? and why did mitigation take as long as it did?
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, capacity reservation and enforcement (after 12, 19) new
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
21. Runbooks, readiness bar and ledger blast-radius reduction (after 5, 8) new
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the single largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel workstream on blast-radius reduction — tenant or function partitioning, read replicas, and isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (after 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days + shadowing): command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Comms Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence for their own domain before independent primary duty; two shadow shifts minimum, never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (after 12, 21, 22) new
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (after 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20) from P4 step 14
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (after 9, 11, 12, 21, 22, 24) from P4 step 22
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (after 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- Response: time to detect (from first impact), time to declare, time to commander, acknowledgement, mitigate, resolve — split by severity, tier, journey, region and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; team scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Wave rollout to all 28 teams with readiness gates (after 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director; failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
28. Alert noise burn-down campaign (after 10, 12, 25) from P4 step 23
Run the noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
29. Change management, fairness and pager culture (after 4, 9, 25)
Run this from day one in parallel. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (after 24, 26, 27) new
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and contingencies (after 1)
Name the ways this programme fails and pre-commit the response. Review it monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (after 26, 27, 30) new
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
- A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 223f0a8a-068a-4034-ad20-8e2ff8553040, Agent: opus5_initial_1, LLM: anthropic/claude-opus-5):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 9 months.
- Customer-first detection falls from 40% of incidents to under 10% within 6 months and under 5% within 12.
- Median time to mitigate falls from 3 h 10 min to under 60 minutes within 12 months.
- An Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes.
- Status page updated within 15 minutes of SEV1 declaration and 30 minutes of SEV2 in 95% of cases.
- Monthly alert volume falls from 3,400 to under 500 pages, with actionability above 75%; out-of-hours pages under 2 per person per week.
- All six legacy alerting tools consolidated into one paging platform, legacy paging paths disabled, by week 16.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by their due date.
- SLA credits fall from $1.3M to under $400k in the first 12 months.
- Customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar.
- All 28 teams onboarded by week 24, with 24x7 rotations of 6+ certified responders for every Tier 0/1 team.
- 30+ certified Incident Commanders and 20+ certified Communications Leads, giving 24x7 primary and secondary command cover.
- Paid on-call policy approved by HR, Legal and Finance and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among engineers on rotation.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked action items.
Steps (29):
1. Program charter, executive mandate and funding
Convert the CEO's frustration into a named program with one accountable owner, a budget and a deadline that is earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, the 28 teams will opt out.
- Appoint a single Incident Management Program Lead (Head of Reliability/SRE) with direct exec sponsorship from CTO and CEO.
- Form a steering group: CTO, VP Eng, Head of Support/CS, CISO/Compliance, Legal, Finance (SLA credits), HR (on-call pay).
- Set the non-negotiables: one severity scale, one paging tool, one postmortem format, mandatory action tracking, paid on-call.
- Fix the timeline: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready by week 28 (four weeks of buffer before the audit).
- Approve budget lines: tooling (~$150–250k/yr), on-call compensation (~$600k–1M/yr), 2–3 dedicated program FTEs. Anchor it against $1.3M of credits plus incident cost.
2. Forensic baseline of the 31 incidents and the alert estate (depends on: 1)
Before designing anything, rebuild the facts. Re-open all 31 incidents and profile the 3,400 monthly alerts so every later design decision is evidence-based.
This also creates the "before" picture the exec and the auditor will compare against.
- Re-code each incident: trigger, service, detection source (customer vs monitor), timestamps for detect/acknowledge/declare/mitigate/resolve, who led, credits paid, root cause family.
- Quantify the 40% customer-first detections: which signal was missing in each case.
- Classify the two "nobody in charge" incidents minute by minute; use them as the burning-platform story.
- Audit the six alerting tools: volume per tool, per team, per alert rule; identify the top 50 rules that produce most of the 85% noise; find rules with no owner and no runbook.
- Baseline the numbers formally: MTTD 22 min, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed. Freeze them as the reference line.
3. Stakeholder listening tour and resistance map (depends on: 1)
Engineer pushback against "carrying a pager for other teams' code" is the main delivery risk. Treat it as a design input, not an attitude problem.
Run structured interviews across all 28 teams, plus Support, CS and Sales, in two weeks.
- Test the real objection: is it unpaid work, night sleep, unfamiliar code, poor runbooks, or fear of blame? Each has a different fix.
- Collect current informal practices — the 12 teams already on-call are the pilot candidates and the source of veterans.
- Document the promise that answers the objection: **you are only paged for services your team owns**, plus a trained commander who runs the incident and pulls in others.
- Map influencers and blockers by team; recruit 10–15 credible engineers as a design working group so the process is co-authored, not imposed.
- Survey baseline sentiment (trust in alerts, willingness to be on-call, burnout) to re-measure at 6 and 12 months.
4. Service ownership catalog and criticality tiering (depends on: 2)
You cannot page the right person across 180 services until each service has a named owning team. This is the foundation of both on-call fairness and severity mapping.
Build a machine-readable catalog (Backstage or equivalent) that is the single source of truth for routing.
- One owning team per service, a named engineering manager, a Slack channel, a paging escalation policy, a dependency list.
- Tier services by business impact: Tier 0 (money movement, ledger, auth, shared PostgreSQL cluster), Tier 1 (customer-facing but degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map each Tier 0/1 service to the customer-visible capability it supports (payment initiation, settlement, reporting, onboarding).
- Flag orphan services and cross-team shared components; force an ownership decision for each within 30 days, or schedule decommissioning.
- Publish coverage gaps to the steering group: any Tier 0 service without an owner is an executive escalation.
5. Severity scale and declaration criteria (depends on: 2, 4)
Define a five-level scale with objective, payments-specific triggers so declaration is a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade.
Each level triggers a fixed bundle of response, comms and postmortem obligations.
- **SEV1**: money movement stopped or incorrect, ledger integrity in doubt, data breach, full region loss, >10% of customers impacted. Triggers: immediate 24x7 page of IC + comms + exec, bridge within 5 min, status page within 15 min, mandatory postmortem, regulator assessment.
- **SEV2**: severe degradation, settlement at risk of missing a window, single large/strategic customer fully down, SLA breach likely. Triggers: IC paged, status page within 30 min, mandatory postmortem.
- **SEV3**: partial or workaround-available degradation, no credit exposure. Team-led, business-hours comms, postmortem optional but encouraged.
- **SEV4/5**: minor or internal-only; ticket-tracked, no paging.
- Add auto-escalation rules: any SEV3 open >2 h, or any incident touching the shared ledger cluster, becomes SEV2 automatically. Include a severity decision tree and 12 worked examples drawn from the 31 real incidents.
6. Incident roles, decision authority and handover rules (depends on: 5)
Solve the "nobody in charge for an hour" failure by making command explicit, transferable and logged.
Define five roles with written responsibilities, entry criteria and explicit authority.
- **Incident Commander**: owns the incident, not the fix. Authority to declare severity, pull any engineer, approve customer-impacting mitigations, invoke failover and authorise spend. The IC never types in the terminal.
- **Communications Lead**: owns status page, internal updates, account-manager briefings and the exec summary. Single voice to customers.
- **Scribe**: maintains the timeline, decisions and open questions; feeds the postmortem and the audit evidence trail.
- **Subject-Matter Responders**: engineers from owning teams; they investigate and remediate, and report to the IC.
- **Executive Liaison** (SEV1 only): shields the IC from exec questions and owns regulator/board escalation.
- Rules: the IC role is assumed within 5 minutes of declaration, stated explicitly in the channel ("I am IC"), and any handover is announced and logged. Roles may be combined below SEV2; never at SEV1.
7. 24x7 incident command coverage model (depends on: 6, 4)
Command is staffed by a small, trained, cross-team pool — not by 28 teams individually. This is what makes 24x7 realistic in one New York time zone.
Design a central rotation that scales with a growing certified pool.
- Create a **Duty Incident Commander** rotation of 25–35 certified volunteers (target ~1 per team, plus managers and senior engineers), giving each person roughly one week per 6–8 months.
- Pair with a Duty Comms Lead rotation (Support/CS leads plus engineering managers, ~15–20 people) and a Scribe pool (rotating, lowest barrier, used as the training entry point).
- Coverage: primary + secondary IC at all times; hard 5-minute acknowledgement SLA with automatic failover to secondary, then to the on-call engineering director.
- Night coverage options to evaluate in writing: US-only rotation with paid night stipend now; a Lisbon/Dublin or APAC follow-the-sun cell as a 12-month option; a 24x7 NOC-style triage desk for first-line detection.
- Eligibility: certification required (S21); commanders are volunteers with manager approval and can step out with 30 days' notice.
8. Team on-call structure, rotations and routing rules (depends on: 4, 6)
Rebuild team on-call around the principle that answers the pushback: **you are only paged for code your team owns**.
Apply a tiered obligation so 28 teams are not treated identically.
- Tier 0/1 owning teams (expected 14–18 teams): 24x7 primary + secondary, minimum 6 people per rotation, one-week shifts, handover Wednesday mornings.
- Tier 2/3 teams: business-hours on-call with a best-effort out-of-hours escalation path, no night paging.
- Platform/Infrastructure and Database teams: 24x7, since they own the shared PostgreSQL ledger cluster and the Kubernetes/regional layer.
- Rotations under 6 people are merged across teams or backfilled by hiring; no rotation of fewer than 4 is approved.
- Routing: every page resolves through the service catalog to the owning team's escalation policy; cross-team pages are made by the IC, never by an alert.
- Guardrails: maximum one week in four, no on-call in the week after a SEV1 you led, protected recovery time after any night page, and a per-person page budget (see S10).
9. On-call compensation, labour compliance and fairness policy (depends on: 8, 3)
Unpaid on-call is both a retention risk and a legal exposure in New York. Paying for it is the fastest way to convert resistance into participation.
Design the scheme with HR, Legal, Finance and Payroll, and publish it before asking anyone to sign up.
- Base stipend per week on rotation, differentiated by tier: e.g. $800–1,200 for 24x7 Tier 0/1, $300–500 for business-hours rotations, with premiums for holidays and weekends.
- Per-incident payment for out-of-hours activation (e.g. $150 per night page plus hourly beyond one hour) and guaranteed time-off-in-lieu after night work.
- Separate Duty IC stipend, since command is a distinct and heavier burden.
- Verify FLSA exempt/non-exempt treatment, NY State wage rules and overtime exposure for non-exempt staff; document the legal review.
- Budget and model the annual cost; get board/CFO approval as a line item, benchmarked against $1.3M of credits.
- Add non-cash elements: on-call time counted as delivery load (teams reduce sprint commitment by ~15%), incident leadership recognised in promotion criteria, and a public quarterly report of on-call load per team.
10. Alert quality standard and page budget (depends on: 2, 4)
3,400 alerts a month at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a contractual condition of paging someone.
Publish a standard, then enforce it mechanically.
- Every paging alert must have: a named owning team, a documented customer impact, a runbook link, a tested threshold, and a severity mapping. Alerts failing this are demoted to ticket or deleted.
- Page only on symptoms that affect customers (SLO burn rate, error budget, queue depth against settlement deadlines); cause-based CPU/memory alerts become dashboards or tickets.
- Set a **page budget**: maximum 2 out-of-hours pages per person per week. Breach triggers a mandatory alert-tuning sprint for the owning team and blocks new alert creation.
- Auto-quarantine: any alert that fires more than 5 times a month without action, or has >70% no-action acknowledgements, is silenced automatically and returned to its owner.
- Monthly alert review per team: kill, tune, or keep, with the numbers on screen.
- Target: 3,400 → under 500 pages a month, with actionability above 75% within six months.
11. Detection uplift: SLOs, synthetic journeys and ledger assurance (depends on: 10, 4)
The goal is to stop customers telling you first. Detection must be driven by customer-visible outcomes, not host metrics.
Instrument the money path end to end and alert on it.
- Define SLOs for each Tier 0/1 customer capability: payment initiation success rate, authorisation latency, settlement file timeliness, API availability and reporting freshness. Tie them to the 99.95% contractual SLA with a stricter internal target.
- Deploy synthetic transactions from outside the platform, in both regions, every 60 seconds, covering the full payment lifecycle including a small real-value canary flow where feasible.
- Add ledger assurance checks: continuous double-entry balance reconciliation, replication lag and failover-readiness alarms on the shared PostgreSQL cluster, and settlement-window countdown alerts.
- Build per-customer anomaly detection for the top 100 accounts (volume drop, error spike) so a single-tenant outage is detected before the account manager calls.
- Create an "inbound signal" bridge: any support ticket or account-manager report matching impact keywords auto-creates a triage incident within 5 minutes.
- Track every incident's detection source; make "customer detected first" a reviewed defect with its own follow-up action.
12. Tool consolidation and incident platform implementation (depends on: 5, 7, 8, 10)
Collapse six alerting tools into one paging and incident platform so there is a single queue, a single timeline and a single audit record.
Run a short, time-boxed selection and migrate within the pilot window.
- Select an integrated stack: paging/on-call scheduling plus an incident management layer (e.g. PagerDuty + incident.io/FireHydrant, or a single vendor) and a hosted status page.
- Implement one-command declaration in Slack (`/incident declare`) that creates the channel and bridge, pages the Duty IC, sets severity, opens the timeline and starts the clock.
- Migrate all monitoring sources to route into the one platform; decommission direct paging from the legacy six and block new integrations that bypass it.
- Automate the evidence trail: timestamps, role assignments, severity changes, comms sent, and postmortem linkage exported for SOC 2.
- Integrate with the service catalog for routing, Jira for actions, Salesforce/CS tooling for affected-customer lists, and Zoom/Slack Huddle for the bridge.
- Hard requirement: the platform must work when AWS in one region is down — verify out-of-band paging (SMS/phone) and a printed/offline fallback runbook.
13. Detection-to-escalation path and the five-minute command rule (depends on: 6, 7, 12)
Write the single path from "something looks wrong" to "someone is in charge", and make it impossible to skip.
The design target is detection to commander in under five minutes, any hour.
- Entry points: automated alert, engineer observation, support ticket, account manager, partner bank, customer-facing SEV hotline. All converge on the same declaration command.
- Anyone in the company may declare up to SEV2; nobody is punished for over-declaring. Publish that rule in writing and repeat it.
- Auto-page ladder: Duty IC (5 min) → secondary IC (5 min) → on-call Director (10 min) → CTO. Same ladder for the owning team's responder.
- Cross-team pull: the IC can page any team's on-call directly, with a 10-minute acknowledgement obligation. This is the reciprocal commitment that makes single-team ownership viable.
- Explicit takeover protocol: if no one claims IC within 5 minutes, the platform assigns it and announces it; the assignee cannot decline, only hand over.
- Define standing severity triggers for immediate regional failover, ledger read-only mode and partner-bank notification, with pre-authorised decision rights so the IC does not wait for an executive.
14. Internal communications protocol (depends on: 6, 12)
Standardise the internal channel so responders, executives and support see the same picture without interrupting the IC.
Separate the working channel from the audience channel.
- One incident channel per incident (auto-created), one bridge, and a read-only broadcast channel for executives, Support and Sales.
- Update cadence by severity: SEV1 every 30 minutes even if nothing has changed; SEV2 every 60 minutes; SEV3 at state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, ETA or next update time, current IC and Comms Lead.
- Exec briefing rule: executives ask questions only to the Executive Liaison; the IC is not interrupted. Publish this as a behavioural expectation signed by the exec team.
- Support/CS enablement: a live affected-customer list and a holding statement within 15 minutes of SEV1/SEV2 so the front line is never guessing.
- Handover protocol for incidents beyond 4 hours: formal IC handover checklist, fatigue rule, and staffing of a second shift.
15. Customer communications and status page policy (depends on: 5, 14)
Customers currently learn of outages from their own monitoring and hear from whoever happens to be around. Replace that with a timed, owned, pre-approved process.
The Comms Lead is the single author; templates remove the need to write under pressure.
- Timing commitments: status page posted within 15 minutes of SEV1 declaration and 30 minutes for SEV2; updates every 30/60 minutes; resolution notice within 30 minutes of mitigation; customer-facing summary within 5 business days for SEV1.
- Pre-approve 12–15 templates with Legal and Comms (degradation, delay in settlement, API errors, security event, third-party failure) so nothing needs legal review mid-incident.
- Subscription-based status page with per-component granularity mapped to the customer capabilities from S11, plus an email/webhook/RSS feed.
- Tiered outreach: top 100 accounts get a direct named call or email from their account manager within 30 minutes of SEV1, with a briefing pack from the Comms Lead; long tail gets the status page and a proactive email.
- Rules of language: state impact and next update time, never speculate on cause, never assign blame to a vendor before facts are confirmed.
- Run a quarterly customer-perception check with the top accounts on whether comms were timely and useful.
16. Regulatory, partner and legal notification playbook (depends on: 5, 15)
In payments, some incidents are reportable and the clock starts at detection. Build the assessment into the process so it is never an afterthought.
Work with Legal, Compliance and the CISO to produce a decision tree and contact matrix.
- Map obligations: NYDFS Part 500 (72-hour cybersecurity event notification), state breach laws, GLBA/FTC Safeguards, PCI DSS if card data is in scope, sponsor-bank and card-network contractual notice windows, and any FinCEN/OFAC implications.
- Add a mandatory regulatory-assessment checkpoint to every SEV1 and every security-related SEV2, owned by the Executive Liaison, completed within 2 hours of declaration and recorded even when the answer is "not reportable".
- Build the contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel, with 24x7 numbers and named backups.
- Pre-draft notification letters and hold them under legal privilege review.
- Check customer contracts for bespoke notification SLAs (often 1–4 hours for enterprise accounts) and encode them in the customer tiering.
- Test the playbook once per quarter as part of the simulation programme.
17. SLA credit and financial impact workflow (depends on: 5, 15)
Link incidents to money so severity, credits and prioritisation stay consistent — and so Finance stops being surprised.
Make credit calculation an automated output of the incident record, not a negotiation.
- Define the availability measurement method per contract, per component, and agree it with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and the per-capability telemetry; generate a proposed credit schedule within 5 business days of resolution.
- Decide the posture: proactive credits for the top tier (reputation upside) versus claims-based for the rest; document the approval chain.
- Track credits per incident and per root-cause family; feed a quarterly report showing which reliability investments would have prevented which credits.
- Set a target: reduce credits from $1.3M to under $400k in year one, and use that delta as the ongoing business case.
18. Postmortem standard and blameless review forum (depends on: 5, 6)
Replace "some incidents, various formats" with a mandatory, single-format, blameless process with fixed deadlines.
The discipline is in the deadlines and the forum, not in the template.
- Mandatory for: every SEV1 and SEV2, every incident where a customer detected it first, every incident over 2 hours, every repeat of a known cause, and every near-miss involving the ledger. Optional but templated for SEV3.
- Fixed timeline: draft within 3 business days, peer review within 5, published company-wide within 10. The IC owns delivery; the owning team's manager is accountable.
- One template: timeline, customer and financial impact, detection analysis (why not sooner), response analysis (why mitigation took as long as it did), contributing factors, what went well, action items with owner and due date.
- Blameless rules in writing: describe systems and decisions in the context available at the time; no individual named as a cause; HR and management commit that postmortems are never used in performance reviews.
- Weekly 60-minute Incident Review Board: reviews all postmortems from the prior week, challenges quality, ratifies severity, and approves or rejects action items. Attendance by engineering directors is mandatory.
- Publish a searchable postmortem library and a quarterly "top five recurring causes" analysis.
19. Action item ownership and tracking system (depends on: 18, 12)
11 of 64 actions closed is the single clearest symptom of a process nobody enforces. Give actions the same status as customer commitments.
Track them where engineering work already lives, with visible escalation.
- Every action gets: a named individual owner (not a team), a priority class, a due date and a Jira ticket auto-created from the postmortem.
- Priority classes with hard SLAs: P0 prevents recurrence of a SEV1, due in 30 days; P1 in 60 days; P2 in 90 days. P0s are committed into the next sprint before any roadmap work.
- Capacity rule: teams reserve a standing 15–20% of sprint capacity for reliability and incident actions. Without reserved capacity, the actions will not land.
- Escalation ladder for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Monthly reporting of closure rate by team in the engineering leadership review; include it in manager performance objectives.
- Target: 90% of P0/P1 actions closed on time within two quarters.
20. Runbooks, major-incident playbooks and the on-call readiness bar (depends on: 4, 8)
Nobody can respond well to unfamiliar systems at 3 a.m. without runbooks — and poor runbooks are a real part of the pager resistance.
Define a minimum readiness bar that a service must meet before it is allowed to page anyone.
- Readiness checklist per Tier 0/1 service: current architecture diagram, dependency map, dashboard link, alert-to-runbook mapping, rollback procedure, feature-flag kill switches, escalation contacts, and a data-loss/latency impact statement.
- Write major-incident playbooks for the top failure modes derived from S2: shared PostgreSQL ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank/third-party outage, settlement-window breach, and suspected security compromise.
- Prioritise the shared ledger: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and a clearly stated data-loss tolerance (RPO/RTO) signed off by the exec.
- Runbooks must be tested at least twice a year in a drill; untested runbooks are marked stale in the catalog.
- Enforcement: a service without readiness sign-off cannot create paging alerts, and the gap is reported to its director.
21. Training, certification and the commander academy (depends on: 6, 13, 14, 18)
Command is a skill, not a title. Build a certification path so 24x7 coverage is staffed by people who have practised.
Use a tiered curriculum with real assessment.
- **Scribe** (2 hours): timeline discipline and tooling. The entry point for everyone.
- **Responder** (half a day): severity scale, declaration, escalation, runbook use, comms hygiene. Mandatory for every engineer joining an on-call rotation.
- **Incident Commander** (two days plus shadowing): command presence, delegation, decision-making under uncertainty, severity calls, handover, exec management. Requires two shadowed incidents and one simulated SEV1 before certification.
- **Communications Lead** (one day): status-page writing, customer tiering, legal boundaries, regulator triggers.
- Certification is valid 12 months and renewed via a simulation; the register of certified people is an audit artefact.
- Add on-call onboarding per team: a new joiner shadows two shifts before holding primary, and never holds primary in their first 90 days.
22. Simulation programme: game days, drills and wheel of misfortune (depends on: 21, 12, 20)
The process must be rehearsed before it meets a real SEV1. Simulations also build the commander pool and expose runbook gaps cheaply.
Run a standing calendar rather than one-off exercises.
- Monthly 60-minute tabletop ("wheel of misfortune") per engineering group, using a real past incident from the 31.
- Quarterly full-scale game day in production or a production-like environment: regional failover, ledger replica promotion, dependency failure, with the whole role structure activated and timed.
- Twice-yearly unannounced paging drill to measure real acknowledgement times at night.
- One security-incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Every exercise produces a lightweight postmortem and action items in the same system as real incidents.
- Measure and publish drill metrics: time to IC, time to first status update, time to correct mitigation decision.
23. Pilot with wave 0 teams (depends on: 22, 9, 11)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Run a six-week pilot with tight measurement and a public verdict.
- Select 5–6 teams: core payments, ledger/database, platform/Kubernetes, API gateway, plus two of the 12 teams already on-call.
- Activate the full stack for them: new severity scale, Duty IC rotation, single paging tool, alert budget, status-page policy, mandatory postmortems, paid on-call.
- Hold a weekly pilot retro; expect and document 20–40 process defects, and fix them in the standard before rollout.
- Validate the hard questions: does the 5-minute IC rule hold at 3 a.m.? Do cross-team pulls get answered? Is the severity tree unambiguous?
- Exit criteria: MTTD under 10 minutes for pilot services, IC assigned within 5 minutes in 95% of incidents, page volume down 50%, all postmortems on time, positive on-call sentiment.
- Publish a one-page pilot result to the whole company — this is the main adoption argument for the remaining teams.
24. Metrics, dashboards and the review cadence (depends on: 12, 18, 23)
Instrument the process itself, so improvement is visible and the audit has evidence of monitoring and review.
Define a small set of metrics with owners and a fixed meeting rhythm.
- Response metrics: MTTD, time to declare, time to IC assigned, MTTA, MTTM, MTTR, incidents per month by severity, % detected by customers first.
- Quality metrics: page volume per person per week, alert actionability rate, page budget breaches, postmortem on-time rate, action closure rate and ageing.
- Business metrics: SLA credits paid, availability against 99.95% per capability, error-budget consumption, repeat-incident rate.
- People metrics: on-call load distribution across teams, out-of-hours pages per person, on-call sentiment and attrition among on-call staff.
- Cadence: weekly Incident Review Board (postmortems and actions), monthly Reliability Review (metrics per team, alert hygiene, on-call load), quarterly Executive/Board review (credits, trends, investment asks), annual policy review.
- Every metric gets a target and a named owner; dashboards are self-serve and public inside the company.
25. Wave rollout across all 28 teams with readiness gates (depends on: 23, 24)
Roll out in four waves of six to eight teams, every three weeks, ordered by criticality. Each wave passes an explicit gate rather than a deadline.
Gates keep quality high and make the standard credible.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with 6+ certified responders, one IC candidate nominated, one drill passed.
- Assign each wave a named coach from the program team for three weeks of hands-on support.
- Gate criteria are checked and signed by the director; teams that fail are re-scheduled, not waived.
- Freeze legacy tooling per wave: after onboarding, the old alerting paths are disabled, not left as a fallback.
- Publish a live adoption scoreboard by team so progress is social, not administrative.
26. SOC 2 control mapping, evidence automation and internal dry-run audit (depends on: 18, 19, 25)
Design the process so audit evidence is a by-product of doing the work, then test it before the auditors do.
Engage the auditor early to confirm the interpretation of controls.
- Map the process to the Trust Services Criteria: CC7.3 and CC7.4 (incident identification, response, recovery), CC7.2 (monitoring), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), plus availability criteria A1.2.
- Produce and approve formal policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard — versioned, signed, annually reviewed.
- Automate evidence: incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, postmortem library, action-item closure reports, training and certification register, drill records.
- Confirm the observation window with the auditor and ensure the process is operating for a minimum of three months before fieldwork.
- Run an internal dry-run audit at month six: sample 15 incidents and walk the full evidence chain; fix gaps with 8 weeks to spare.
- Keep a remediation log for any incident where the process was not followed, with the corrective action — auditors respond better to documented exceptions than to a claim of perfection.
27. Change management, incentives and communications campaign (depends on: 3, 9, 23)
Run this in parallel from day one. The process will be judged by engineers on fairness, and by executives on visible results.
Communicate the deal explicitly and repeatedly.
- The deal in one sentence: **you are paid for on-call, you are paged only for what you own, a trained commander runs the incident, and your postmortem actions get real sprint capacity**.
- Launch communications: CTO all-hands, per-team roadshows, a one-page process card for laptops, an internal wiki hub and a Slack support channel with a 4-hour answer SLA.
- Recognition: incident-response contribution in promotion criteria and performance frameworks, quarterly awards for best postmortem and biggest alert-noise reduction, public thanks after every SEV1.
- Manager accountability: adoption, alert hygiene, action closure and on-call load in each engineering manager's quarterly objectives.
- Handle the exceptions: a written path for engineers who cannot do nights (caring responsibilities, health), covered by stipended volunteers elsewhere.
- Track sentiment quarterly and publish the results, including bad news, to keep credibility.
28. Program risk register and contingency planning (depends on: 1)
Name the ways this program fails and pre-commit the response. Review it monthly in the steering group.
The main risks are predictable.
- **Volunteer shortfall for the IC pool**: contingency is to make command a rostered duty for engineering managers and senior engineers until the pool reaches 25.
- **Compensation not approved in time**: fall back to time-off-in-lieu plus a phased stipend, but do not launch mandatory night on-call without some compensation.
- **Tool migration slipping**: keep the single-queue requirement and cut scope on the incident-management layer, not on paging consolidation.
- **Alert pruning causing a missed incident**: prune from paging to ticket first, observe for 30 days, then delete; keep a recovery path.
- **Burnout or attrition among the 12 experienced on-call teams**: monitor load weekly and cap individual page counts.
- **A major SEV1 mid-rollout**: pre-agree that the program lead becomes a full-time responder and the wave schedule slips by one wave, with the steering group informed the same day.
29. Continuous improvement, maturity roadmap and post-audit sustainability (depends on: 24, 25, 26)
Protect against the classic failure: the process decays once the audit is passed. Build the second-year plan before the first year ends.
Set a maturity model and a forward roadmap with owners.
- Quarterly process retrospective with the IC pool: what in the process itself slowed us down, what needs simplifying, is the severity scale calibrated?
- Re-baseline targets every six months; a process that hits all targets is under-ambitious.
- Year-two roadmap candidates: follow-the-sun coverage cell, automated mitigation and self-healing for the top three recurring causes, error-budget policy that gates releases, per-customer real-time impact reporting, and blast-radius reduction for the shared ledger cluster (the largest single structural risk).
- Move from lagging metrics (MTTR) to leading ones (error-budget burn, near-miss rate, drill performance).
- Make the annual policy review, certification renewal and drill calendar permanent calendar items owned by the Head of Reliability, independent of the audit cycle.
- Report to the board quarterly on availability, credits and incident trends so the process keeps executive attention after SOC 2 is signed.
Previous Proposal 2 (ID: 7ee53cee-6357-4ed7-bd99-535e92e8b1e6, Agent: gpt5.6-sol_initial_2, LLM: openai/gpt-5.6-sol):
Estimated Complexity: high
Success Metrics: - Within 7 days, every suspected SEV0–SEV2 has one incident record, one channel, and a named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named incident commander for more than 10 minutes.
- By day 30, 100% of Tier 0 and Tier 1 services have a named owner, primary escalation, secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and subject-matter coverage.
- By day 120, 100% of production services have sustainable ownership and tested escalation paths.
- At least 95% of SEV0 and SEV1 pages are acknowledged within 5 minutes by month 3.
- At least 95% of SEV2 pages are acknowledged within 10 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to 10 minutes by day 90 and 5 minutes by month 6.
- The proportion of incidents first detected by customers falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to below 90 minutes by day 120 and below 60 minutes by month 6.
- At least 95% of qualifying incidents meet their initial customer-communication deadline by month 3.
- At least 95% of published incidents meet their required update cadence by month 3.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without loss of Tier 0 or Tier 1 detection coverage.
- Monthly pages fall from 3,400 to no more than 1,500 by day 90, with alert actionability and missed-detection reviews used as countermeasures against unsafe suppression.
- 100% of new paging alerts satisfy the owner, runbook, dashboard, action, severity, and escalation quality rules by day 60.
- 100% of required SEV0 and SEV1 postmortems are drafted within 3 business days and reviewed within 5 business days by month 2.
- At least 90% of postmortem actions are completed by their approved due dates by month 6.
- All 53 currently open historical actions are triaged within 30 days; all unaccepted high-risk items are completed within 90 days.
- Repeat incidents with the same unaddressed contributing factor decline by at least 50% within 6 months.
- Every primary rotation has at least six trained responders or a documented, time-limited executive exception by day 120.
- No responder is routinely scheduled more frequently than one primary week in six by day 120.
- Two end-to-end cross-company exercises, including regional and ledger scenarios, are completed before the audit, with all critical findings assigned and tracked.
- Monthly availability meets or exceeds the 99.95% contractual target by month 6, with exceptions reviewed at the executive reliability meeting.
- SLA credits decline by at least 50% on an annualized trailing basis by month 8.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (19):
1. Establish ownership, authority, and funding
Launch the program within 48 hours under an executive sponsor. Give one program owner authority to standardize incident management across all 28 teams.
- Name the CTO or equivalent as executive sponsor and a Head of Incident Management or Reliability as directly accountable owner.
- Form a working group with Engineering, SRE or Platform, Product, Support, Customer Success, Communications, Security, Legal, Compliance, Risk, HR, Finance, and Internal Audit.
- Approve authority for an incident commander to stop deployments, roll back releases, disable features, shift traffic, invoke continuity plans, and pause payment processing when integrity is at risk.
- Preserve financial controls. The incident commander may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access requirements.
- Fund paging tools, compensation, training, observability work, exercises, and dedicated reliability capacity.
- Reserve engineering capacity for incident remediation. Start with 10% of capacity and adjust through quarterly risk reviews.
- Record the current baselines: 31 customer-impacting incidents, 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M in credits, 3,400 monthly alerts, 85% noise, and 11 of 64 actions closed.
- Maintain a risk register for staffing gaps, shared-ledger concentration, regional failover, alert coverage, third parties, and audit readiness.
2. Install immediate minimum controls (depends on: 1)
Put an interim process in place during the first seven days. Do not wait for tool consolidation, policy perfection, or the SOC 2 audit.
- Publish a one-page interim severity guide and incident declaration procedure.
- Establish one continuously monitored incident declaration path through chat, telephone, and the paging system.
- Create a standard incident channel, conference bridge, incident document, and event naming convention.
- Staff an interim primary and backup incident commander at all times. Compensate this duty retroactively under the final compensation policy.
- Give trained duty personnel access to the status page, paging system, dashboards, support queue, service catalog, and emergency contacts.
- Require an incident commander to be named within 10 minutes for every suspected major incident.
- Direct Support to escalate credible customer reports immediately rather than waiting for engineering confirmation.
- Triage all 53 open historical postmortem actions. Complete, re-plan, or formally risk-accept the items affecting ledger integrity, payment duplication, regional resilience, security, and detection first.
- Hold a daily 15-minute operational review until permanent controls are working.
3. Create the service and dependency catalog (depends on: 1)
Build a reliable ownership map for all production services and customer journeys. This is the basis for paging, escalation, impact assessment, and audit evidence.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, data stores, queues, external processors, banking partners, and customer-facing endpoints.
- Assign each component a single accountable team, primary responder group, secondary escalation group, engineering manager, and product owner.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 according to financial integrity, customer impact, dependency centrality, and contractual obligations.
- Treat the ledger, payment orchestration, authentication, settlement, reconciliation, and critical shared infrastructure as Tier 0 or Tier 1.
- Map every important customer journey to its service, database, cloud-region, and third-party dependencies.
- Record SLOs, RTOs, RPOs, data classification, dashboards, runbooks, deployment controls, feature flags, and failover methods.
- Assign separate but coordinated responders for the ledger application and the shared PostgreSQL platform.
- Document whether each service is active-active, active-passive, or region-bound. Identify dependencies that make nominal regional redundancy ineffective.
- Make missing ownership or missing runbooks a release-blocking risk for Tier 0 and Tier 1 services.
4. Adopt severity and incident lifecycle standards (depends on: 1)
Approve one impact-based severity model for operational, security, data, and third-party incidents. When evidence is incomplete, start at the higher credible severity and downgrade later.
- SEV0, crisis: Use for actual or credible unauthorized, lost, duplicated, or corrupted movement of money; ledger integrity loss; material security compromise; material data exposure; both-region failure; or an event likely to require crisis or regulatory management. Page all roles immediately, engage executives, Security, Legal, Compliance, and Risk, and consider pausing payment activity.
- SEV1, critical: Use for widespread inability to initiate, process, settle, or reconcile payments; a core journey failing without a viable workaround; material regional impact; fast SLA-budget exhaustion; or an imminent integrity risk. Staff all incident roles, notify the executive duty officer, and publish customer communications.
- SEV2, major: Use for a material customer subset, one or more critical customers, significant degradation with a workaround, partial transaction failure, or a likely contractual impact. Assign an incident commander and subject-matter responders; add communications and scribe roles whenever customers are affected.
- SEV3, minor: Use for localized, low-impact degradation with no financial-integrity, security, regulatory, or material contractual risk. The owning team leads the response and keeps an internal record; external communication is not normally required.
- Base severity on actual or credible impact, not the seniority of the reporter, number of alerts, or presumed complexity of the fix.
- Permit any employee to declare an incident. Only the incident commander may lower severity after recording the evidence and rationale.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of customer impact. Define resolution only after stability, backlog processing, transaction recovery, and required ledger reconciliation are complete.
- Start measurement from the earliest reliable indication of impact, including telemetry, customer reports, and partner notifications.
5. Define roles and sustainable 24x7 staffing (depends on: 3, 4)
Separate command from technical remediation. This allows trained commanders to coordinate any incident without asking engineers to debug code they do not own.
- Incident commander: Owns severity, priorities, role assignment, escalation, decision cadence, mitigation strategy, handoffs, and final closure. One person has command at a time.
- Communications lead: Owns internal notices, status-page updates, account-manager briefs, approved customer language, and coordination with Legal or regulators.
- Scribe: Maintains a timestamped timeline of observations, decisions, commands, owners, and status changes. Automation may assist but does not replace human validation for SEV0 and SEV1.
- Subject-matter responders: Diagnose and mitigate only services or domains for which they have accepted ownership, training, access, and runbooks.
- Executive duty officer: Removes organizational obstacles and approves exceptional business decisions. This role does not take command unless a formal transfer occurs.
- Security, Legal, Compliance, Support, Vendor Management, and Business Continuity join according to predefined triggers.
- Create a company-wide incident-command rotation with at least eight certified primary commanders and eight qualified backups. Use weekly rotations with explicit handoffs.
- Create similarly sustainable communications and scribe pools using Engineering Operations, Support, Customer Operations, and Communications personnel.
- Group service responders into approximately 8–12 coherent product or platform domains rather than creating 28 fragile rotations. Each domain rotation should normally contain at least six trained responders.
- Do not place an engineer into another team's responder pool without training, access, runbooks, shadow shifts, and explicit acceptance by both teams.
- Maintain dedicated database-platform and ledger-application escalation coverage for the shared PostgreSQL environment.
- Require distinct people for commander, communications, and primary technical lead during SEV0 and SEV1 incidents.
- Require a verbal and written handoff when any incident role changes. Record the exact time and the new role owner.
6. Implement compensation and fatigue safeguards (depends on: 5)
End unpaid on-call before expanding coverage. Treat availability, interrupted personal time, and overnight recovery as compensable work.
- Pay a fixed stipend for each primary on-call week and a secondary stipend equal to a defined percentage of the primary amount.
- Pay a higher holiday stipend. Apply overtime and call-out rules to non-exempt employees as required by law.
- Give exempt employees a minimum call-out credit or equivalent paid recovery time for material after-hours work.
- Provide a paid recovery day after prolonged overnight work, a SEV0, or a qualifying SEV1. Managers must arrange daytime coverage rather than expecting normal output.
- Have HR, Finance, and employment counsel publish dollar amounts, tax treatment, eligibility, and payroll procedures within 14 days. Apply the policy consistently across teams and locations.
- Target rotations no more frequent than one week in six. Exceptions require a time-limited staffing plan and executive risk acceptance.
- Avoid consecutive primary and secondary weeks. A person must not be primary for two simultaneous domain rotations.
- Track after-hours pages, sleep interruptions, swaps, missed acknowledgements, and reported burnout by rotation.
- Trigger a staffing or alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
- Permit responders to declare themselves temporarily unfit after overnight work without performance penalty.
7. Consolidate incident and paging tooling (depends on: 4)
Create one operational system of record while migrating safely from the six current alerting tools. Consolidation must reduce ambiguity without creating a monitoring gap.
- Select one enterprise paging and escalation platform and one integrated incident record.
- Initially ingest events from all six tools. Deduplicate, correlate, and route them through the new platform before retiring sources.
- Integrate paging with chat, conference bridges, ticket tracking, the service catalog, observability tools, and the customer status page.
- Automatically capture declaration time, acknowledgements, role assignments, severity changes, messages, decisions, mitigated time, and resolved time.
- Use role-based access, multifactor authentication, break-glass controls, immutable audit logs, and periodic access reviews.
- Provide mobile and telephone fallback paths if chat, identity, or the primary paging tool is unavailable.
- Test paging, escalation, status publication, and conference access every week.
- Retire a legacy alert path only after its signals have named owners, successful end-to-end tests, and at least two weeks of verified operation in the new platform.
8. Improve detection and enforce alert quality (depends on: 3, 7)
Shift detection toward customer journeys, payment outcomes, and ledger integrity. Infrastructure metrics alone will not solve the current customer-first detection problem.
- Instrument payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, and reporting with SLOs and business-level success metrics.
- Run external synthetic transactions and API checks from outside the production boundary and from both AWS regions.
- Monitor transaction failure rates, processing latency, queue age, unprocessed volume, reconciliation breaks, unexpected ledger balances, duplicate identifiers, regional asymmetry, and third-party response quality.
- Correlate application telemetry with Kubernetes, AWS, PostgreSQL, network, deployment, and feature-flag events.
- Route high-priority support cases and credible partner notifications into the same incident declaration path within five minutes.
- Define noise as a page that is duplicate, informational, unactionable, non-production, or requires no timely human action.
- Require every paging alert to name an owner, affected service, urgency, customer or SLO risk, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Send non-urgent conditions to a ticket queue rather than a pager.
- Run new alerts in shadow mode for at least seven days unless an emergency risk exception is approved. Test both firing and recovery behavior.
- Review any alert with less than 50% actionability or more than three firings in seven days within two business days.
- Never silently disable a noisy alert. Verify compensating detection, record the decision, and assign a correction owner first.
- Review alert actionability, false positives, missed detection, and page load with every responder group each month.
9. Codify acknowledgement and escalation paths (depends on: 3, 4, 5, 7, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts when a qualifying signal or customer report enters the system.
- For SEV0 and SEV1, page the owning primary immediately; page the secondary after five unacknowledged minutes; page the domain manager and company incident commander at 10 minutes; and engage the executive duty officer by 15 minutes.
- For SEV2, require primary acknowledgement within 10 minutes and incident-command assignment within 15 minutes. Escalate to the secondary and manager when either target is missed.
- For SEV3, require acknowledgement within 30 minutes when immediate production action is needed. Otherwise create a prioritized work item.
- Automatically page the company incident commander for any credible integrity or security concern, cross-team event, customer-visible Tier 0 failure, regional event, or unresolved ownership question.
- If impact remains unknown after 15 minutes, raise severity rather than waiting for certainty.
- Let the incident commander summon dependency owners, cloud support, database support, payment processors, banking partners, and vendors through maintained escalation contacts.
- Test vendor contacts and premium-support entitlements quarterly.
- Route alerts with no valid owner to the central command rotation, then treat the missing ownership record as a control defect.
- Require human acknowledgement. Delivery to a device or chat channel does not count.
- Record every missed acknowledgement, failed escalation, and manual contact workaround for review.
10. Standardize live incident execution (depends on: 4, 5, 7, 9)
Give responders one concise operating procedure for the first minutes through resolution. Prioritize limiting customer and financial harm before proving a root cause.
- Open a dedicated channel, bridge, incident record, and timeline immediately for SEV0 through SEV2.
- Have the incident commander state severity, known impact, current hypothesis, immediate objective, assigned roles, and next update time.
- Freeze unrelated production changes during SEV0 and SEV1 incidents. Record exceptions approved by the incident commander.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Use pre-approved runbooks for region failover, Kubernetes recovery, PostgreSQL failover, credential rotation, queue recovery, and payment suspension.
- Guard against split-brain, replay, duplication, and out-of-order processing during regional or database recovery.
- Require reconciliation and controlled backlog processing before declaring payment or ledger incidents resolved.
- Keep diagnosis and mitigation workstreams separate when enough responders are available.
- State decisions and owners aloud and in the incident record. Avoid unrecorded direct-message command paths.
- Require stability for a severity-specific observation period before closure. Reopen the incident if impact recurs during that period.
- Conduct an explicit operational handback to the owning team, Support, and Customer Success.
11. Standardize internal, customer, and regulatory communications (depends on: 4, 5, 7, 10)
Communicate known impact early without waiting for a root cause. Use approved facts, acknowledge uncertainty, and give the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish an initial customer status within 15 minutes when disclosure is operationally and legally appropriate, then update every 15 minutes.
- For SEV1, notify internal stakeholders within 15 minutes, publish an initial customer status within 15 minutes, and update at least every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send an initial notice within 30 minutes. Update at least every 60 minutes.
- Do not normally publish SEV3 events. Notify specifically affected customers if contracts or material impact require it.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. State affected capabilities, customer symptoms, workarounds, regions, and next update time.
- Do not speculate about root cause, blame, security scope, recovery time, or data integrity.
- Give account managers a single approved briefing and an affected-customer list. Prohibit contradictory or improvised incident explanations.
- Maintain templates for outages, delays, data-integrity investigation, third-party failure, regional failure, security events, and resolution.
- Issue a resolution notice only after operational recovery and required reconciliation. Provide a customer-facing incident summary within five business days for qualifying events.
- Have Legal and Compliance maintain a jurisdiction, regulator, sponsor-bank, network, cyber-insurer, partner, and contract notification matrix.
- Where applicable, explicitly track the current New York cybersecurity-event notification clock, including the 72-hour requirement, without assuming every incident is reportable.
- Have Legal record the reportability decision, decision time, evidence, approver, deadline, and submission confirmation.
- Allow Security or Legal to limit public detail during an active threat, but require the reason and an alternative stakeholder plan to be recorded.
- Coordinate service-credit calculations and contractual notices with Finance and Customer Success from the same incident record.
12. Make postmortems mandatory and actionable (depends on: 4, 7, 10)
Use postmortems to improve systems and controls, not to assign personal blame. Keep performance or misconduct processes separate from the learning review.
- Require a postmortem for every SEV0 and SEV1.
- Require one for a SEV2 that affected customers, incurred credits, breached an SLO or contract, involved financial or data integrity, repeated a prior failure, exposed a control gap, or lasted more than two hours.
- Permit incident command, Security, Compliance, or the service owner to require a review for a near miss.
- Produce a factual draft within three business days and hold the cross-functional review within five business days.
- Use one template covering executive summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, what worked, what failed, and lessons.
- Include why monitoring did or did not detect the event before customers.
- Avoid a single-root-cause assumption. Examine technical, organizational, process, dependency, testing, and incentive factors.
- Give every action one owner, due date, priority, expected risk reduction, verification method, and linked engineering item.
- Classify actions as containment due within 7 days, corrective work due within 30 days, or strategic work normally due within 90 days.
- Require director approval and documented residual-risk acceptance for overdue high-risk actions.
- Verify effectiveness after implementation. Closing a ticket without evidence does not close the action.
- Publish broadly useful reviews internally. Maintain access-restricted versions for security, privacy, personnel, or legally privileged details.
13. Measure performance and review it routinely (depends on: 7, 8, 11, 12)
Use outcome, process, quality, and human-sustainability measures together. Do not reward teams for suppressing declarations or hiding incidents.
- Measure detection time from first impact to first internal signal, declaration time, acknowledgement time, role-staffing time, mitigation time, resolution time, and recurrence.
- Report both median and 90th percentile. Break results down by severity, service tier, customer journey, region, detection source, and owning domain.
- Track customer-first detection, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, and role conflicts.
- Track availability, error-budget consumption, failed-payment volume, delayed value, reconciliation breaks, impacted customers, contractual breaches, and service credits.
- Track alert volume, actionability, duplicates, after-hours pages, missed pages, pages per responder, and tool-source distribution.
- Track required postmortems completed on time, actions completed by due date, action age, verified effectiveness, and repeat contributing factors.
- Track rotation size, on-call frequency, swaps, recovery days, attrition signals, and quarterly responder sentiment.
- Hold a weekly operational review for recent incidents, overdue actions, alert problems, and upcoming risk.
- Hold a monthly executive reliability review covering trends, investment decisions, accepted risks, and SLA exposure.
- Hold a quarterly resilience and control review with Security, Compliance, Risk, Internal Audit, and Product leadership.
- Use team scorecards to direct investment and assistance, not individual performance penalties.
- Reconcile dashboard data against a monthly sample of incident records and customer cases to detect metric gaming or missing incidents.
14. Train and certify participants (depends on: 4, 5, 9, 10, 11, 12)
Train people before assigning full independent duty. Use paid working time for training, shadowing, exercises, and certification.
- Train all employees to recognize impact, declare an incident, and find the incident channel and status page.
- Train engineers and Support on severity, escalation, customer-report handling, evidence preservation, and financial-integrity precautions.
- Certify incident commanders through instruction, tabletop exercises, shadow incidents, and observed command performance.
- Train communications leads in status writing, contractual communications, regulator escalation, and avoiding unsupported claims.
- Train scribes in timestamping, decision capture, evidence hygiene, and separating fact from hypothesis.
- Require subject-matter responders to demonstrate dashboard, runbook, rollback, failover, and access competence for their assigned domain.
- Add training to new-hire onboarding and repeat role-specific certification annually.
- Appoint an incident-management champion in each of the 28 teams to collect feedback and support local adoption.
- Conduct listening sessions focused on pager fairness, cross-team boundaries, tooling friction, and psychological safety.
- Publish that command duty is process coordination, not responsibility for understanding or repairing another team's code.
15. Pilot and expand the on-call model (depends on: 3, 5, 6, 8, 9, 14)
Pilot the model on the highest-risk customer journeys before expanding it. Correct staffing, alert, access, and compensation defects at each stage gate.
- Start with the ledger, payment orchestration, Kubernetes platform, PostgreSQL platform, authentication, settlement, and Support intake.
- Run the command, communications, and domain rotations in parallel with existing paths for two weeks.
- Verify primary and secondary coverage, handoffs, access, runbooks, paging, conference access, status publication, and compensation processing.
- Require at least two shadow shifts before independent primary duty.
- Review every pilot page within one business day for routing accuracy, actionability, responder load, and missing context.
- Expand by customer journey and dependency domain, not by arbitrary team order.
- Provide central first-line triage if useful, but keep technical remediation with the accepted service owner.
- Do not use contractors or a managed service as the sole incident commander or sole owner of payment and ledger remediation.
- Permit temporary shared domain rotations only after service owners document training, access, runbooks, and escalation boundaries.
- Set an executive-reviewed deadline and remediation plan for any production service that cannot provide sustainable 24x7 ownership.
16. Exercise regional, ledger, and communication failures (depends on: 10, 11, 14, 15)
Validate the process under realistic conditions before relying on it. Begin in tabletop and staging environments, then use controlled production tests where risk permits.
- Run a company-wide incident-command tabletop within 30 days of policy approval.
- Exercise loss of one AWS region, Kubernetes control-plane degradation, shared PostgreSQL failure, payment-processor failure, queue backlog, credential compromise, and suspected duplicate payments.
- Exercise a simultaneous operational and security event to test command boundaries and disclosure control.
- Exercise status-page failure and loss of the primary chat or paging provider.
- Exercise overnight staffing, role handoff, executive escalation, account-manager messaging, and a potential regulator-notification decision.
- Validate backups, restore procedures, RPO, RTO, failover prerequisites, and post-recovery reconciliation.
- Do not inject uncontrolled changes into the production ledger. Use replicas, staging, simulations, or tightly governed production tests.
- Record exercise observations as tracked actions under the same ownership and due-date rules as incident actions.
- Run at least one domain exercise per quarter and two cross-company exercises before the SOC 2 audit.
17. Execute a time-boxed enterprise rollout (depends on: 2, 6, 7, 11, 12, 13, 15)
Use fixed implementation waves so the audit deadline does not become the start date. Report progress weekly and escalate missed stage gates as business risks.
- Days 0–7: Establish governance, interim command coverage, one declaration path, provisional severity, and daily operational reviews.
- By day 14: Approve the core policy, role definitions, communications timings, compensation design, and historical-action triage.
- By day 30: Complete Tier 0 ownership, certify the first command roster, begin the on-call pilot, enable standard incident records, and run the first tabletop.
- By day 60: Provide 24x7 coverage for all Tier 0 and Tier 1 customer journeys, integrate the six alert sources, and enforce postmortem tracking.
- By day 90: Migrate critical paging, implement customer-journey detection, complete status and regulatory playbooks, and materially reduce alert noise.
- By day 120: Assign sustainable ownership and escalation for every production service and complete the first controlled regional or continuity exercise.
- By day 180: Complete tool retirement decisions, verify action closure, rerun weak scenarios, and demonstrate improving detection and mitigation trends.
- In month 7: Conduct a mock audit and executive readiness review, leaving at least one month to correct evidence or operating defects.
- Use exception records with owners, expiry dates, compensating controls, and executive approval. Do not allow indefinite verbal exceptions.
18. Build SOC 2 evidence as the process operates (depends on: 1)
Design evidence collection at the start rather than reconstructing it before the audit. Demonstrate both control design and sustained operation.
- Map the incident process to applicable SOC 2 criteria with Compliance and the auditor, including detection, response, communication, change management, access, availability, and corrective action.
- Maintain approved, version-controlled policies, procedures, severity definitions, role descriptions, and exception records.
- Preserve rotation schedules, compensation activation, training attendance, certification, paging tests, access reviews, and exercise results.
- Preserve incident declarations, timestamps, role assignments, communications, decisions, status updates, postmortems, and corrective-action evidence.
- Record regulatory and contractual notification assessments, including decisions that no notification was required.
- Define retention, confidentiality, legal-hold, and access requirements for operational and security records.
- Sample evidence monthly and trace incidents from initial signal through action verification.
- Have Internal Audit or an independent control owner test the process in months 4 and 6.
- Correct control failures through tracked actions rather than editing historical records.
- Conduct the formal mock audit in month 7 using the same evidence populations expected for the external audit.
19. Sustain accountability and continuous improvement (depends on: 13, 17, 18)
Make incident management an operating discipline rather than an audit project. Keep policy, staffing, tools, and investment aligned with changing customer and system risk.
- Assign permanent owners for the incident policy, paging platform, status page, service catalog, training program, and metrics.
- Review severity thresholds, communication timings, compensation, and staffing at least annually and after material incidents.
- Use incident trends to prioritize architectural work on the shared ledger, regional independence, deployment safety, dependency isolation, and graceful degradation.
- Review repeat incidents and repeat contributing factors quarterly. Require executive action when remediation repeatedly loses priority.
- Survey responders quarterly and publish actions addressing fatigue, fairness, psychological safety, and tool friction.
- Recognize effective incident leadership, early declaration, useful postmortems, and preventive work.
- Prohibit retaliation for good-faith incident declaration or escalation.
- Provide the board or risk committee a quarterly summary of severe incidents, SLA exposure, regulatory events, overdue high-risk actions, and resilience investment.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
Previous Proposal 3 (ID: c051663c-fdf6-4324-bfbe-a5f42ed62eed, Agent: qwen3.8-max_initial_3, LLM: alibaba/qwen3.8-max):
Estimated Complexity: high
Success Metrics: - **MTTD reduced from 22 minutes to <5 minutes** within 6 months of full rollout.
- **Customer-first detection rate reduced from 40% to <10%** within 6 months.
- **MTTM for SEV1 incidents reduced from 3 h 10 min to <60 minutes** within 9 months.
- **Monthly SLA credit payouts reduced from ~$325K to <$100K per quarter** within 12 months.
- **Alert volume reduced from 3,400/month to <600 actionable alerts/month** within 6 months; signal-to-noise ratio >80%.
- **Postmortem completion rate: 100% of SEV1/SEV2 incidents** have a blameless postmortem within 5 business days.
- **Postmortem action-item completion rate >90% within 30 days** of the postmortem (up from ~17%).
- **Zero incidents with >15 minutes of unowned command** (down from 2 incidents with >1 hour).
- **100% on-call coverage**: all 28 teams staffed with primary + secondary on-call 24×7 within 14 weeks.
- **On-call compensation adopted**: 100% of on-call engineers receiving stipends and page pay; on-call satisfaction score ≥4/5 in quarterly survey.
- **Status-page first-update within 15 minutes for SEV1 and 30 minutes for SEV2**, 100% compliance.
- **SOC 2 Type II audit passed** at month 8 with zero incident-response findings.
- **All 260 engineers trained**; 56+ certified ICs and 28+ certified CLs active within 14 weeks.
- **Six alerting tools consolidated to one** within 6 months; legacy tools decommissioned.
- **Regulator notification process tested**: at least one tabletop exercise includes a NY DFS / FinCEN notification drill, and the legal/compliance playbook is documented and approved.
- **Quarterly IMC reviews held consistently** with published KPI dashboards and action-item tracking.
- **On-call participation resistance resolved**: <10% of engineers report 'unwilling to participate' in the 6-month pulse survey (baseline to be measured in S1).
Steps (13):
1. Assess Current State and Baseline Metrics
Build an evidence-based picture of the current incident management reality before designing anything new.
Collect and catalog the last 12 months of incident data: all 31 customer-impacting incidents, 3,400 monthly alerts, on-call coverage gaps across the 28 teams, and the 11 of 64 closed postmortem action items. Interview one lead from each of the 28 teams to surface pain points, political concerns (the 'carrying a pager for other teams' pushback), and tool sprawl.
Deliverables to produce:
- **Alert inventory**: which of the six alert tools feed which teams, alert volume per team, noise rate per tool, and overlap between tools.
- **Incident timeline analysis**: median detection-to-notification-to-mitigation-to-resolution times, who detected first (internal vs. customer), who mitigated, and where handoff gaps occurred.
- **On-call coverage map**: which 12 teams have on-call, which 16 do not, rotation length, compensation status, and escalation paths (or lack thereof).
- **Postmortem audit**: format variance, action-item tracking gaps, and the two incidents with no clear owner for over an hour.
- **Tooling and integration audit**: Kubernetes observability stack, six alerting tools, status-page provider, communication channels (Slack, email, phone), and any existing runbooks.
- **Compliance gap analysis**: SOC 2 Type II CC7.3/CC7.4 requirements vs. current practice, with a risk register for the eight-month window.
- **Peer benchmarking**: incident management practices at 3–4 comparable B2B fintech platforms (e.g., Plaid, Stripe, Adyen) for severity scales, on-call comp, and MTTR targets.
2. Secure Executive Sponsorship and Form the IM Governance Body (depends on: 1)
Anchor the program with visible, top-down authority so that 28 teams adopt changes they did not individually request.
The CEO email about 'outages we hear about from clients' is a ready-made mandate. Convert it into a formal sponsorship structure.
- Appoint an **executive sponsor** (CTO or VP Engineering) who owns the program end-to-end and reports to the CEO monthly.
- Create an **Incident Management Office (IMO)**: one dedicated senior incident management lead, one tooling/platform engineer, and one part-time data analyst.
- Establish an **Incident Management Council (IMC)**: one engineering manager from each of the 28 teams, plus the VP of Customer Success, a compliance lead, and a security lead. The IMC meets bi-weekly during rollout, monthly thereafter.
- Draft and circulate an **executive mandate memo** that states: incident response is a shared operational obligation, not a per-team favor; participation in on-call rotations is a condition of employment for production-facing roles; and the program is not optional pending the SOC 2 audit.
- Allocate a dedicated budget line for on-call compensation, tooling consolidation, status-page licensing, training, and external facilitation.
3. Define Severity Levels and Automatic Triggers (depends on: 1)
Replace the current ad-hoc triage with a five-level severity taxonomy that every engineer, support agent, and account manager can apply in under 60 seconds.
- **SEV1 – Critical**: Ledger data corruption or loss, complete payment processing halt, confirmed data breach affecting customer PII or funds, or regulatory reporting breach. Triggers: automatic all-hands page to on-call, CTO + CEO paged within 10 minutes, dedicated bridge call within 5 minutes, status-page update within 15 minutes, regulator notification assessment within 1 hour, customer comms within 30 minutes.
- **SEV2 – High**: Payment processing degraded >30% throughput or >5% error rate, single-region failover failure, ledger read-only mode, or any condition likely to breach the 99.95% SLA within the current window. Triggers: primary + secondary on-call paged, incident commander assigned within 10 minutes, bridge call within 15 minutes, status-page update within 30 minutes, VP Engineering notified within 20 minutes.
- **SEV3 – Medium**: Non-critical service degradation affecting <30% of customers, single non-ledger microservice outage with working fallback, or elevated latency above SLA threshold but below halt. Triggers: primary on-call paged, IM notified within 30 minutes, status-page update within 1 hour if customer-visible, daily standup update.
- **SEV4 – Low**: Degraded internal tooling, minor UX bug with workaround, non-customer-facing alert. Triggers: next-business-day response, ticket created, no page unless on-call agrees.
- **SEV5 – Informational / Noise**: Cosmetic issues, planned-maintenance notifications, alert misfires logged for tuning. Triggers: no page, logged for weekly alert-quality review.
Define **escalation rules**: any SEV3 unresolved after 4 hours auto-escalates to SEV2; any SEV2 unresolved after 2 hours auto-escalates to SEV1. Severity can be **downgraded** only by the incident commander with IMC notification.
Publish the taxonomy as a one-page decision tree, a Slack slash-command (`/sev`), and an integration into the alerting tool so that every alert carries a suggested severity.
4. Define Incident Roles and Staffing Model (depends on: 3)
Codify four mandatory roles for every SEV1/SEV2 incident and optional roles for SEV3, then solve the 24×7 staffing problem across 28 teams.
**Roles**
- **Incident Commander (IC)**: owns the incident end-to-end, declares severity, assigns tasks, authorizes mitigations, decides when to escalate or stand down. Never writes code during the incident.
- **Communications Lead (CL)**: owns status-page updates, internal Slack channels, account-manager briefings, and regulator notifications. Separate from the IC so the IC can focus on mitigation.
- **Scribe / Timeline Keeper**: logs every decision, action, and timestamp in the incident channel and the incident-management tool. Produces the raw timeline for the postmortem.
- **Subject-Matter Responders (SMRs)**: 1–3 engineers from the owning team(s) who diagnose and fix. For the shared PostgreSQL ledger, a dedicated DBA responder is always required.
**24×7 Staffing via a Three-Tier Follow-the-Sun Model**
- **Tier 1 – Front-line on-call**: Primary + secondary responder per team, paged first. Covers the team's own services.
- **Tier 2 – Platform / SRE on-call**: A dedicated 6-person SRE rotation covering cross-cutting infrastructure: Kubernetes, the shared PostgreSQL ledger, networking, and the two AWS regions. This tier directly addresses the 'carrying a pager for other teams' concern by absorbing infrastructure incidents.
- **Tier 3 – IMC escalation**: Engineering managers and the IMO on-call for multi-team or SEV1 incidents. Provides the IC and CL when no team-level IC is available.
**Follow-the-Sun**: If any engineering hub exists in a second timezone, use it for overnight Tier 1 coverage. If not, partner with a managed on-call service for overnight first-response triage (severity declaration + paging the correct team), reducing 3 a.m. pages for NY-based engineers.
**IC and CL pools**: Nominate at least 2 ICs and 1 CL per team (56 ICs, 28 CLs minimum). ICs are trained and certified before they rotate. For SEV1 incidents, the IC must be a certified IC from the IMC pool, not just 'whoever is around'.
**Ledger-specific rule**: Because the PostgreSQL ledger is shared, a **Ledger Duty Officer** from the SRE Tier 2 is always on the bridge for any incident touching ledger services, regardless of which team owns the failing microservice.
5. Design On-Call Rotations, Compensation, and Alert-Quality Rules (depends on: 4)
Make on-call sustainable, fairly compensated, and free of alert noise so engineers stop resisting participation.
**Rotation Design**
- 7-day rotations, one primary + one secondary per team per week. No engineer is on-call more than one week in four.
- Minimum 48-hour rest between rotations. No on-call during approved PTO.
- All 28 teams participate. Teams without current on-call get a 90-day ramp with a shadow rotation before going live.
- Tier 2 SRE rotation: 6 engineers, one week on / five weeks off, with a dedicated backup.
**Compensation Package**
- **Base on-call stipend**: $500 per week of primary on-call, $250 for secondary, paid regardless of whether pages fire.
- **Page pay**: $75 per acknowledged page outside business hours; $150 if the page leads to active incident work.
- **Time-off-in-lieu (TOIL)**: Any engineer who works >4 hours overnight (00:00–06:00 local) gets a full TOIL day. >8 hours in a single incident gets 1.5 TOIL days.
- **SEV1 bonus**: $300 flat bonus for every engineer who actively works a SEV1 incident, paid within the next pay cycle.
- **Annual on-call cap**: No engineer exceeds 13 weeks of on-call per year. Exceeding the cap triggers a mandatory team-staffing review.
- Budget estimate: ~$420K/year for stipends and page pay across 28 teams; present this to the CFO as a fraction of the $1.3M annual SLA credit cost.
**Alert-Quality Rules (the '85% noise' problem)**
- Every alert must carry: owning team, suggested severity, runbook link, and a 30-day noise score.
- **Alert budget**: each team gets a maximum of 100 actionable alerts per month. Exceeding the budget triggers a mandatory alert-tuning session with the IMO.
- **Noise threshold**: any alert that fires >10 times in 7 days with no human action is auto-flagged for suppression or tuning within 14 days.
- **Alert review cadence**: weekly 30-minute alert-quality review per team; monthly cross-team alert review in the IMC.
- **Sunset rule**: alerts with no runbook are demoted to SEV5 after 30 days and suppressed after 60 days unless a runbook is written.
- Target: reduce monthly alert volume from 3,400 to <600 actionable alerts within 6 months.
6. Build Detection, Escalation, and Communication Paths (depends on: 3, 4)
Eliminate the 22-minute median detection gap and the 40% customer-first-detection rate with layered monitoring and a single escalation spine.
**Detection Layers**
- **Synthetic transactions**: run a payment end-to-end through the full stack (API → service → ledger → confirmation) every 60 seconds from both AWS regions. Alert if latency >2× baseline or any step fails. This catches what per-service metrics miss.
- **Customer-traffic anomaly detection**: monitor API error rates, payment success rates, and latency percentiles per customer cohort. Alert on >2σ deviation.
- **SLO-based alerting**: define SLIs for the 99.95% SLA (availability, latency p99, ledger consistency). Alert when error budget burn rate exceeds threshold, before the SLA actually breaches.
- **Infrastructure health**: Kubernetes node/pod health, PostgreSQL replication lag, disk I/O, and cross-region latency.
- **Support-ticket spike detection**: if >5 customers open tickets about the same symptom within 10 minutes, auto-create a SEV3 candidate.
**Escalation Path**
- Alert fires → PagerDuty routes to Tier 1 primary → 5-min no-ack → Tier 1 secondary → 10-min no-ack → Tier 2 SRE → 15-min no-ack → IMC on-call manager → 20-min no-ack → VP Engineering auto-page.
- Any SEV1 declaration auto-pages the CTO, opens a dedicated Slack channel + Zoom bridge, and notifies the CL.
- **No incident goes unowned for >15 minutes.** If no IC is assigned by minute 15, the IMC on-call manager assumes IC role by default.
**Internal Communications**
- Dedicated Slack channels: `#inc-sev1`, `#inc-sev2`, `#inc-sev3` (auto-created per incident), plus `#inc-updates` for broadcast.
- IC posts a structured update every 15 minutes (SEV1), 30 minutes (SEV2), 1 hour (SEV3) into the incident channel.
- CL posts a summary to `#inc-updates` and notifies relevant engineering managers.
**Customer Communications**
- **Status page**: auto-updated via API. SEV1: first update within 15 minutes, then every 30 minutes until resolved. SEV2: first update within 30 minutes, then every hour. SEV3: within 1 hour if customer-visible.
- **Account managers**: CL briefs AMs via a dedicated Slack channel within 30 minutes (SEV1) or 1 hour (SEV2). AMs contact their top-20 revenue accounts directly.
- **Customer email/SMS**: for SEV1 and SEV2, automated notification to all 2,100 customers via the status-page subscription system within 30 minutes.
- **Regulator notification**: Legal/Compliance assesses within 1 hour whether NY DFS, FinCEN, or card-network notification is required. If yes, file within the regulatory deadline (typically 72 hours for NY DFS cybersecurity events). Log the decision and filing in the incident record.
**Post-resolution**: CL publishes a 'resolved' update within 30 minutes of mitigation. For SEV1/SEV2, a preliminary customer-facing RCA summary is published within 5 business days.
7. Standardize Postmortems with Tracking and Accountability (depends on: 3, 6)
Fix the '11 of 64 action items closed' problem with a mandatory, uniform, blameless postmortem process backed by engineering-manager accountability.
**When Mandatory**
- All SEV1 and SEV2 incidents: postmortem required within 5 business days.
- SEV3 incidents: postmortem required if >50 customers affected, if the incident lasted >4 hours, or if it was customer-detected.
- SEV4/SEV5: optional, but any recurring SEV4 (≥3 times in 30 days) triggers a mandatory review.
**Format (single template, enforced by tooling)**
- Incident summary (severity, duration, customers affected, revenue impact, SLA credit exposure).
- Timeline (auto-generated from scribe notes + alert timestamps).
- Detection analysis: how was it detected, why did it take X minutes, could it have been faster.
- Root cause analysis using 5-Whys or fault-tree, not blame.
- Contributing factors (process, tooling, staffing, knowledge gaps).
- Impact quantification: customers affected, transactions failed, SLA credits triggered.
- Action items: each with a **named owner**, **due date**, **priority**, and **ticket in Jira**.
- Lessons learned and what went well.
**Blameless Review Meeting**
- Held within 5 business days, facilitated by the IMO or a trained facilitator (never the IC of that incident).
- All responders, the CL, relevant engineering managers, and an IMC representative attend.
- Ground rules: focus on system and process failures, not individual mistakes. The facilitator enforces this.
- Meeting recorded; notes published to the engineering-wide wiki within 48 hours.
**Action-Item Tracking and Accountability**
- Every action item is created as a Jira ticket with a due date and a named owner.
- **Engineering managers are accountable**: action-item completion is a standing agenda item in the bi-weekly IMC meeting. Any action item >7 days overdue is escalated to the VP Engineering.
- **Completion gate**: no team may close its postmortem until 100% of its action items have Jira tickets. Postmortem is 'closed' only when all tickets are resolved.
- **Quarterly audit**: the IMO audits action-item completion rates and reports to the IMC and the executive sponsor. Target: >90% completion within 30 days of the postmortem.
- Link postmortem quality and action-item completion to team health scores and engineering-manager performance reviews.
8. Define KPIs, Dashboards, and Governance Reviews (depends on: 7)
Create a measurable feedback loop so leadership can see whether the program is working and where to intervene.
**Primary KPIs (tracked weekly, reported monthly)**
- **MTTD** (Median Time to Detect): target <5 minutes (from 22).
- **MTTA** (Median Time to Acknowledge): target <5 minutes.
- **MTTM** (Median Time to Mitigate): target <60 minutes for SEV1 (from 3 h 10 min), <4 hours for SEV2.
- **Customer-first detection rate**: target <10% (from 40%).
- **SLA compliance**: maintain 99.95%; track monthly SLA credit payouts, target <$100K/quarter (from ~$325K/quarter).
- **Alert signal-to-noise ratio**: target >80% actionable (from ~15%).
- **Alert volume**: target <600/month (from 3,400).
- **Postmortem completion rate**: 100% for SEV1/SEV2 within 5 business days.
- **Action-item completion rate**: >90% within 30 days (from ~17%).
- **On-call health**: pages per engineer per week (target <5), TOIL usage, on-call satisfaction survey score.
- **Unowned incident duration**: target 0 incidents with >15 minutes without an IC.
**Dashboards**
- Real-time operational dashboard (Grafana): current incidents, active alerts, on-call roster, SLA error-budget burn.
- Weekly leadership dashboard (auto-generated): KPI trends, open action items, alert-noise report, on-call load distribution.
- Quarterly IMC scorecard per team.
**Review Cadence**
- **Weekly**: IMO publishes KPI snapshot to `#inc-updates`.
- **Bi-weekly IMC**: review open incidents, overdue action items, alert-quality exceptions, and on-call load.
- **Monthly executive review**: CTO presents KPI trends, SLA credit cost, and risk register to the CEO.
- **Quarterly incident-management review**: deep-dive into trends, training gaps, tooling needs, and process improvements. Output fed into the next quarter's roadmap.
9. Consolidate Tooling and Build the Incident Management Platform (depends on: 2, 3)
Replace six alerting tools and ad-hoc status-page updates with a single, integrated incident management stack.
**Target Tool Architecture**
- **Single alerting and on-call platform** (e.g., PagerDuty or Opsgenie): ingest all alerts, apply severity routing, manage on-call schedules, handle escalations, and send pages. Retire the other five tools within 6 months.
- **Observability consolidation**: standardize on one APM/metrics stack (e.g., Datadog or Grafana Cloud) for all 180 Kubernetes services across both AWS regions. Ensure the shared PostgreSQL ledger has dedicated dashboards.
- **Status page**: a dedicated, branded status page (e.g., Statuspage.io or Instatus) with API integration for auto-updates. Subscribe all 2,100 customers.
- **Incident coordination tool**: integrate incident-management workflows into Slack (auto-create channels, invite responders, post templates) and a dedicated incident record system (e.g., Jira Service Management, incident.io, or Rootly) for timelines, postmortems, and action-item tracking.
- **Runbook repository**: a central wiki (Confluence or Notion) with mandatory runbooks for every alert. No alert goes live without a linked runbook.
**Implementation Tasks**
- Migrate all 28 teams' alert rules into the single platform in three waves (highest-volume teams first).
- Build the severity-based routing rules and escalation policies per S3 and S6.
- Automate status-page updates triggered by severity declaration.
- Build the synthetic-transaction monitor and SLO-based alerting per S6.
- Integrate Jira for automatic action-item ticket creation from postmortems.
- Decommission legacy tools only after all teams have completed training on the new stack.
- Budget: allocate $150K–$250K/year for licensing, plus engineering time for migration.
10. Prepare for the SOC 2 Type II Audit (depends on: 7, 8, 9)
Ensure the incident management process produces the evidence the auditor will need, well before the audit window opens in eight months.
**SOC 2 Requirements to Address (CC7.3, CC7.4, CC7.5)**
- Documented incident response procedures (the severity taxonomy, role definitions, communication templates).
- Evidence of incident detection, response, and recovery for every SEV1/SEV2 incident during the audit period.
- Postmortem records with action-item tracking.
- On-call schedules, training records, and escalation evidence.
- Status-page update logs and customer notification records.
- Regulator notification logs (if any).
**Preparation Tasks**
- The IMO maintains a **SOC 2 evidence folder**: every incident record, postmortem, action-item ticket, status-page update, and training completion certificate is stored and indexed.
- Conduct a **mock SOC 2 audit** at month 5: an internal or external auditor reviews the incident management process end-to-end and identifies gaps.
- Remediate mock-audit findings before month 7.
- Ensure the incident management tool retains all records for at least 12 months (the SOC 2 Type II observation window).
- Document the **chain of custody** for incident records: who accessed, modified, or closed each record.
- Prepare a **narrative document** describing the incident management process, roles, and controls for the auditor.
- Coordinate with the compliance lead to align incident management evidence with the broader SOC 2 scope (access controls, change management, etc.).
11. Design and Deliver Training, Runbooks, and Change Management (depends on: 4, 5, 9)
Equip all 260 engineers, 28 team leads, account managers, and support staff with the knowledge and muscle memory to execute the new process.
**Training Tracks**
- **All 260 engineers** (2-hour session): severity taxonomy, how to acknowledge a page, how to join an incident bridge, how to hand off to an IC, and how to write a postmortem contribution. Delivered in team-level sessions over 4 weeks.
- **IC pool (56+ engineers)** (8-hour certification): incident command techniques, severity declaration, escalation decision-making, bridge facilitation, and blameless postmortem facilitation. Includes two tabletop exercises. Certification valid for 12 months, renewed annually.
- **CL pool (28+ staff)** (4-hour session): status-page writing, customer communication templates, regulator notification triggers, and AM briefing protocol.
- **Account managers and support staff** (1-hour session): how to read the status page, how to escalate a customer report into an incident, and what information to collect.
- **SRE Tier 2** (16-hour onboarding): Kubernetes and PostgreSQL ledger deep-dive, cross-region failover runbooks, and escalation authority.
**Runbooks**
- Every alert must have a runbook before it is routed to on-call. The IMO provides a runbook template and audits compliance weekly.
- Priority runbooks to write first: shared PostgreSQL ledger failover, Kubernetes cluster degradation, payment-processing pipeline failure, cross-region failover, and ledger data-integrity check.
- Runbooks are peer-reviewed and version-controlled.
**Change Management for Adoption**
- Address the 'carrying a pager for other teams' concern directly: publish an FAQ explaining the three-tier model, the SRE Tier 2 absorbing cross-team infrastructure, the compensation package, and the TOIL policy.
- Run **office hours** weekly for the first 8 weeks where any engineer can ask questions or raise concerns.
- Identify **team champions**: one engineer per team who volunteers as an early adopter and peer mentor.
- Publish a **weekly 'incident management newsletter'** during rollout: what changed, what improved, KPI trends, and success stories.
- Make on-call participation a documented expectation in job descriptions and performance reviews for production-facing roles.
12. Execute Phased Rollout, Tabletop Exercises, and Continuous Improvement (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11)
Introduce the process in three waves so teams are not overwhelmed, then validate with exercises and iterate continuously.
**Phase 1 – Weeks 1–6: Foundation**
- Publish the severity taxonomy, role definitions, and communication protocols (S3, S4, S6).
- Launch the single alerting platform for the 12 teams already on-call; begin migration for the other 16.
- Activate the SRE Tier 2 rotation for the shared PostgreSQL ledger and cross-cutting infrastructure.
- Deploy the status page and test the API integration.
- Begin IC and CL training (first cohort of 20 ICs, 10 CLs).
- Publish the on-call compensation package; HR integrates stipends into payroll.
- Write the top 10 priority runbooks.
**Phase 2 – Weeks 7–14: Expansion**
- All 28 teams live on the single alerting platform; legacy tools in read-only mode.
- All 28 teams on the on-call rotation schedule (the 16 new teams in shadow mode for the first 4 weeks).
- Second and third IC/CL training cohorts completed.
- First **tabletop exercise**: simulate a SEV1 ledger corruption scenario with all roles, test the escalation path, status-page updates, and AM briefings. Debrief and fix gaps.
- Postmortem template and Jira integration live; all new incidents use the standard process.
- Alert-tuning sprint: each team reduces its alert volume by 50%.
**Phase 3 – Weeks 15–24: Optimization**
- All teams fully live; legacy alerting tools decommissioned.
- Second **tabletop exercise**: simulate a SEV1 cross-region failure with regulator notification.
- First quarterly IMC review with full KPI dashboard.
- Mock SOC 2 audit (month 5) and remediation.
- Retrospective on the rollout: survey all 260 engineers for feedback, adjust compensation or rotation rules if needed.
- Establish the **continuous improvement cadence**: quarterly process review, annual severity-taxonomy review, and annual on-call compensation benchmarking.
**Ongoing Governance**
- The IMC owns the process document and approves changes.
- The IMO tracks all KPIs and reports to the CTO monthly.
- Any process change requires IMC approval and a 2-week notice period before enforcement.
- Annual external benchmarking against peer B2B payments platforms.
13. Establish Ongoing Governance, Annual Review, and Audit Readiness Cycle (depends on: 12)
Embed incident management as a permanent organizational capability, not a one-time project.
- **Annual process review**: the IMC reviews the severity taxonomy, role definitions, on-call structure, and compensation against industry benchmarks and internal KPIs. Update as needed.
- **Bi-annual tabletop exercises**: one SEV1 infrastructure scenario, one SEV1 data-breach/regulator scenario. Rotate the IC and CL assignments so everyone gets practice.
- **Quarterly alert-quality audit**: the IMO reviews alert volumes, noise rates, and runbook coverage across all 28 teams.
- **On-call health survey**: quarterly anonymous survey measuring burnout, fairness, and compensation satisfaction. Results reviewed by the IMC.
- **SOC 2 readiness cycle**: begin evidence collection immediately after each audit ends. The IMO maintains a rolling evidence folder. Mock audit at month 5 of every 12-month audit cycle.
- **Postmortem maturity tracking**: track the action-item completion rate monthly. If it drops below 80%, the VP Engineering intervenes.
- **Incident management maturity model**: adopt a 5-level maturity model (ad-hoc → defined → managed → optimized → predictive). Assess annually. Target: Level 3 within 12 months, Level 4 within 24 months.
- **Budget review**: annually review on-call compensation, tooling costs, and training budget against the reduction in SLA credits and incident frequency.
Previous Proposal 4 (ID: 31fbbc6c-8c88-4abb-9c6d-b67706d9da75, Agent: grok4.6_initial_4, LLM: xai/grok-4.6):
Estimated Complexity: high
Success Metrics: - Median time to detect customer-impacting incidents ≤ 5 minutes within 6 months of go-live.
- Share of SEV-1/SEV-2 incidents first detected by customers ≤ 5% (from 40%).
- Median time to mitigate SEV-1/SEV-2 ≤ 45 minutes (from 3 h 10 min).
- Named Incident Commander assigned within 5 minutes for ≥ 95% of SEV-1/SEV-2.
- First status-page update within policy time for ≥ 95% of SEV-1/SEV-2.
- SLA credits down ≥ 80% versus the trailing $1.3M within 12 months.
- Paging volume ≤ 500 per month and noise ≤ 15% (from 3,400 and 85%).
- 100% of production services have a named owning team and a paging policy.
- Postmortems filed within 5 business days for 100% of SEV-1/SEV-2; action-item close rate ≥ 80% within 30 days.
- 24×7 IC and critical-path coverage with zero unfilled shifts per quarter.
- Paid on-call live for every rotation before that rotation pages humans.
- SOC 2 Type II incident-response controls evidenced for ≥ 5 months before the auditor's report.
- On-call pulse: ≥ 70% of engineers agree rotations are fair and limited to their services.
Steps (30):
1. Secure executive mandate and budget
Get a written CEO/CTO mandate that incident command is a company process, not a team hobby.
The mandate must state that **paid on-call** is required for production ownership. "No pager for other teams' code" is solved by named ownership, not by refusing coverage.
- Approve budget for tooling, stipends, training, and a dedicated program lead for six months.
- Name an executive sponsor (CTO or VP Engineering) who will chair the weekly incident review.
- Tie the clock to SOC 2 Type II: the process must be live in about 10 weeks so ~6 months of evidence remain.
- Commit that the CEO will hear about outages from this process, not from customers.
2. Form the working group and decision rights (depends on: 1)
Stand up a small group that can decide. Do not form a 28-team committee.
**Core seats:** SRE/platform lead, payments/ledger engineering manager, support lead, legal/compliance, HR, one rotating team EM, and a program manager.
- Meet twice a week for 10 weeks, then weekly.
- RACI: the group proposes; the sponsor decides in 48 hours; teams implement.
- Publish one Slack channel and one source-of-truth doc on day one.
- Time-box design to four weeks. Ship v1 rather than wait for consensus.
3. Inventory services, owners, and on-call gaps (depends on: 1)
Build a living catalog of all ~180 services: owning team, criticality, current on-call, alert sources, and runbook link.
Walk the last 31 customer-impacting incidents and the two events where **nobody was in charge**. Record who detected, who led, time to mitigate, and which alerts fired.
- Tag each service as critical-path, customer-visible, or internal.
- List the 16 teams with no on-call and every orphan service with no owner.
- Map all six alerting tools and the 3,400 monthly alerts onto services.
- Flag the shared PostgreSQL ledger and two-region failover as named special cases.
4. Map regulatory and contractual notification duties (depends on: 1)
Legal and compliance list every duty an incident can trigger. Do not invent clocks that violate a contract.
Cover SOC 2 CC7, customer MSA/SLA credit terms, money-transmitter rules, NYDFS 23 NYCRR 500 if applicable, PCI if in scope, and breach clocks.
- Extract **notification timings** from the largest customer contracts (status page, named AM, written notice).
- Define when legal, regulators, insurers, or the board must be told.
- Feed these clocks into severity triggers and the communications playbook.
5. Approve paid on-call and incident pay (depends on: 1, 4)
Unpaid on-call is why 16 teams refuse the pager and why nights are uncovered. Fix the money before asking for coverage.
HR, legal, and finance design a New York–compliant package: weekly stipend for primary and secondary, extra stipend for company IC and comms, and **after-hours incident pay** or comp time.
- Treat exempt vs non-exempt staff explicitly under NY wage-hour rules.
- Put stipend in the next pay cycle after policy publish, not "later."
- Cap consecutive night weeks. Fund hiring if a team cannot rotate fairly (minimum six people for 24x7 primary plus secondary).
- Publish the package before any new rotation starts. This is the main answer to pager pushback.
6. Ratify four severity levels and their triggers (depends on: 3, 4)
Adopt a business-impact scale. Engineers do not invent severity in the moment.
**SEV-1:** material payments failure; ledger down or inconsistent; security or customer-data incident; both regions impaired; or many customers already in SLA-credit territory.
**SEV-2:** degraded payments or a contracted feature down for multiple customers; SLA at risk.
**SEV-3:** narrow or single-customer impact with a workaround; no fleet-wide SLA risk.
**SEV-4:** no customer impact; ticket only.
- SEV-1 pages IC, comms, scribe, owning SMEs, and an exec; war room in 5 minutes; status page in 10; AM outreach in 20.
- SEV-2 pages IC and owning SMEs; comms may be the IC; status page in 15 minutes; updates every 30 minutes.
- SEV-3 pages the owning team only; customer notice only if that customer is affected.
- Anyone may declare. Only the IC may downgrade. When unsure, start high.
7. Define incident roles and their authority (depends on: 6)
Four roles. Separate coordination from debugging so "who is in charge" cannot stall for an hour again.
- **Incident Commander:** owns severity, the room, the clock, and the next action. Does not write code. May page anyone, freeze deploys, and invoke failover. Staffed from a company-wide trained pool, not from the failing team.
- **Communications Lead:** status page, customers, AMs, execs, regulators. Speaks only from IC-approved facts.
- **Scribe:** timeline in the incident tool. Required for SEV-1 and SEV-2.
- **SME responders:** the owning team's on-call. They mitigate. They do not run the room.
Publish a one-page authority card. The IC stays in charge even if a VP joins.
8. Design 24x7 coverage without 28 night rotations (depends on: 3, 5, 7)
Do not put 28 teams on 24x7. That is what engineers are rejecting.
Use **three layers** so people page for their own code, plus a trained commander.
- Layer A — company IC and SEV-1 comms: 24x7; about 24 trained people; week-long primary and secondary.
- Layer B — critical-path team on-call (ledger, payments processing, auth, API edge, platform/Kubernetes, data stores): 24x7 primary plus secondary.
- Layer C — all other teams: business-hours on-call; after hours the IC pages the team EM, who has a written escalation list.
Platform on-call is the safety net for unknown-owner pages, never the permanent owner. Every service must have a named team within 60 days or be scheduled to shut off.
9. Set rotation, handoff, and load rules (depends on: 8)
Write mechanical rules so rotations are fair and load is visible.
Primary week, then secondary week, then at least two weeks off. No one holds primary on two rotations at once.
- Handoff is a 30-minute overlap covering open incidents, silenced alerts, and upcoming changes.
- Page-load SLO: p50 ≤ 4 pages per 12-hour night shift; p95 ≤ 10. A breach opens an alert-quality action.
- Require a shadow week before a first IC shift or a first critical-path rotation.
- Swaps live in the paging tool. Managers own coverage gaps, not the last person on the roster.
10. Write detection and escalation paths (depends on: 6, 9)
Customers currently detect 40% of incidents and median time to detect is 22 minutes. That is the first failure mode.
Detection path: synthetic full-payment probes in both regions, SLO burn-rate alerts, support-to-incident intake, and one customer callback path that can create a SEV.
- A page must be acked in **5 minutes** or it auto-escalates to secondary, then IC, then the EM, then the VP.
- Support may declare SEV-2 or higher without engineering permission.
- If ownership is unclear for 10 minutes, the IC keeps the incident and assigns a temporary owner. Never wait.
- An exec bridge auto-opens for every SEV-1 at T+15 minutes.
11. Write internal, customer, and regulator communications (depends on: 4, 6, 7)
Stop "whoever is around" from writing the status page. Comms follow the clock, not convenience.
Timings from declaration:
- Internal war room: immediate. Exec summary for SEV-1/2 at 15 minutes, then every 30 minutes.
- **Public status page:** SEV-1 in 10 minutes, SEV-2 in 15. Updates at least every 30 minutes until resolve. Templates only. No speculation.
- Account managers get an affected-customer list and a script at T+20 minutes for SEV-1/2.
- Resolve notice and credit assessment within one business day.
- Comms pages legal on SEV-1 security, ledger integrity, or any outage that will breach contractual notice. Legal owns outbound regulatory letters; the IC owns facts.
12. Standardize blameless postmortems and action tracking (depends on: 6)
A written postmortem is mandatory for every SEV-1 and SEV-2 within 5 business days. SEV-3 if the IC or EM requests it.
Use one template: timeline, customer impact (volume, duration, credits), detection gap, what went well, what did not, process-focused five whys, and numbered actions with owner and due date.
- Review is **blameless** and scheduled. The IC attends. The exec sponsor reads every SEV-1.
- Actions live in one tracker, not in the doc. No action without an owner and a date. Default due date 14 days; 30 days max unless architecture work with a milestone.
- Close rate is a published metric. The old 11-of-64 pattern is a process failure.
13. Set alert quality rules that make paging acceptable (depends on: 3, 6)
3,400 alerts a month and 85% noise is why on-call feels like punishment. Pages are a product with a quality bar.
A page (not a ticket) must map to a customer-facing SLO or a hard dependency of one. It must have an owner team, a runbook link, and a default severity. It must be actionable at 3 a.m. by the person who is paged.
- Ban parallel paging from six tools. One paging policy: symptom-based; burn-rate preferred over raw thresholds.
- Every team gets a monthly noise budget. Exceeding it is a sprint task, not heroics.
- A human may silence a flapping alert only with a linked ticket.
14. Publish Incident Management Policy v1 (depends on: 5, 6, 7, 8, 9, 10, 11, 12, 13)
Collapse the design into a short policy people will open during an outage.
Ten pages or fewer, plus one-page cards for severity, roles, and comms timings. Host it where the incident tool can link it.
- Include the compensation summary and the rule: you are **not on-call for other teams' services**.
- Version it. v1 is mandatory from the pilot start date.
- Legal, HR, and the exec sponsor sign. Announce in all-hands, not only in Slack.
15. Implement a single incident command tool (depends on: 7, 10, 11)
Put one tool in the path that creates the room, pages roles from severity, records the timeline, and prompts status-page updates.
Requirements: Slack (or equivalent) incident bot, severity in one click, role assignment, stakeholder groups, and timeline export for postmortems and auditors.
- Integrate with the pager so IC, comms, and SME pages are automatic.
- Retain artifacts at least one year for SOC 2.
- Ad-hoc Zoom/Slack threads are no longer the system of record.
16. Consolidate six alerting tools onto one pager (depends on: 9, 13)
Pick one paging product. Connect existing monitors to it. Migrate **pages** first, tickets second.
- Inventory every page-producing rule. Delete or downgrade the noisy majority in S23.
- Route by service label → owning team schedule → the escalation policy from S10.
- IC and comms schedules live in the same product.
- Set a hard date after which pages outside the chosen tool are not valid on-call obligations.
17. Operationalize the status page and AM path (depends on: 11, 15)
Put the status page behind the Comms Lead role. Use templates for investigating, identified, mitigating, and resolved.
Subscribe AMs and customers whose contracts require it. Generate the affected-customer list from the incident (tenant, region, payment method).
- Dry-run a SEV-2 update before the pilot goes live.
- Record every public update in the incident timeline for the audit.
- Define partial vs full-outage wording so impact cannot be understated.
18. Stand up one postmortem repo and action board (depends on: 12)
Create the template, the filing location, and one Jira/Linear board with states: open, in progress, blocked, done, and won't-do (with exec reason).
Wire the incident tool so a SEV-1/2 automatically opens a draft postmortem and action tickets.
- Each week the program manager reports actions past due to the exec sponsor.
- If it is not in the board, it does not exist.
19. Assign owners and write critical-path runbooks (depends on: 3, 9)
Close the ownership gaps that cause "pager for other people's code."
Every production service gets a team in the catalog. Unowned services get an owner in 30 days or a decommission date.
- Write runbooks for ledger Postgres, regional failover, payments API, auth, and the Kubernetes control plane: symptoms, dashboards, mitigate vs escalate, customer impact.
- Link runbooks from alerts. If there is no runbook, the alert cannot page at night unless the EM accepts the gap in writing.
20. Detect payments failures before customers do (depends on: 3, 13)
Build or finish end-to-end synthetics: create payment, ledger write, webhook, both regions, both critical payment methods.
Alert on SLO burn, not on a single 500. Page SEV-2 or SEV-1 from these probes. This is the fastest lever on the 22-minute MTTD and the 40% customer-detected rate.
- Add ledger lag, replication, disk, and failover-readiness as first-class pages to the ledger team.
- Every postmortem asks first: **why did a customer see this first?**
21. Train the first cadre of ICs, comms leads, and scribes (depends on: 7, 14, 15)
Train about 24 ICs and 12 comms leads before the pilot. Classroom plus a recorded shadow of a simulated SEV-1.
Curriculum: severity, authority card, tool, comms timings, when to call legal, how to run a room of 20, how to hand off at 2 a.m.
- Certification: pass a tabletop. No certificate, no rotation.
- Recertify yearly and after any SEV-1 where process failed.
- Managers of ICs protect calendar time. This is part of the job.
22. Pilot on the payments critical path for six weeks (depends on: 5, 14, 15, 16, 17, 18, 19, 21)
Go live with policy, tool, paid rotations, and IC coverage for ledger, payments, API, platform, and support intake.
Keep old paths as backup for one week, then cut over. Real incidents use the new process only.
- Staff the program lead in every SEV-2+ as coach, not as secret IC.
- Collect friction daily. Fix tooling and wording in 48 hours.
- Expansion gate: named IC in under 5 minutes, first status update on time, no unpaid pages, postmortem filed.
23. Cut alert noise with a forced burn-down (depends on: 13, 16)
Give every team a numbered list of their noisiest alerts. Move noise from 85% to **under 15%**, and monthly pages from 3,400 toward 500.
Each sprint, critical-path teams must delete, debounce, or convert to ticket a fixed quota. Platform provides burn-rate and grouping libraries.
- Publish a weekly noise leaderboard. Shame systems, not people.
- After eight weeks, any page without a runbook or with >30% false pages in 14 days is auto-downgraded until fixed.
24. Roll remaining teams onto the model by risk (depends on: 22)
After the pilot gate, add teams in waves of four to five every two weeks. Highest customer-impact first.
Layer C teams get business-hours schedules and the EM night list. Do not surprise anyone with a pager.
- Each wave: ownership confirmed, alerts routed, runbooks for paging alerts, paid rotation in HR, one tabletop.
- Finish all 28 teams at least five months before the SOC 2 report date so the observation window covers the company.
- Orphan services still unowned at wave end are escalated to the sponsor for shutdown or reassignment.
25. Run tabletops and multi-region game days (depends on: 21, 22)
Schedule a monthly tabletop: SEV-1 ledger, SEV-1 region loss, SEV-2 degraded payments, customer-detected incident, and a "who is in charge" chaos drill.
Quarterly game day: fail a region or a ledger replica in staging or a controlled production drill.
- Include support, AMs, legal, and an exec. Process fails if only engineers show up.
- Capture actions on the same board as real postmortems.
- Use results as SOC 2 evidence that IR is tested.
26. Resolve ownership fights and pager culture (depends on: 5, 8, 22)
Treat pushback as design input, not defiance. Repeat the contract in office hours: you carry a pager for **your** services; IC is coordination; nights are paid; noise is a defect.
- EMs who cannot staff a fair rotation get headcount or have services reassigned. Do not run two-person 24x7.
- Publicly close the two historical "nobody in charge" incidents with what would be different now.
- Pulse-survey on-call at 60 and 120 days. If load or fairness is red, stop expansion until fixed.
27. Launch metrics, weekly review, and error budgets (depends on: 15, 22)
The CEO email exists because there was no operating rhythm. Stand up a dashboard and use it.
Track incident count by SEV, MTTD, MTTA, MTTM, customer-detected %, time-to-IC, time-to-status, pages per shift, noise %, SLA credits, postmortem SLA, and action close rate.
- **Weekly 30-minute review** chaired by the exec sponsor. Last week's incidents, open actions, alert SLO breaches. Mandatory for EMs of teams that had SEV-2+.
- Error budgets on critical SLOs: burn too fast and the team stops features and pays down reliability.
- Monthly report to the CEO so customer-first news does not recur.
28. Collect SOC 2 Type II evidence as a side effect (depends on: 4, 14, 22, 27)
Do not create a parallel "audit process." The real process is the evidence.
From pilot start, retain policy versions, on-call schedules, incident records, status-page history, postmortems, action tickets, training attendance, tabletop notes, and access reviews of the incident and pager tools.
- Map each artifact to SOC 2 CC7.2–CC7.5. Compliance owns the matrix; engineering owns completeness.
- Internal audit or a consultant runs readiness reviews at week 12 and week 20, not the week before the auditor.
29. Inspect and adapt the process after 90 days (depends on: 22, 27)
After 90 days live, revise v1 using data, not opinions.
Look at MTTD/MTTM, uncovered shifts, IC burnout, missed status updates, action close rate, and survey results.
- Change severity definitions if teams inflate or deflate them.
- Adjust Layer B vs C membership from real pages.
- Cut steps nobody follows. Add only what the last 90 days proved missing.
- Freeze v2 as the SOC 2 described process unless a defect is found.
30. Rehearse the audit and freeze the observation window (depends on: 24, 28, 29)
Six to eight weeks before the auditor, run a mock interview with an IC, a random engineer, support, and compliance.
Show one SEV-1 and one SEV-2 end-to-end: detect, page, command, comms, postmortem, closed actions.
- Fix evidence gaps immediately.
- Freeze process wording for the rest of the observation window; log exceptions.
- Brief the CEO and customer success with the metrics so the "outages we hear about from clients" story is retired.
Previous Proposal 5 (ID: 53260d57-1d36-433f-83b4-7205e06869b2, Agent: deepseek-v4-pro_initial_5, LLM: deepseek/deepseek-v4-pro):
Estimated Complexity: high
Success Metrics: - Customer-detected incidents decrease from 40% to less than 15% within six months.
- Median time to detect (MTTD) is under 5 minutes for SEV1 and SEV2 incidents.
- Median time to mitigate (MTTM) is under 60 minutes for SEV1 and under 2 hours for SEV2.
- Alert noise decreases from 85% to below 10% within three months.
- 100% of SEV1 and SEV2 incidents have a completed blameless postmortem within 5 business days.
- 100% of postmortem action items are tracked with an owner and due date; 90% are completed on time.
- 24x7 on-call coverage achieved across all 28 teams with no unpaid on-call.
- 99% of pages are acknowledged within 5 minutes.
- SLA credits paid reduce by at least 50% over the next 12 months.
- SOC 2 readiness: all incident response controls are documented, tested, and evidence is produced by month 7.
Steps (14):
1. Baseline current incident response and align stakeholders
Collect data from the last 12 months of incidents, all six alert tools, on-call practices, and team interviews. Identify gaps against the target incident management process and secure executive sponsorship.
- Gather incident timeline, detection source, mitigation time, customer impact, SLA credit, and postmortem status for all 31 incidents.
- Survey 28 teams on on-call burden, alert quality, and operational pain.
- Map current tools, escalation paths, and communication workflows.
- Create baseline metrics and a stakeholder map with executive sponsor and audit owner.
2. Define severity levels and response triggers (depends on: 1)
Define a four-level severity scale with objective business-impact criteria so any engineer can classify an incident consistently.
- SEV1: widespread transaction processing outage, data breach, security incident, or severe SLA breach; triggers full incident command, executive notification, and a 5-minute status page update.
- SEV2: major feature outage, significant degradation without workaround, or customer financial risk; triggers incident commander, full communications role, and status page updates.
- SEV3: partial impairment with workaround or limited customer impact; triggers on-call response, internal communication, and optional status page update.
- SEV4: minor or internal issue, no customer impact; handled during business hours through ticketing.
- Include an escalation matrix showing who can declare, downgrade, and invoke regulatory or legal involvement.
3. Define incident roles and decision authority (depends on: 1, 2)
Define incident roles, responsibilities, and decision authority using RACI to remove ambiguity about who is in charge.
- Incident Commander: owns the incident, declares severity, and coordinates resolution.
- Communications Lead: owns internal and external messaging, status page updates, and account manager notifications.
- Scribe: maintains timeline, incident log, and postmortem notes.
- Subject-matter responders: diagnose and fix the incident; may come from multiple teams.
- Executive sponsor: optional for SEV1; customer liaison: handles account managers.
- Define decision rights for severity declaration, escalation, rollback, customer communications, and incident closure.
4. Design 24x7 staffing model across 28 teams (depends on: 3)
Design 24x7 coverage across 28 teams without overloading engineers. Use service-based on-call plus a central incident command pool.
- Each service or domain team assigns primary and secondary on-call for its own services.
- Create central incident commander, communications, and scribe rotations staffed from a trained incident response guild across all teams; use follow-the-sun between the two AWS regions and time zones.
- Define escalation layers: service on-call to team lead or manager to service owner to executive.
- Define handoff times, shadow shifts, and load balancing; target at most one week of on-call per engineer per month.
- Bridge the current 12-team paid on-call to 28-team paid coverage; no team remains uncovered.
5. Define on-call rotations, compensation, and alert quality rules (depends on: 4)
Define sustainable rotations, pay, and rules that eliminate noisy pages.
- Rotations: weekly or biweekly, at least one primary and one secondary, with 12-hour shifts where possible or 24-hour for low-volume services.
- Compensation: monthly on-call stipend for all on-call engineers, additional incident response bonus for after-hours work, and time off in lieu; align with market rates.
- Alert quality rules: every page must be actionable, have a runbook link, specify a service owner, include severity, and be based on SLO burn or known failure signals; no dashboard-only alerts.
- Noise budget: reject or downgrade non-actionable alerts; all pages must go to on-call only after suppression and deduplication.
- Weekly alert review removes the top noisy alerts.
6. Design detection, escalation, and alert routing (depends on: 2, 4, 5)
Define how incidents are detected, routed, and escalated so nothing waits on a human to notice.
- Consolidate the six alert tools into one alerting and paging platform with routing by service, severity, and tags.
- Detection sources: infrastructure metrics, application synthetic transactions, log-based anomalies, business transaction SLI monitoring, and customer-reported issues through support or account managers.
- Routing: alert is paged to service on-call within 30 seconds; primary must acknowledge within 5 minutes; if no ack, page secondary then on-call manager.
- Escalation timeouts: unresolved SEV1 escalates to service owner at 15 minutes and to leadership at 30 minutes; any engineer can escalate to the incident commander.
- Define customer-reported incident intake and classification in the same tool.
7. Define internal and external communication protocols (depends on: 2, 3)
Define communication channels, templates, and timing for internal, customer, and regulator audiences.
- Internal: dedicated incident Slack channel, internal status page mirror, and war room bridge for SEV1; incident commander and communications lead own these channels.
- Status page: SEV1 post within 5 minutes, updates every 30 minutes or on material change, resolution within 60 minutes of mitigation; SEV2 post within 15 minutes, updates hourly; SEV3 optional.
- Account managers: SEV1 and SEV2 notify account managers within 15 minutes with an approved customer-facing description and expected impact.
- Regulators: legal or compliance determines notification for data breaches, security incidents, funds availability issues, or regulatory reportable events; criteria and timing follow legal and regulatory requirements; communications lead coordinates.
- Use pre-approved message templates and an approval chain; no ad-hoc wording.
8. Define postmortem policy and action tracking (depends on: 3)
Define mandatory blameless postmortems and action tracking.
- Mandatory for all SEV1 and SEV2 incidents, and any SEV3 that breaches SLA or is customer-detected.
- Format: impact, timeline, root causes, contributing factors, detection and response gaps, what worked well, and action items.
- Blameless: focus on system and process causes, not individual blame; use trained facilitators.
- Ownership: each action has an owner, due date, and tracking ID in a single backlog.
- Review postmortems at the weekly incident review; track action closure; expect 100% completion.
- Complete postmortems within 5 business days for SEV1 and SEV2 incidents.
9. Define metrics, dashboards, and review cadence (depends on: 2, 3, 8)
Define metrics and review cadence to measure process health.
- Metrics: MTTD, MTTM, customer detected percentage, alert noise percentage, on-call response time, on-call load, SLA credits paid, and postmortem action completion.
- Dashboards: real-time operational dashboard for on-call engineers and management.
- Weekly incident review: review all SEV1 and SEV2 incidents, action items, and noisy alerts.
- Monthly trends with leadership; quarterly review against SLOs and audit controls.
- Success thresholds: MTTD under 5 minutes, MTTM under 60 minutes for SEV1, customer detected under 15%, and alert noise under 10%.
10. Configure incident tooling and integrations (depends on: 5, 6, 7, 8, 9)
Implement and integrate the tools that automate the defined process.
- Aggregate alerts from the existing six tools into PagerDuty, Opsgenie, or a similar platform.
- Configure on-call schedules, escalation policies, and paging targeted at service owners.
- Integrate status page API for automated or one-click updates.
- Add Slack commands to declare incidents, start war rooms, assign roles, and post status updates.
- Integrate runbook and service catalog access; create postmortem templates in Jira or Notion with action item tracking.
- Ensure audit trails and role assignments are logged for SOC 2.
11. Pilot with 2-3 volunteer teams and iterate (depends on: 10)
Run a controlled pilot before full rollout to validate and refine the process.
- Select 2-3 volunteer teams with representative services and on-call patterns.
- Run the new severity, roles, on-call, alerting, and communication process for 2 weeks.
- Track metrics and gather feedback from on-call engineers, incident commanders, and communications leads.
- Iterate severity thresholds, alert rules, templates, and runbooks based on findings.
- Exit criteria: no SEV1 without a declared incident commander, alert noise below target, and positive on-call survey results.
12. Train and certify all 28 teams (depends on: 11)
Prepare all 28 teams so the process is not a paper exercise.
- Create role-specific training: all engineers, incident commanders, communications leads, scribes, subject-matter responders, and managers.
- Deliver interactive workshops and online modules.
- Run at least two game-day drills per quarter simulating SEV1 and SEV2 incidents, testing escalation, communication, and postmortem.
- Certify incident commanders and communications leads; only certified staff can hold those roles in rotations.
- Include scenario-specific training for account managers and legal or regulatory handling.
13. Full rollout and enforce process (depends on: 12)
Move every team onto the new process and retire old habits.
- Migrate all 28 teams into consolidated on-call schedules and alert routing.
- Retire old alerting tools or methods that bypass the process.
- Activate status page templates and postmortem action tracking globally.
- Monitor the first 30 days with a daily operational review and quick fixes.
- Communicate rollout status to leadership and all engineers; provide office hours and support.
14. SOC 2 readiness and continuous improvement (depends on: 13)
Map the process to SOC 2 and keep improving after go-live.
- Map each process component, including severity, roles, detection, communications, postmortems, and metrics, to SOC 2 control requirements.
- Produce runbooks, evidence of training, incident logs, postmortems, and action item tracking.
- Conduct an internal readiness assessment at month 6 and remediate gaps before the audit at month 8.
- Maintain weekly and monthly metrics reviews to drive continuous improvement.
- Refresh on-call compensation and alert quality rules based on operational data.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Executive mandate, single owner, funding and non-negotiables", "description": "Convert the CEO email into a chartered program with one accountable owner and authority over all 28 teams. Incident response becomes a **company operating process**, not a per-team choice.\n\n- Name an executive sponsor (CTO) and one accountable owner (Head of Reliability / Incident Management) with a small permanent office: 1 program lead, 1 platform engineer, 1 analyst.\n- Form a decision group with Engineering, SRE/Platform, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours. It is not a 28-person committee.\n- Fix the non-negotiables now: one severity scale, one paging tool, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.\n- Set the clock deliberately earlier than the audit: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready week 28.\n- Approve budget against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises.", "dependencies": []}, {"step_id": "S2", "title": "Seven-day interim command bridge", "description": "Do not let design work leave the company unprotected for six weeks. Put a crude but real process in place within seven days and improve it later.\n\n- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.\n- Stand up an interim Duty Incident Commander roster from engineering managers and senior SREs, primary plus backup, 24x7. Pay it retroactively under the final policy.\n- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.\n- Tell Support to escalate credible customer reports immediately, without waiting for engineering confirmation.\n- Triage the 53 open historical action items: complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.\n- Hold a 15-minute daily operations review until the permanent process is live.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Forensic baseline of incidents, alerts and money lost", "description": "Rebuild the facts before designing anything. This becomes both the design input and the \"before\" picture for the executive and the auditor.\n\n- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start / detect / declare / commander-assigned / mitigate / resolve, who led, credits paid, contributing factors.\n- For each of the 40% customer-first detections, name the specific signal that was missing. This list drives the detection backlog.\n- Reconstruct the two \"nobody in charge\" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.\n- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome. Identify the top 50 rules producing most of the noise and every rule with no owner or runbook.\n- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Listening tour, resistance map and the on-call deal", "description": "Engineer pushback is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.\n\n- Interview all 28 teams plus Support, CS and Sales in two weeks. Test what the objection actually is: unpaid work, lost sleep, unfamiliar code, missing runbooks, or fear of blame. Each has a different remedy.\n- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.\n- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.\n- Recruit 12–15 credible engineers as a design working group so the process is co-authored.\n- Run a baseline sentiment survey (fairness, trust in alerts, willingness, burnout) to re-measure at 60, 120 and 365 days.", "dependencies": ["S1"]}, {"step_id": "S5", "title": "Service catalog, ownership and money-path tiering", "description": "You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that is the single source of truth for paging, impact and audit.\n\n- One accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list.\n- Tier by business impact, not by technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).\n- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions and third parties, including sponsor banks and processors.\n- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that make nominal redundancy fake.\n- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.\n- Every orphan service gets an owner in 30 days or a decommission date approved by the sponsor. Tier 0 without an owner is an executive escalation.", "dependencies": ["S3"]}, {"step_id": "S6", "title": "Severity scale, declaration rules and incident lifecycle", "description": "Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.\n\n- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.\n- **SEV2 (critical):** material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.\n- **SEV3 (major, contained):** narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.\n- **SEV4:** no customer impact; ticket only, never pages.\n- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.\n- Anyone may declare and nobody is penalised for over-declaring; only the commander may downgrade, with the evidence recorded.\n- Lifecycle: Detected → Declared → Triaged → **Mitigated** (customer impact ends) → Monitoring → **Resolved** (backlog processed and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.\n- Publish a decision tree with 12 worked examples taken from the real 31 incidents.", "dependencies": ["S3", "S5"]}, {"step_id": "S7", "title": "Incident roles, authority and handover discipline", "description": "Solve \"nobody in charge for an hour\" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.\n\n- **Incident Commander:** owns severity, priorities, roles, cadence and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit spend. Retains command when a VP joins.\n- **Communications Lead:** single voice for status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.\n- **Scribe:** timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates for SEV1/SEV2.\n- **Subject-Matter Responders:** engineers of the owning team; they mitigate, they do not run the room.\n- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.\n- Security, Legal, Compliance and Vendor Management join on defined triggers.\n- Rules: command claimed within 5 minutes, stated in channel (\"I am IC\"), distinct people for command, comms and technical lead at SEV1/SEV2, and every handover announced verbally and in writing with the exact time.\n- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation or privileged-access rules.", "dependencies": ["S6"]}, {"step_id": "S8", "title": "24x7 coverage model: central command corps, local expertise", "description": "Do not create 28 night rotations. Centralise coordination in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.\n\n- **Layer A — Incident Command corps:** ~30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of ~18 from Support, CS and engineering management; a scribe pool used as the training entry point.\n- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.\n- **Layer C — everyone else:** business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.\n- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.\n- Evaluate in writing, with cost and timeline: a US-only paid night rotation now, a Lisbon or APAC follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.\n- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.", "dependencies": ["S5", "S7"]}, {"step_id": "S9", "title": "On-call compensation, labour compliance and fatigue safeguards", "description": "Unpaid on-call in New York is both a retention problem and a legal exposure. Pay for it before asking anyone to sign up, and publish the numbers.\n\n- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 for business-hours rotations, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.\n- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover rather than expecting normal output.\n- HR, Finance and employment counsel publish amounts, eligibility, tax treatment and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.\n- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.\n- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.\n- Non-cash elements: on-call counted as delivery load with a ~15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.", "dependencies": ["S8", "S4"]}, {"step_id": "S10", "title": "Alert quality standard, page budget and noise burn-down", "description": "3,400 alerts at 85% noise is the reason detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human, and burn the backlog down deliberately rather than by mass silencing.\n\n- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.\n- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.\n- Run new alerts in shadow mode for seven days; test both firing and recovery.\n- Set a **page budget** of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.\n- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements; return it to its owner with a correction deadline.\n- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.\n- Target 3,400 → under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.", "dependencies": ["S3", "S5"]}, {"step_id": "S11", "title": "Detection uplift on the money path", "description": "The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.\n\n- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.\n- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.\n- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.\n- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls; five similar support tickets in ten minutes auto-creates a triage incident.\n- Route credible partner and processor notifications into the same declaration path within five minutes.\n- Record detection source on every incident. \"Customer detected first\" becomes a named defect class with a mandatory tracked action.", "dependencies": ["S10", "S5"]}, {"step_id": "S12", "title": "Consolidate to one pager, one incident record, one status page", "description": "Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrating without creating a monitoring gap.\n\n- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.\n- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation in the new platform.\n- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline and starts the clock.\n- Route every page through the service catalog: service label → owning domain schedule → escalation policy.\n- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.\n- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack or the identity provider is unavailable. Test paging, escalation, status publication and bridge access weekly.\n- Set a hard date after which pages outside this tool create no on-call obligation.", "dependencies": ["S6", "S8", "S10"]}, {"step_id": "S13", "title": "Escalation ladder and the five-minute command rule", "description": "Write one unskippable path from \"something looks wrong\" to \"someone is in charge\", and make the default action never be waiting.\n\n- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.\n- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.\n- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.\n- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.\n- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.\n- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification, so the commander never waits for an executive.\n- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.", "dependencies": ["S7", "S8", "S12"]}, {"step_id": "S14", "title": "Live incident execution doctrine", "description": "Give responders one short operating procedure for the first minutes through closure. Priority is limiting customer and financial harm, not proving a root cause.\n\n- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.\n- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.\n- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.\n- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.\n- Payments-specific guards: protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; controlled backlog drain; **reconciliation completed before any payment or ledger incident is declared resolved**.\n- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.\n- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and Customer Success, and reopening if impact recurs.", "dependencies": ["S7", "S12", "S13"]}, {"step_id": "S15", "title": "Internal communications protocol", "description": "Standardise the internal picture so executives, Support and Sales are informed without interrupting the person running the incident.\n\n- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.\n- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.\n- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.\n- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.\n- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.\n- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.", "dependencies": ["S7", "S12"]}, {"step_id": "S16", "title": "Customer communications and status page policy", "description": "Replace \"whoever is around\" with a timed, owned, pre-approved process. The Comms Lead is the single author and never writes from scratch under pressure.\n\n- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.\n- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.\n- Component-level status page mapped to customer journeys, with subscriptions, email, webhook and RSS; all 2,100 customers subscribed by default.\n- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack; the long tail gets the status page and a proactive email.\n- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity or blame; state affected regions and workarounds.\n- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.", "dependencies": ["S6", "S15"]}, {"step_id": "S17", "title": "Regulatory, partner and legal notification playbook", "description": "In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.\n\n- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.\n- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded **even when the answer is \"not reportable\"**, with evidence, decision-maker and timestamp.\n- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.\n- Pre-draft notification letters held under privilege review; Legal owns outbound regulatory text, the commander owns the facts.\n- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.\n- Rehearse the playbook once a quarter as part of the exercise programme.", "dependencies": ["S6", "S16"]}, {"step_id": "S18", "title": "SLA credit and financial impact workflow", "description": "Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops being surprised.\n\n- Agree with Legal and Finance the availability measurement method per contract and per component.\n- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.\n- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.\n- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.\n- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.\n- Target: credits from $1.3M to under $400k in year one, and use the delta as the standing business case for on-call pay and reliability capacity.", "dependencies": ["S6", "S16"]}, {"step_id": "S19", "title": "Blameless postmortem standard and Incident Review Board", "description": "Replace \"some incidents, various formats\" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.\n\n- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.\n- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.\n- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.\n- Two mandatory questions in every review: **why did a customer see this first?** and **why did mitigation take as long as it did?**\n- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.\n- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.\n- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly \"top five recurring causes\" analysis, with restricted versions for security or privileged content.", "dependencies": ["S6", "S7"]}, {"step_id": "S20", "title": "Action ownership, capacity reservation and enforcement", "description": "Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.\n\n- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.\n- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.\n- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.\n- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.\n- Verify effectiveness before closing. A closed ticket without evidence does not close the action.\n- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.", "dependencies": ["S19", "S12"]}, {"step_id": "S21", "title": "Runbooks, readiness bar and ledger blast-radius reduction", "description": "Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.\n\n- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.\n- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.\n- Treat the **shared ledger cluster as the single largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel workstream on blast-radius reduction — tenant or function partitioning, read replicas, and isolation of non-critical readers.\n- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.\n- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.", "dependencies": ["S5", "S8"]}, {"step_id": "S22", "title": "Training and certification academy", "description": "Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.\n\n- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.\n- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.\n- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.\n- **Incident Commander (2 days + shadowing):** command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.\n- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.\n- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence for their own domain before independent primary duty; two shadow shifts minimum, never primary in the first 90 days.\n- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.", "dependencies": ["S7", "S13", "S15", "S19"]}, {"step_id": "S23", "title": "Exercise programme: tabletops, game days and unannounced drills", "description": "The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.\n\n- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.\n- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.\n- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.\n- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.\n- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.\n- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.\n- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update and time-to-correct-mitigation.", "dependencies": ["S22", "S12", "S21"]}, {"step_id": "S24", "title": "Publish Incident Management Policy v1", "description": "Collapse the design into a document people will actually open mid-outage, and make it official.\n\n- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.\n- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.\n- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.\n- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in Slack.\n- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.", "dependencies": ["S6", "S7", "S8", "S9", "S10", "S13", "S15", "S16", "S17", "S19", "S20"]}, {"step_id": "S25", "title": "Pilot on the payments critical path", "description": "Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.\n\n- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake — including two of the 12 teams already on-call.\n- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.\n- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.\n- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.\n- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold fixes into the standard before rollout.\n- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment. Publish a one-page result to the whole company.", "dependencies": ["S24", "S11", "S12", "S21", "S22", "S9"]}, {"step_id": "S26", "title": "Metrics, dashboards, review cadence and anti-gaming", "description": "Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.\n\n- **Response:** time to detect (from first impact), time to declare, time to commander, acknowledgement, mitigate, resolve — split by severity, tier, journey, region and detection source.\n- **Quality:** customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.\n- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.\n- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.\n- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.\n- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.\n- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.\n- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; team scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.", "dependencies": ["S12", "S19", "S25"]}, {"step_id": "S27", "title": "Wave rollout to all 28 teams with readiness gates", "description": "Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.\n\n- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.\n- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.\n- Each wave gets a named coach from the program office for three weeks.\n- Gates are signed by the director; failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.\n- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.\n- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.", "dependencies": ["S25", "S26"]}, {"step_id": "S28", "title": "Alert noise burn-down campaign", "description": "Run the noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.\n\n- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.\n- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.\n- Publish a weekly noise leaderboard that names systems, never people.\n- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.\n- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness; review missed detections monthly with the same seriousness as noise.\n- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.", "dependencies": ["S10", "S12", "S25"]}, {"step_id": "S29", "title": "Change management, fairness and pager culture", "description": "Run this from day one in parallel. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.\n\n- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.\n- Publicly close the two historical \"nobody in charge\" incidents with a written account of what would be different now.\n- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.\n- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.\n- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.\n- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.", "dependencies": ["S4", "S9", "S25"]}, {"step_id": "S30", "title": "SOC 2 evidence by design, internal testing and mock audit", "description": "Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.\n\n- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, and A1.2 for availability — and confirm the observation window early.\n- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including \"not reportable\", postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.\n- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.\n- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support and Compliance.\n- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.\n- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.", "dependencies": ["S24", "S26", "S27"]}, {"step_id": "S31", "title": "Risk register and contingencies", "description": "Name the ways this programme fails and pre-commit the response. Review it monthly with the sponsor.\n\n- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.\n- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.\n- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.\n- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.\n- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.\n- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.\n- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.", "dependencies": ["S1"]}, {"step_id": "S32", "title": "Ninety-day inspect-and-adapt, then year-two sustainability", "description": "Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.\n\n- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.\n- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.\n- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.\n- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.\n- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.\n- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.", "dependencies": ["S26", "S27", "S30"]}], "estimated_complexity": "high", "success_metrics": "- A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.\n- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.\n- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.\n- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.\n- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.\n- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.\n- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.\n- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.\n- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.\n- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.\n- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.\n- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.\n- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.\n- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.\n- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.\n- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.\n- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.\n- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of \"not reportable\".\n- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.\n- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.\n- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders."}Grew from 19 to 24 steps by adding the pieces it was missing — readiness/runbook standard (S9), detection uplift (S12), pilot (S20), waves (S21), exercises (S22) — while keeping its distinctive strengths: compliance designed on day one (S4, before severity in S5), a balanced anti-gaming scorecard (S18), and now numeric severity guardrails.
- S5 is the round's most operational severity definition: >10% payment failures for 5 minutes = SEV1, 1–10% = SEV2, with the explicit caveat that thresholds are guardrails, not permission to under-classify integrity or settlement risk.
- S4 sequences evidence and retention design before the operating process is finalised, so the audit trail is not retrofitted.
- S7 schedules an Engineering Director as 24x7 Executive Duty Officer — the only plan to staff that role on a roster.
- S9 readiness gate links every paging alert to the exact runbook step expected of the responder.
- S17 adds the escalation ladder (manager +7, director +14, CTO +30) and effectiveness verification the R0 postmortem step lacked.
- Most honest paging target in the round: 3,400 → 1,500 by day 90 → 700 by month 6, rather than a uniform <500.
- Dropped the risk register from R0 S1; nothing in R1 names failure modes or pre-commits contingencies.
- Dropped R0 S15's explicit prohibition on contractors or a managed service as sole incident commander or ledger owner — a useful guardrail.
- Still no dollar amounts or budget envelope for compensation (S8 only defines structure and a 14-day publication deadline), making the CFO ask unquantified.
- 28 success metrics with no headline set; the executive view is buried.
- No short field-usable policy card equivalent to P1 S24 / P4 R0 S14 — policies exist only as versioned documents in S4.
- Proposal 1 : Service readiness bar and major-incident playbooks as a precondition for paging.
- Proposal 1 : Action escalation ladder with capacity reservation and residual-risk acceptance.
- Proposal 1 : Automated SLA credit computation from the incident record, proactive versus claims-based posture.
- Proposal 1 : Critical-path pilot with numeric exit gates and gated waves by criticality.
- Proposal 4 : Publish the compensation package before any new rotation starts.
- Proposal 1 : Time-based auto-escalation (SEV3 open >2h becomes SEV2).
- Proposal 1 : Publishing indicative stipend dollar amounts in the plan.
- Proposal 3 : Automated email/SMS notification to all 2,100 customers for SEV1 and SEV2.
+ Install an interim process in seven days+ Build the baseline, ownership catalog, and risk map+ Design compliance and evidence controls from day one+ Define roles, authority, and handoffs+ Create sustainable 24×7 coverage across the service estate+ Set the service readiness and runbook standard+ Establish one paging and incident system of record+ Enforce alert quality and burn down noise safely+ Detect payment and ledger failures before customers+ Operationalize regulatory, partner, contract, and credit decisions+ Enforce action ownership and effectiveness tracking+ Measure outcomes, controls, business impact, and human load+ Train participants and address pager resistance+ Pilot on the payment critical path+ Roll out by customer journey and risk+ Exercise command, regional resilience, and ledger recovery+ Test SOC 2 operating effectiveness before the auditorInstall immediate minimum controlsCreate the service and dependency catalogDefine roles and sustainable 24x7 staffingConsolidate incident and paging toolingImprove detection and enforce alert qualityCodify acknowledgement and escalation pathsMeasure performance and review it routinelyTrain and certify participantsPilot and expand the on-call modelExercise regional, ledger, and communication failuresExecute a time-boxed enterprise rolloutBuild SOC 2 evidence as the process operates
The plan produced
1. Create the mandate, ownership, and funding
Launch incident management as a company operating program within 48 hours. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable program owner.
- Form a small steering group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Fund two to three implementation staff, paging and incident tooling, observability work, training, exercises, and on-call compensation.
- Reserve 10%–15% of engineering capacity for alert remediation, runbooks, and incident actions.
- Authorize incident commanders to freeze deployments, order rollback, disable features, shift traffic, and invoke continuity plans.
- Preserve financial controls. Incident commanders may coordinate ledger recovery but may not bypass dual approval, privileged-access controls, or reconciliation.
- Set the delivery target: critical controls operational within 60 days, enterprise rollout within 120 days, and a mock audit in month 7.
2. Install an interim process in seven days (after 1) from P4 step 29
Do not wait for new tools or the final policy. Put a minimum viable incident process into operation immediately and start collecting evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communications schedule.
- Establish one monitored declaration path through chat, telephone, and the existing paging tools.
- Create a standard incident channel, bridge, incident document, and naming convention.
- Staff temporary primary and backup incident commanders 24×7 from the existing on-call teams and engineering leadership.
- Compensate interim duties retroactively under the final compensation policy.
- Require a named incident commander within 10 minutes for every suspected major incident.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Hold a daily 15-minute control review until the permanent process is live.
3. Build the baseline, ownership catalog, and risk map (after 1) from P1 step 4
Establish the facts behind the current failures and assign every production component an owner. Use the resulting catalog as the source for routing, escalation, and audit evidence.
- Reconstruct all 31 customer-impacting incidents, including first impact, detection source, declaration, command assignment, mitigation, resolution, customer communications, and credits.
- Analyze the two incidents with no clear leader and every case detected first by customers.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, databases, queues, regional dependencies, payment processors, and banking partners.
- Assign one accountable team, engineering manager, product owner, primary escalation, secondary escalation, dashboard, and runbook to each service.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 based on financial integrity, customer impact, contractual exposure, and dependency centrality.
- Map critical customer journeys to their application, PostgreSQL, Kubernetes, regional, and third-party dependencies.
- Inventory the six alert sources and all 3,400 monthly alerts by owner, volume, actionability, and duplication.
- Interview representatives from all 28 teams and baseline on-call sentiment, fatigue, and objections.
- Give orphan services an owner or decommissioning decision within 30 days.
4. Design compliance and evidence controls from day one (after 1) new
Map the operating process to audit, legal, contractual, and record-retention requirements before finalizing it. Confirm the expected SOC 2 observation period with the auditor immediately.
- Map controls to applicable SOC 2 criteria for monitoring, incident identification, response, recovery, communications, corrective action, access, and availability.
- Define evidence required for declarations, pages, acknowledgements, role assignments, decisions, status updates, postmortems, actions, training, drills, and exceptions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve the Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, and Postmortem Standard.
- Record control exceptions with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing evidence before the audit.
5. Adopt one severity and lifecycle standard (after 2, 3, 4)
Use four impact-based severity levels across operational, security, data, and third-party incidents. Start at the highest credible severity when facts are uncertain, then downgrade with recorded evidence.
- SEV0 — financial or security crisis: suspected ledger corruption, unauthorized or duplicated funds movement, material data compromise, both-region loss, or a decision to suspend payment processing. Page all roles immediately; engage executives, Security, Legal, Compliance, and Risk; assess regulatory duties within one hour; use distinct role holders; require a postmortem.
- SEV1 — critical availability event: a core payment journey is unavailable, payment failures exceed an initial 10% guardrail for five minutes, regional loss has impaired failover, a settlement deadline is at imminent risk, or rapid error-budget burn makes a material SLA breach likely. Assign all roles; notify internal stakeholders within 10 minutes; publish customer status within 15 minutes; update every 30 minutes; require a postmortem.
- SEV2 — major bounded event: approximately 1%–10% of payment attempts fail, a material customer subset or critical customer is down, degradation has a workaround, or contractual impact is likely. Assign an incident commander and responders; add communications and scribe roles for customer impact; publish status within 30 minutes; update every 60 minutes; require a postmortem for customer-visible events.
- SEV3 — limited event: localized impact, a safe workaround, and no financial-integrity, security, regulatory, or material contractual risk. The owning team leads; page only if immediate action is necessary; use a ticket otherwise.
- Treat the percentage thresholds as declaration guardrails, not reasons to under-classify integrity, settlement, security, or strategic-customer risk.
- Permit any employee to declare an incident. Only the incident commander may lower severity, with the rationale logged.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of active customer harm. Declare resolution only after stability, backlog recovery, and required ledger reconciliation.
6. Define roles, authority, and handoffs (after 5) from P1 step 6
Separate command, communication, recordkeeping, and technical repair. One named person must hold command at every moment of a major incident.
- Incident Commander: owns severity, priorities, role assignment, escalation, decision cadence, mitigation coordination, and closure. The commander does not act as the primary technical operator.
- Communications Lead: owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language. Legal or Compliance retains ownership of regulatory submissions.
- Scribe: maintains the timestamped record of facts, hypotheses, decisions, commands, owners, and status changes.
- Subject-Matter Responders: diagnose and mitigate services for which they have accepted ownership, access, training, and runbooks.
- Executive Duty Officer: removes organizational obstacles and approves exceptional business decisions without displacing the incident commander.
- Require separate commander, communications, scribe, and primary technical lead for SEV0 and SEV1.
- For customer-visible SEV2, keep the commander separate from the primary technical responder; communications and scribe may be combined if workload permits.
- Announce every role assignment and transfer in the incident channel. Require verbal and written handoff with current impact, decisions, risks, and next actions.
- Keep executives and account managers out of the technical command path; questions flow through the Executive Duty Officer or Communications Lead.
7. Create sustainable 24×7 coverage across the service estate (after 3, 6) new
Use central command coverage and service-domain responder coverage rather than creating 28 fragile night rotations. Engineers remain responsible only for services they own or have formally trained to support.
- Build a company incident-command pool of approximately 18–24 certified senior engineers and managers, with a primary and backup scheduled at all times.
- Build 12–16-person communications and scribe pools using Support, Customer Operations, Engineering Operations, and qualified engineering managers.
- Schedule an Engineering Director or equivalent as the 24×7 Executive Duty Officer.
- Group related service owners into approximately 8–12 coherent responder domains only where members have training, access, and explicit acceptance.
- Require Tier 0 and Tier 1 domains to provide 24×7 primary and secondary responders, normally with at least six trained people in each sustainable rotation.
- Give Tier 2 services business-hours coverage plus a maintained manager escalation path. Treat Tier 3 conditions as tickets unless their impact changes.
- Maintain distinct but coordinated coverage for the ledger application and the PostgreSQL platform.
- Route unknown-owner events to the duty commander and platform triage temporarily. Treat every such event as an ownership control defect.
- If a small team cannot staff a fair rotation, merge coverage only after training or provide headcount, service reassignment, or decommissioning.
8. Approve compensation and fatigue protections (after 7)
End unpaid on-call before expanding mandatory coverage. Publish the policy through HR, Legal, Finance, and Payroll within 14 days.
- Pay fixed weekly stipends for primary and secondary service rotations.
- Pay separate stipends for duty commander, communications, and scribe assignments.
- Provide additional call-out compensation or equivalent paid recovery time for material after-hours work.
- Apply overtime and reporting rules correctly for non-exempt employees under federal and New York requirements.
- Pay higher rates for company holidays and provide a protected recovery day after qualifying overnight work, SEV0 events, or prolonged SEV1 response.
- Target no more than one primary week in six and prohibit simultaneous primary assignments.
- Avoid consecutive primary weeks and make all swaps visible in the paging system.
- Reduce sprint commitments for people carrying primary duty rather than expecting normal delivery capacity.
- Provide a documented accommodation path for health, disability, or caregiving constraints without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
9. Set the service readiness and runbook standard (after 3, 5, 7) new
A team cannot respond effectively at night without ownership, access, telemetry, and rehearsed recovery procedures. Apply a formal readiness gate to every Tier 0 and Tier 1 service.
- Require a current architecture diagram, dependency map, dashboards, SLOs, runbooks, rollback method, feature-control method, contacts, and tested escalation path.
- Record RTO, RPO, data-integrity requirements, regional mode, and customer-facing capabilities in the service catalog.
- Link every paging alert to the exact runbook step expected from the responder.
- Prevent new paging alerts for services that fail readiness review. Preserve existing critical detection through a documented exception until a safe replacement exists.
- Write priority playbooks for PostgreSQL failure, ledger-integrity investigation, regional failover, Kubernetes control-plane degradation, payment-processor failure, queue backlog, credential compromise, and payment suspension.
- For the ledger, document read-only or stop-processing modes, failover controls, replay and duplicate protections, backlog handling, and post-recovery reconciliation.
- Require runbook review after material incidents and at least twice per year through exercises.
10. Establish one paging and incident system of record (after 4, 5, 6, 7) new
Consolidate paging and incident coordination without requiring an unsafe big-bang replacement of every monitoring system. Monitoring sources may remain specialized, but all human pages must enter one controlled platform.
- Select one enterprise paging platform and one integrated incident record, with chat, telephone, SMS, conference, status-page, ticketing, and service-catalog integrations.
- Initially ingest events from all six current tools, then deduplicate, correlate, and route by service ownership.
- Automate incident channel and bridge creation, role paging, timeline capture, severity changes, communications reminders, and postmortem creation.
- Preserve immutable records of declarations, acknowledgements, role assignments, decisions, and messages.
- Use role-based access, multifactor authentication, break-glass controls, and periodic access reviews.
- Provide telephone and offline fallback procedures for loss of chat, identity, the paging vendor, or an AWS region.
- Test paging and fallback paths weekly.
- Retire a legacy paging route only after its signals have owners, quality review, successful end-to-end tests, and at least two weeks of verified operation in the new path.
11. Enforce alert quality and burn down noise safely (after 3, 10) new
Treat paging alerts as production products with owners and quality requirements. Do not reduce noise by silently disabling detection.
- Require every page to identify the service, owner, customer or SLO risk, urgency, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Define an actionable page as one requiring prompt human judgment or intervention that materially reduces customer, financial, security, or contractual risk.
- Route informational, capacity-planning, and non-urgent conditions to dashboards or ticket queues.
- Prefer symptom and error-budget burn alerts over raw CPU, memory, pod, or log-volume thresholds.
- Run new alerts in shadow mode for at least seven days unless an emergency exception is approved.
- Review alerts with less than 50% actionability, more than three firings in seven days, or repeated no-action acknowledgements within two business days.
- Require compensating detection and an owner before suppressing or removing an alert.
- Set a responder load target of no more than two after-hours pages per person per week. A breach creates a mandatory alert-remediation plan.
- Review alert actionability, duplication, missed detection, and page load monthly by domain.
- Prioritize the small number of rules producing most of the current 85% noise.
12. Detect payment and ledger failures before customers (after 3, 11) from P4 step 20
Shift detection from infrastructure symptoms to customer journeys and financial outcomes. Set internal objectives stricter than the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, webhooks, and reporting freshness.
- Run external synthetic transactions from outside the production boundary and through both regions at least every minute for critical paths.
- Monitor payment failure rates, latency, queue age, delayed value, settlement-window risk, regional asymmetry, and third-party response quality.
- Add continuous ledger controls for reconciliation breaks, unexpected balances, duplicate identifiers, replication lag, backup health, and failover readiness.
- Add tenant or cohort anomaly detection for high-value customers and common payment methods.
- Convert high-priority Support, account-manager, bank, and processor reports into incident candidates within five minutes.
- Review every customer-first incident as a missed-detection defect and create a corrective action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Codify escalation and live incident execution (after 6, 10)
Create one time-bound path from signal to ownership and mitigation. Delivery of a notification does not count as acknowledgement.
- Page the owning Tier 0 or Tier 1 primary immediately; page the secondary after five unacknowledged minutes; page the domain manager at 10 minutes; escalate to the engineering director at 15 minutes.
- For SEV0–SEV2, page the duty commander immediately. Page the backup after five minutes and require the Engineering Director to assume command if no certified commander owns the event by 10 minutes.
- Automatically involve command for integrity concerns, security concerns, regional events, cross-team impact, critical customer journeys, or unresolved ownership.
- If impact remains materially unknown after 15 minutes, increase response posture rather than waiting for certainty.
- Open one incident channel, bridge, and system record. State severity, known impact, assigned roles, current objective, and next update time.
- Freeze unrelated changes during SEV0 and SEV1 unless the commander records an exception.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require controlled backlog processing and reconciliation before resolving payment or ledger incidents.
- Require a formal command handoff for incidents extending beyond four hours or when fatigue impairs a role holder.
14. Standardize internal and customer communications (after 5, 6, 10)
Communicate known impact early without waiting for root cause. The Communications Lead uses approved facts and always states the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish customer status within 15 minutes when customer-visible and legally safe, then update at least every 30 minutes.
- For SEV1, brief internal stakeholders and publish status within 15 minutes, then update every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send notice within 30 minutes, then update every 60 minutes.
- For SEV3, communicate directly to affected customers only when impact or contract terms require it.
- Give account managers an approved statement and affected-customer list within 30 minutes for SEV0 or SEV1 and within 60 minutes for SEV2.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. Describe affected capabilities, symptoms, workarounds, and the next update.
- Do not speculate about root cause, blame, data integrity, security scope, or recovery time.
- Post a Monitoring update within 30 minutes of mitigation. Post Resolved only after stability and required reconciliation.
- Provide a customer-facing incident summary within five business days for qualifying incidents.
- Record and approve any delay or restriction of public details during an active security threat.
15. Operationalize regulatory, partner, contract, and credit decisions (after 4, 5, 14) new
Give Legal and Compliance a timed decision process while keeping technical command with the incident commander. Record both reportable and non-reportable determinations.
- Build a jurisdiction and obligation matrix covering applicable NYDFS requirements, state breach laws, GLBA or FTC requirements, PCI obligations, money-transmitter rules, sponsor banks, payment networks, cyber insurance, and customer contracts.
- Validate applicability and deadlines with counsel rather than assuming every operational incident is reportable.
- Begin reportability assessment immediately for every SEV0, security event, suspected ledger-integrity event, and relevant SEV1.
- Target a documented initial legal assessment within one hour for SEV0 and within two hours for other potentially reportable events.
- Record the decision, evidence, approver, legal deadline, submission owner, and confirmation of delivery.
- Maintain tested 24×7 contacts for regulators, banks, networks, insurers, outside counsel, and critical vendors.
- Encode customer-specific notification clocks and channels in the customer record used by the Communications Lead.
- Have Finance calculate affected minutes, likely credits, and contractual exposure from the incident record within five business days.
- Review whether proactive credits or claims-based handling applies by customer segment and contract.
16. Make postmortems mandatory, consistent, and blameless (after 5, 6, 10)
Use one review standard to learn from incidents and test whether controls worked. Keep learning reviews separate from performance or misconduct processes.
- Require a postmortem for every SEV0 and SEV1.
- Require one for customer-visible SEV2, customer-first detection, incidents lasting more than two hours, contractual breaches, repeat failures, control gaps, and ledger-integrity near misses.
- Produce the factual draft within three business days, conduct the review within five, and publish the approved version within 10.
- Use one template covering summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, lessons, and actions.
- Analyze why detection was not earlier and why mitigation took as long as it did.
- Examine technical, organizational, process, testing, dependency, and incentive factors rather than forcing a single root cause.
- Use a trained facilitator and written blameless rules. Describe decisions in the context and information available at the time.
- Publish broadly useful findings internally while maintaining restricted versions for security, privacy, personnel, or privileged material.
- Review postmortem quality and recurring factors in a weekly Incident Review Board.
17. Enforce action ownership and effectiveness tracking (after 16) from P1 step 19
Treat corrective actions as risk commitments rather than suggestions. Closing a ticket is insufficient without evidence that the control or system behavior improved.
- Give every action one named individual owner, manager, priority, due date, expected risk reduction, verification method, and linked work item.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with intermediate milestones.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Place recurrence-prevention actions for SEV0 and SEV1 ahead of discretionary feature work unless an executive accepts the residual risk.
- Escalate overdue high-risk actions to the manager after seven days, director after 14 days, and CTO after 30 days.
- Require written residual-risk acceptance, compensating controls, and a new review date when high-risk work is deferred.
- Verify completed actions through tests, telemetry, drills, or production evidence.
- Triage the 53 currently open historical actions within 30 days. Complete, re-plan, or formally accept the risk, prioritizing ledger, regional, security, and detection items.
18. Measure outcomes, controls, business impact, and human load (after 10, 11, 14, 17) new
Use a balanced scorecard so teams are not rewarded for suppressing alerts or avoiding incident declarations. Report medians and 90th percentiles, not averages alone.
- Measure time from first impact to internal detection, declaration, acknowledgement, command assignment, mitigation, and resolution.
- Track customer-first detection, missed escalations, status-page timeliness, update-cadence compliance, and role conflicts.
- Track incident count, recurrence, availability, error-budget burn, affected payment value, delayed transactions, reconciliation breaks, and SLA credits.
- Track page volume, actionability, duplicates, after-hours pages, missed acknowledgements, and load per responder.
- Track postmortem timeliness, action completion, action age, verified effectiveness, and repeated contributing factors.
- Track rotation size, duty frequency, recovery days, swaps, attrition signals, and quarterly responder sentiment.
- Hold a weekly Incident Review Board for incidents, actions, missed controls, and noisy alerts.
- Hold a monthly executive reliability review for trends, investment decisions, contractual exposure, and accepted risks.
- Hold a quarterly resilience and controls review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile dashboard data against customer cases and sampled incident records monthly to identify missing incidents or metric gaming.
19. Train participants and address pager resistance (after 6, 7, 8, 13, 14, 16) new
Introduce the operating model as a fair exchange, not an audit mandate. The message is that people are paid, paged only for accepted services, supported by trained command, and given capacity to fix defects.
- Train all employees to recognize impact, declare an incident, and find the incident channel and customer status page.
- Give all 260 engineers role-based training in severity, escalation, evidence preservation, handoff, and financial-integrity precautions.
- Certify incident commanders through instruction, simulation, two shadowed events or exercises, and observed performance.
- Train Communications Leads in status writing, account-manager briefing, contractual clocks, and legal escalation.
- Train scribes in timeline quality, decision capture, fact-versus-hypothesis labeling, and evidence handling.
- Require responders to demonstrate access, dashboards, rollback, runbooks, and escalation competence before primary duty.
- Require two shadow shifts before independent primary on-call.
- Appoint one adoption champion in each team and hold weekly office hours during rollout.
- Publish compensation, fatigue protections, service boundaries, and the accommodation process before assigning shifts.
- Use paid working time for training, exercises, shadowing, runbook work, and certification.
- Survey engineers at baseline, day 60, day 120, and quarterly thereafter.
20. Pilot on the payment critical path (after 8, 9, 11, 12, 13, 14, 17, 19) from P4 step 22
Run a four-to-six-week pilot across the highest-risk customer journey. Use real incidents and exercises to correct the process before wider rollout.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, authentication or API edge, settlement, and Support intake.
- Activate paid primary and secondary rotations, central command, communications, the incident record, status templates, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new system authoritative.
- Review every pilot page within one business day for routing, actionability, responder load, and missing context.
- Have the program team coach incidents without silently taking command from assigned role holders.
- Hold a weekly pilot retrospective and correct critical process or tool defects within 48 hours.
- Require 95% command assignment within five minutes, 95% timely communications, no unpaid pages, complete postmortems, and at least a 50% page-noise reduction before expansion.
21. Roll out by customer journey and risk (after 18, 20) new
Expand in fixed waves rather than waiting for every team to become perfect. Apply explicit readiness gates and time-limited exceptions.
- Days 0–7: operate the interim declaration, command, and communications process.
- By day 30: complete Tier 0 ownership, activate the first certified command roster, approve compensation, and begin the pilot.
- By day 60: provide compensated 24×7 coverage for every Tier 0 and Tier 1 customer journey and route all critical pages through the new platform.
- By day 90: complete critical detection upgrades, customer communications, regulatory playbooks, and the first major alert-noise reduction.
- By day 120: assign every production service a response tier, owner, tested escalation, and appropriate coverage model.
- Roll out remaining teams in two-to-three-week waves ordered by customer and dependency risk.
- Require each wave to pass ownership, alert, runbook, access, training, compensation, and tabletop gates.
- Disable legacy paging paths after verified cutover rather than leaving ambiguous parallel obligations.
- Publish a weekly adoption dashboard by team and escalate failed gates as business risks.
- Never start mandatory night coverage before compensation, staffing, training, and access are ready.
22. Exercise command, regional resilience, and ledger recovery (after 9, 13, 14, 15, 19) new
Validate the process under realistic conditions before depending on it during a crisis. Use the same action-tracking rules for exercises and real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops covering payment failure, customer-first detection, third-party failure, and ambiguous ownership.
- Exercise loss of one AWS region, Kubernetes degradation, PostgreSQL failure, suspected duplicate payments, queue backlog, and settlement risk.
- Exercise simultaneous operational and security events to test command and disclosure boundaries.
- Exercise loss of chat, status-page, identity, or paging providers using telephone and offline fallbacks.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors.
- Validate backups, restoration, RPO, RTO, failover prerequisites, financial controls, and post-recovery reconciliation.
- Avoid uncontrolled production-ledger experiments; use staging, replicas, simulations, or tightly governed production tests.
- Complete at least two cross-company exercises before the audit, including one overnight or unannounced paging test.
23. Test SOC 2 operating effectiveness before the auditor (after 4, 18, 21, 22) new
Demonstrate that the controls operate consistently, not merely that policies exist. Correct failures through tracked remediation rather than rewriting historical records.
- Preserve policy approvals, service ownership, schedules, compensation activation, access reviews, training, certifications, incidents, communications, postmortems, actions, and exercises.
- Sample evidence monthly from initial signal through verified action closure.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Use exceptions to document missed acknowledgements, late communications, incomplete records, and compensating controls.
- Conduct a formal mock audit in month 7 using the populations and interview roles expected by the external auditor.
- Trace at least one SEV0 or SEV1, one SEV2, a customer-reported event, and an exercise end to end.
- Remediate evidence and operating gaps before external fieldwork.
- Brief commanders, responders, Support, and Compliance for auditor interviews without scripting inaccurate answers.
24. Institutionalize continuous improvement (after 18, 21, 23)
Keep incident management active after the audit by assigning permanent owners, budgets, and review cycles. Use incident trends to drive architectural investment.
- Assign permanent owners for the policy, service catalog, paging platform, status page, training program, metrics, and evidence repository.
- Review severity thresholds, communications timing, staffing, and compensation annually and after material process failures.
- Recertify commanders and communications leads annually through observed exercises.
- Review recurring failure families quarterly and require executive action when remediation repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fatigue, fairness, tool friction, and psychological safety.
- Report severe incidents, SLA exposure, overdue risks, resilience investments, and customer-first detection to the board or risk committee quarterly.
- Prioritize reduction of the shared-ledger concentration risk, stronger regional independence, deployment safety, graceful degradation, and automated mitigation.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
- By day 7, every suspected SEV0–SEV2 uses one incident record, one coordination channel, and one named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named commander for more than 10 minutes; at least 95% are assigned within 5 minutes.
- By day 30, 100% of Tier 0 services have a named owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, every Tier 0 and Tier 1 customer journey has compensated 24×7 command and subject-matter coverage.
- By day 120, all 180 production services have an owner, response tier, tested escalation path, and appropriate coverage model.
- No mandatory night rotation starts before its compensation, training, access, and staffing controls are active.
- Every critical primary rotation has at least six qualified responders or a documented, expiring executive exception by day 120.
- No responder is routinely scheduled for primary duty more often than one week in six by day 120.
- At least 95% of critical pages are acknowledged within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by day 120 and under 60 minutes by month 8.
- At least 95% of SEV0 and SEV1 customer notices are issued within 15 minutes of declaration by month 3.
- At least 95% of customer-visible SEV2 notices are issued within 30 minutes by month 3.
- At least 95% of incidents meet their required update cadence by month 3.
- Monthly paging volume falls from 3,400 to no more than 1,500 by day 90 and no more than 700 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6 without loss of critical detection coverage.
- All new paging alerts satisfy the owner, impact, action, dashboard, runbook, deduplication, and escalation standard by day 60.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 currently open historical actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date by month 6, with effectiveness evidence.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Two end-to-end cross-company exercises, including regional failure and ledger recovery, are completed before the audit.
- Monthly availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- SLA credits decline by at least 50% on an annualized trailing basis within 12 months.
- Quarterly on-call surveys show improving fairness and sustainability, with at least 75% favorable responses by month 6.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 223f0a8a-068a-4034-ad20-8e2ff8553040, Agent: opus5_initial_1, LLM: anthropic/claude-opus-5):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 9 months.
- Customer-first detection falls from 40% of incidents to under 10% within 6 months and under 5% within 12.
- Median time to mitigate falls from 3 h 10 min to under 60 minutes within 12 months.
- An Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes.
- Status page updated within 15 minutes of SEV1 declaration and 30 minutes of SEV2 in 95% of cases.
- Monthly alert volume falls from 3,400 to under 500 pages, with actionability above 75%; out-of-hours pages under 2 per person per week.
- All six legacy alerting tools consolidated into one paging platform, legacy paging paths disabled, by week 16.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by their due date.
- SLA credits fall from $1.3M to under $400k in the first 12 months.
- Customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar.
- All 28 teams onboarded by week 24, with 24x7 rotations of 6+ certified responders for every Tier 0/1 team.
- 30+ certified Incident Commanders and 20+ certified Communications Leads, giving 24x7 primary and secondary command cover.
- Paid on-call policy approved by HR, Legal and Finance and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among engineers on rotation.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked action items.
Steps (29):
1. Program charter, executive mandate and funding
Convert the CEO's frustration into a named program with one accountable owner, a budget and a deadline that is earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, the 28 teams will opt out.
- Appoint a single Incident Management Program Lead (Head of Reliability/SRE) with direct exec sponsorship from CTO and CEO.
- Form a steering group: CTO, VP Eng, Head of Support/CS, CISO/Compliance, Legal, Finance (SLA credits), HR (on-call pay).
- Set the non-negotiables: one severity scale, one paging tool, one postmortem format, mandatory action tracking, paid on-call.
- Fix the timeline: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready by week 28 (four weeks of buffer before the audit).
- Approve budget lines: tooling (~$150–250k/yr), on-call compensation (~$600k–1M/yr), 2–3 dedicated program FTEs. Anchor it against $1.3M of credits plus incident cost.
2. Forensic baseline of the 31 incidents and the alert estate (depends on: 1)
Before designing anything, rebuild the facts. Re-open all 31 incidents and profile the 3,400 monthly alerts so every later design decision is evidence-based.
This also creates the "before" picture the exec and the auditor will compare against.
- Re-code each incident: trigger, service, detection source (customer vs monitor), timestamps for detect/acknowledge/declare/mitigate/resolve, who led, credits paid, root cause family.
- Quantify the 40% customer-first detections: which signal was missing in each case.
- Classify the two "nobody in charge" incidents minute by minute; use them as the burning-platform story.
- Audit the six alerting tools: volume per tool, per team, per alert rule; identify the top 50 rules that produce most of the 85% noise; find rules with no owner and no runbook.
- Baseline the numbers formally: MTTD 22 min, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed. Freeze them as the reference line.
3. Stakeholder listening tour and resistance map (depends on: 1)
Engineer pushback against "carrying a pager for other teams' code" is the main delivery risk. Treat it as a design input, not an attitude problem.
Run structured interviews across all 28 teams, plus Support, CS and Sales, in two weeks.
- Test the real objection: is it unpaid work, night sleep, unfamiliar code, poor runbooks, or fear of blame? Each has a different fix.
- Collect current informal practices — the 12 teams already on-call are the pilot candidates and the source of veterans.
- Document the promise that answers the objection: **you are only paged for services your team owns**, plus a trained commander who runs the incident and pulls in others.
- Map influencers and blockers by team; recruit 10–15 credible engineers as a design working group so the process is co-authored, not imposed.
- Survey baseline sentiment (trust in alerts, willingness to be on-call, burnout) to re-measure at 6 and 12 months.
4. Service ownership catalog and criticality tiering (depends on: 2)
You cannot page the right person across 180 services until each service has a named owning team. This is the foundation of both on-call fairness and severity mapping.
Build a machine-readable catalog (Backstage or equivalent) that is the single source of truth for routing.
- One owning team per service, a named engineering manager, a Slack channel, a paging escalation policy, a dependency list.
- Tier services by business impact: Tier 0 (money movement, ledger, auth, shared PostgreSQL cluster), Tier 1 (customer-facing but degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map each Tier 0/1 service to the customer-visible capability it supports (payment initiation, settlement, reporting, onboarding).
- Flag orphan services and cross-team shared components; force an ownership decision for each within 30 days, or schedule decommissioning.
- Publish coverage gaps to the steering group: any Tier 0 service without an owner is an executive escalation.
5. Severity scale and declaration criteria (depends on: 2, 4)
Define a five-level scale with objective, payments-specific triggers so declaration is a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade.
Each level triggers a fixed bundle of response, comms and postmortem obligations.
- **SEV1**: money movement stopped or incorrect, ledger integrity in doubt, data breach, full region loss, >10% of customers impacted. Triggers: immediate 24x7 page of IC + comms + exec, bridge within 5 min, status page within 15 min, mandatory postmortem, regulator assessment.
- **SEV2**: severe degradation, settlement at risk of missing a window, single large/strategic customer fully down, SLA breach likely. Triggers: IC paged, status page within 30 min, mandatory postmortem.
- **SEV3**: partial or workaround-available degradation, no credit exposure. Team-led, business-hours comms, postmortem optional but encouraged.
- **SEV4/5**: minor or internal-only; ticket-tracked, no paging.
- Add auto-escalation rules: any SEV3 open >2 h, or any incident touching the shared ledger cluster, becomes SEV2 automatically. Include a severity decision tree and 12 worked examples drawn from the 31 real incidents.
6. Incident roles, decision authority and handover rules (depends on: 5)
Solve the "nobody in charge for an hour" failure by making command explicit, transferable and logged.
Define five roles with written responsibilities, entry criteria and explicit authority.
- **Incident Commander**: owns the incident, not the fix. Authority to declare severity, pull any engineer, approve customer-impacting mitigations, invoke failover and authorise spend. The IC never types in the terminal.
- **Communications Lead**: owns status page, internal updates, account-manager briefings and the exec summary. Single voice to customers.
- **Scribe**: maintains the timeline, decisions and open questions; feeds the postmortem and the audit evidence trail.
- **Subject-Matter Responders**: engineers from owning teams; they investigate and remediate, and report to the IC.
- **Executive Liaison** (SEV1 only): shields the IC from exec questions and owns regulator/board escalation.
- Rules: the IC role is assumed within 5 minutes of declaration, stated explicitly in the channel ("I am IC"), and any handover is announced and logged. Roles may be combined below SEV2; never at SEV1.
7. 24x7 incident command coverage model (depends on: 6, 4)
Command is staffed by a small, trained, cross-team pool — not by 28 teams individually. This is what makes 24x7 realistic in one New York time zone.
Design a central rotation that scales with a growing certified pool.
- Create a **Duty Incident Commander** rotation of 25–35 certified volunteers (target ~1 per team, plus managers and senior engineers), giving each person roughly one week per 6–8 months.
- Pair with a Duty Comms Lead rotation (Support/CS leads plus engineering managers, ~15–20 people) and a Scribe pool (rotating, lowest barrier, used as the training entry point).
- Coverage: primary + secondary IC at all times; hard 5-minute acknowledgement SLA with automatic failover to secondary, then to the on-call engineering director.
- Night coverage options to evaluate in writing: US-only rotation with paid night stipend now; a Lisbon/Dublin or APAC follow-the-sun cell as a 12-month option; a 24x7 NOC-style triage desk for first-line detection.
- Eligibility: certification required (S21); commanders are volunteers with manager approval and can step out with 30 days' notice.
8. Team on-call structure, rotations and routing rules (depends on: 4, 6)
Rebuild team on-call around the principle that answers the pushback: **you are only paged for code your team owns**.
Apply a tiered obligation so 28 teams are not treated identically.
- Tier 0/1 owning teams (expected 14–18 teams): 24x7 primary + secondary, minimum 6 people per rotation, one-week shifts, handover Wednesday mornings.
- Tier 2/3 teams: business-hours on-call with a best-effort out-of-hours escalation path, no night paging.
- Platform/Infrastructure and Database teams: 24x7, since they own the shared PostgreSQL ledger cluster and the Kubernetes/regional layer.
- Rotations under 6 people are merged across teams or backfilled by hiring; no rotation of fewer than 4 is approved.
- Routing: every page resolves through the service catalog to the owning team's escalation policy; cross-team pages are made by the IC, never by an alert.
- Guardrails: maximum one week in four, no on-call in the week after a SEV1 you led, protected recovery time after any night page, and a per-person page budget (see S10).
9. On-call compensation, labour compliance and fairness policy (depends on: 8, 3)
Unpaid on-call is both a retention risk and a legal exposure in New York. Paying for it is the fastest way to convert resistance into participation.
Design the scheme with HR, Legal, Finance and Payroll, and publish it before asking anyone to sign up.
- Base stipend per week on rotation, differentiated by tier: e.g. $800–1,200 for 24x7 Tier 0/1, $300–500 for business-hours rotations, with premiums for holidays and weekends.
- Per-incident payment for out-of-hours activation (e.g. $150 per night page plus hourly beyond one hour) and guaranteed time-off-in-lieu after night work.
- Separate Duty IC stipend, since command is a distinct and heavier burden.
- Verify FLSA exempt/non-exempt treatment, NY State wage rules and overtime exposure for non-exempt staff; document the legal review.
- Budget and model the annual cost; get board/CFO approval as a line item, benchmarked against $1.3M of credits.
- Add non-cash elements: on-call time counted as delivery load (teams reduce sprint commitment by ~15%), incident leadership recognised in promotion criteria, and a public quarterly report of on-call load per team.
10. Alert quality standard and page budget (depends on: 2, 4)
3,400 alerts a month at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a contractual condition of paging someone.
Publish a standard, then enforce it mechanically.
- Every paging alert must have: a named owning team, a documented customer impact, a runbook link, a tested threshold, and a severity mapping. Alerts failing this are demoted to ticket or deleted.
- Page only on symptoms that affect customers (SLO burn rate, error budget, queue depth against settlement deadlines); cause-based CPU/memory alerts become dashboards or tickets.
- Set a **page budget**: maximum 2 out-of-hours pages per person per week. Breach triggers a mandatory alert-tuning sprint for the owning team and blocks new alert creation.
- Auto-quarantine: any alert that fires more than 5 times a month without action, or has >70% no-action acknowledgements, is silenced automatically and returned to its owner.
- Monthly alert review per team: kill, tune, or keep, with the numbers on screen.
- Target: 3,400 → under 500 pages a month, with actionability above 75% within six months.
11. Detection uplift: SLOs, synthetic journeys and ledger assurance (depends on: 10, 4)
The goal is to stop customers telling you first. Detection must be driven by customer-visible outcomes, not host metrics.
Instrument the money path end to end and alert on it.
- Define SLOs for each Tier 0/1 customer capability: payment initiation success rate, authorisation latency, settlement file timeliness, API availability and reporting freshness. Tie them to the 99.95% contractual SLA with a stricter internal target.
- Deploy synthetic transactions from outside the platform, in both regions, every 60 seconds, covering the full payment lifecycle including a small real-value canary flow where feasible.
- Add ledger assurance checks: continuous double-entry balance reconciliation, replication lag and failover-readiness alarms on the shared PostgreSQL cluster, and settlement-window countdown alerts.
- Build per-customer anomaly detection for the top 100 accounts (volume drop, error spike) so a single-tenant outage is detected before the account manager calls.
- Create an "inbound signal" bridge: any support ticket or account-manager report matching impact keywords auto-creates a triage incident within 5 minutes.
- Track every incident's detection source; make "customer detected first" a reviewed defect with its own follow-up action.
12. Tool consolidation and incident platform implementation (depends on: 5, 7, 8, 10)
Collapse six alerting tools into one paging and incident platform so there is a single queue, a single timeline and a single audit record.
Run a short, time-boxed selection and migrate within the pilot window.
- Select an integrated stack: paging/on-call scheduling plus an incident management layer (e.g. PagerDuty + incident.io/FireHydrant, or a single vendor) and a hosted status page.
- Implement one-command declaration in Slack (`/incident declare`) that creates the channel and bridge, pages the Duty IC, sets severity, opens the timeline and starts the clock.
- Migrate all monitoring sources to route into the one platform; decommission direct paging from the legacy six and block new integrations that bypass it.
- Automate the evidence trail: timestamps, role assignments, severity changes, comms sent, and postmortem linkage exported for SOC 2.
- Integrate with the service catalog for routing, Jira for actions, Salesforce/CS tooling for affected-customer lists, and Zoom/Slack Huddle for the bridge.
- Hard requirement: the platform must work when AWS in one region is down — verify out-of-band paging (SMS/phone) and a printed/offline fallback runbook.
13. Detection-to-escalation path and the five-minute command rule (depends on: 6, 7, 12)
Write the single path from "something looks wrong" to "someone is in charge", and make it impossible to skip.
The design target is detection to commander in under five minutes, any hour.
- Entry points: automated alert, engineer observation, support ticket, account manager, partner bank, customer-facing SEV hotline. All converge on the same declaration command.
- Anyone in the company may declare up to SEV2; nobody is punished for over-declaring. Publish that rule in writing and repeat it.
- Auto-page ladder: Duty IC (5 min) → secondary IC (5 min) → on-call Director (10 min) → CTO. Same ladder for the owning team's responder.
- Cross-team pull: the IC can page any team's on-call directly, with a 10-minute acknowledgement obligation. This is the reciprocal commitment that makes single-team ownership viable.
- Explicit takeover protocol: if no one claims IC within 5 minutes, the platform assigns it and announces it; the assignee cannot decline, only hand over.
- Define standing severity triggers for immediate regional failover, ledger read-only mode and partner-bank notification, with pre-authorised decision rights so the IC does not wait for an executive.
14. Internal communications protocol (depends on: 6, 12)
Standardise the internal channel so responders, executives and support see the same picture without interrupting the IC.
Separate the working channel from the audience channel.
- One incident channel per incident (auto-created), one bridge, and a read-only broadcast channel for executives, Support and Sales.
- Update cadence by severity: SEV1 every 30 minutes even if nothing has changed; SEV2 every 60 minutes; SEV3 at state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, ETA or next update time, current IC and Comms Lead.
- Exec briefing rule: executives ask questions only to the Executive Liaison; the IC is not interrupted. Publish this as a behavioural expectation signed by the exec team.
- Support/CS enablement: a live affected-customer list and a holding statement within 15 minutes of SEV1/SEV2 so the front line is never guessing.
- Handover protocol for incidents beyond 4 hours: formal IC handover checklist, fatigue rule, and staffing of a second shift.
15. Customer communications and status page policy (depends on: 5, 14)
Customers currently learn of outages from their own monitoring and hear from whoever happens to be around. Replace that with a timed, owned, pre-approved process.
The Comms Lead is the single author; templates remove the need to write under pressure.
- Timing commitments: status page posted within 15 minutes of SEV1 declaration and 30 minutes for SEV2; updates every 30/60 minutes; resolution notice within 30 minutes of mitigation; customer-facing summary within 5 business days for SEV1.
- Pre-approve 12–15 templates with Legal and Comms (degradation, delay in settlement, API errors, security event, third-party failure) so nothing needs legal review mid-incident.
- Subscription-based status page with per-component granularity mapped to the customer capabilities from S11, plus an email/webhook/RSS feed.
- Tiered outreach: top 100 accounts get a direct named call or email from their account manager within 30 minutes of SEV1, with a briefing pack from the Comms Lead; long tail gets the status page and a proactive email.
- Rules of language: state impact and next update time, never speculate on cause, never assign blame to a vendor before facts are confirmed.
- Run a quarterly customer-perception check with the top accounts on whether comms were timely and useful.
16. Regulatory, partner and legal notification playbook (depends on: 5, 15)
In payments, some incidents are reportable and the clock starts at detection. Build the assessment into the process so it is never an afterthought.
Work with Legal, Compliance and the CISO to produce a decision tree and contact matrix.
- Map obligations: NYDFS Part 500 (72-hour cybersecurity event notification), state breach laws, GLBA/FTC Safeguards, PCI DSS if card data is in scope, sponsor-bank and card-network contractual notice windows, and any FinCEN/OFAC implications.
- Add a mandatory regulatory-assessment checkpoint to every SEV1 and every security-related SEV2, owned by the Executive Liaison, completed within 2 hours of declaration and recorded even when the answer is "not reportable".
- Build the contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel, with 24x7 numbers and named backups.
- Pre-draft notification letters and hold them under legal privilege review.
- Check customer contracts for bespoke notification SLAs (often 1–4 hours for enterprise accounts) and encode them in the customer tiering.
- Test the playbook once per quarter as part of the simulation programme.
17. SLA credit and financial impact workflow (depends on: 5, 15)
Link incidents to money so severity, credits and prioritisation stay consistent — and so Finance stops being surprised.
Make credit calculation an automated output of the incident record, not a negotiation.
- Define the availability measurement method per contract, per component, and agree it with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and the per-capability telemetry; generate a proposed credit schedule within 5 business days of resolution.
- Decide the posture: proactive credits for the top tier (reputation upside) versus claims-based for the rest; document the approval chain.
- Track credits per incident and per root-cause family; feed a quarterly report showing which reliability investments would have prevented which credits.
- Set a target: reduce credits from $1.3M to under $400k in year one, and use that delta as the ongoing business case.
18. Postmortem standard and blameless review forum (depends on: 5, 6)
Replace "some incidents, various formats" with a mandatory, single-format, blameless process with fixed deadlines.
The discipline is in the deadlines and the forum, not in the template.
- Mandatory for: every SEV1 and SEV2, every incident where a customer detected it first, every incident over 2 hours, every repeat of a known cause, and every near-miss involving the ledger. Optional but templated for SEV3.
- Fixed timeline: draft within 3 business days, peer review within 5, published company-wide within 10. The IC owns delivery; the owning team's manager is accountable.
- One template: timeline, customer and financial impact, detection analysis (why not sooner), response analysis (why mitigation took as long as it did), contributing factors, what went well, action items with owner and due date.
- Blameless rules in writing: describe systems and decisions in the context available at the time; no individual named as a cause; HR and management commit that postmortems are never used in performance reviews.
- Weekly 60-minute Incident Review Board: reviews all postmortems from the prior week, challenges quality, ratifies severity, and approves or rejects action items. Attendance by engineering directors is mandatory.
- Publish a searchable postmortem library and a quarterly "top five recurring causes" analysis.
19. Action item ownership and tracking system (depends on: 18, 12)
11 of 64 actions closed is the single clearest symptom of a process nobody enforces. Give actions the same status as customer commitments.
Track them where engineering work already lives, with visible escalation.
- Every action gets: a named individual owner (not a team), a priority class, a due date and a Jira ticket auto-created from the postmortem.
- Priority classes with hard SLAs: P0 prevents recurrence of a SEV1, due in 30 days; P1 in 60 days; P2 in 90 days. P0s are committed into the next sprint before any roadmap work.
- Capacity rule: teams reserve a standing 15–20% of sprint capacity for reliability and incident actions. Without reserved capacity, the actions will not land.
- Escalation ladder for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Monthly reporting of closure rate by team in the engineering leadership review; include it in manager performance objectives.
- Target: 90% of P0/P1 actions closed on time within two quarters.
20. Runbooks, major-incident playbooks and the on-call readiness bar (depends on: 4, 8)
Nobody can respond well to unfamiliar systems at 3 a.m. without runbooks — and poor runbooks are a real part of the pager resistance.
Define a minimum readiness bar that a service must meet before it is allowed to page anyone.
- Readiness checklist per Tier 0/1 service: current architecture diagram, dependency map, dashboard link, alert-to-runbook mapping, rollback procedure, feature-flag kill switches, escalation contacts, and a data-loss/latency impact statement.
- Write major-incident playbooks for the top failure modes derived from S2: shared PostgreSQL ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank/third-party outage, settlement-window breach, and suspected security compromise.
- Prioritise the shared ledger: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and a clearly stated data-loss tolerance (RPO/RTO) signed off by the exec.
- Runbooks must be tested at least twice a year in a drill; untested runbooks are marked stale in the catalog.
- Enforcement: a service without readiness sign-off cannot create paging alerts, and the gap is reported to its director.
21. Training, certification and the commander academy (depends on: 6, 13, 14, 18)
Command is a skill, not a title. Build a certification path so 24x7 coverage is staffed by people who have practised.
Use a tiered curriculum with real assessment.
- **Scribe** (2 hours): timeline discipline and tooling. The entry point for everyone.
- **Responder** (half a day): severity scale, declaration, escalation, runbook use, comms hygiene. Mandatory for every engineer joining an on-call rotation.
- **Incident Commander** (two days plus shadowing): command presence, delegation, decision-making under uncertainty, severity calls, handover, exec management. Requires two shadowed incidents and one simulated SEV1 before certification.
- **Communications Lead** (one day): status-page writing, customer tiering, legal boundaries, regulator triggers.
- Certification is valid 12 months and renewed via a simulation; the register of certified people is an audit artefact.
- Add on-call onboarding per team: a new joiner shadows two shifts before holding primary, and never holds primary in their first 90 days.
22. Simulation programme: game days, drills and wheel of misfortune (depends on: 21, 12, 20)
The process must be rehearsed before it meets a real SEV1. Simulations also build the commander pool and expose runbook gaps cheaply.
Run a standing calendar rather than one-off exercises.
- Monthly 60-minute tabletop ("wheel of misfortune") per engineering group, using a real past incident from the 31.
- Quarterly full-scale game day in production or a production-like environment: regional failover, ledger replica promotion, dependency failure, with the whole role structure activated and timed.
- Twice-yearly unannounced paging drill to measure real acknowledgement times at night.
- One security-incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Every exercise produces a lightweight postmortem and action items in the same system as real incidents.
- Measure and publish drill metrics: time to IC, time to first status update, time to correct mitigation decision.
23. Pilot with wave 0 teams (depends on: 22, 9, 11)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Run a six-week pilot with tight measurement and a public verdict.
- Select 5–6 teams: core payments, ledger/database, platform/Kubernetes, API gateway, plus two of the 12 teams already on-call.
- Activate the full stack for them: new severity scale, Duty IC rotation, single paging tool, alert budget, status-page policy, mandatory postmortems, paid on-call.
- Hold a weekly pilot retro; expect and document 20–40 process defects, and fix them in the standard before rollout.
- Validate the hard questions: does the 5-minute IC rule hold at 3 a.m.? Do cross-team pulls get answered? Is the severity tree unambiguous?
- Exit criteria: MTTD under 10 minutes for pilot services, IC assigned within 5 minutes in 95% of incidents, page volume down 50%, all postmortems on time, positive on-call sentiment.
- Publish a one-page pilot result to the whole company — this is the main adoption argument for the remaining teams.
24. Metrics, dashboards and the review cadence (depends on: 12, 18, 23)
Instrument the process itself, so improvement is visible and the audit has evidence of monitoring and review.
Define a small set of metrics with owners and a fixed meeting rhythm.
- Response metrics: MTTD, time to declare, time to IC assigned, MTTA, MTTM, MTTR, incidents per month by severity, % detected by customers first.
- Quality metrics: page volume per person per week, alert actionability rate, page budget breaches, postmortem on-time rate, action closure rate and ageing.
- Business metrics: SLA credits paid, availability against 99.95% per capability, error-budget consumption, repeat-incident rate.
- People metrics: on-call load distribution across teams, out-of-hours pages per person, on-call sentiment and attrition among on-call staff.
- Cadence: weekly Incident Review Board (postmortems and actions), monthly Reliability Review (metrics per team, alert hygiene, on-call load), quarterly Executive/Board review (credits, trends, investment asks), annual policy review.
- Every metric gets a target and a named owner; dashboards are self-serve and public inside the company.
25. Wave rollout across all 28 teams with readiness gates (depends on: 23, 24)
Roll out in four waves of six to eight teams, every three weeks, ordered by criticality. Each wave passes an explicit gate rather than a deadline.
Gates keep quality high and make the standard credible.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with 6+ certified responders, one IC candidate nominated, one drill passed.
- Assign each wave a named coach from the program team for three weeks of hands-on support.
- Gate criteria are checked and signed by the director; teams that fail are re-scheduled, not waived.
- Freeze legacy tooling per wave: after onboarding, the old alerting paths are disabled, not left as a fallback.
- Publish a live adoption scoreboard by team so progress is social, not administrative.
26. SOC 2 control mapping, evidence automation and internal dry-run audit (depends on: 18, 19, 25)
Design the process so audit evidence is a by-product of doing the work, then test it before the auditors do.
Engage the auditor early to confirm the interpretation of controls.
- Map the process to the Trust Services Criteria: CC7.3 and CC7.4 (incident identification, response, recovery), CC7.2 (monitoring), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), plus availability criteria A1.2.
- Produce and approve formal policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard — versioned, signed, annually reviewed.
- Automate evidence: incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, postmortem library, action-item closure reports, training and certification register, drill records.
- Confirm the observation window with the auditor and ensure the process is operating for a minimum of three months before fieldwork.
- Run an internal dry-run audit at month six: sample 15 incidents and walk the full evidence chain; fix gaps with 8 weeks to spare.
- Keep a remediation log for any incident where the process was not followed, with the corrective action — auditors respond better to documented exceptions than to a claim of perfection.
27. Change management, incentives and communications campaign (depends on: 3, 9, 23)
Run this in parallel from day one. The process will be judged by engineers on fairness, and by executives on visible results.
Communicate the deal explicitly and repeatedly.
- The deal in one sentence: **you are paid for on-call, you are paged only for what you own, a trained commander runs the incident, and your postmortem actions get real sprint capacity**.
- Launch communications: CTO all-hands, per-team roadshows, a one-page process card for laptops, an internal wiki hub and a Slack support channel with a 4-hour answer SLA.
- Recognition: incident-response contribution in promotion criteria and performance frameworks, quarterly awards for best postmortem and biggest alert-noise reduction, public thanks after every SEV1.
- Manager accountability: adoption, alert hygiene, action closure and on-call load in each engineering manager's quarterly objectives.
- Handle the exceptions: a written path for engineers who cannot do nights (caring responsibilities, health), covered by stipended volunteers elsewhere.
- Track sentiment quarterly and publish the results, including bad news, to keep credibility.
28. Program risk register and contingency planning (depends on: 1)
Name the ways this program fails and pre-commit the response. Review it monthly in the steering group.
The main risks are predictable.
- **Volunteer shortfall for the IC pool**: contingency is to make command a rostered duty for engineering managers and senior engineers until the pool reaches 25.
- **Compensation not approved in time**: fall back to time-off-in-lieu plus a phased stipend, but do not launch mandatory night on-call without some compensation.
- **Tool migration slipping**: keep the single-queue requirement and cut scope on the incident-management layer, not on paging consolidation.
- **Alert pruning causing a missed incident**: prune from paging to ticket first, observe for 30 days, then delete; keep a recovery path.
- **Burnout or attrition among the 12 experienced on-call teams**: monitor load weekly and cap individual page counts.
- **A major SEV1 mid-rollout**: pre-agree that the program lead becomes a full-time responder and the wave schedule slips by one wave, with the steering group informed the same day.
29. Continuous improvement, maturity roadmap and post-audit sustainability (depends on: 24, 25, 26)
Protect against the classic failure: the process decays once the audit is passed. Build the second-year plan before the first year ends.
Set a maturity model and a forward roadmap with owners.
- Quarterly process retrospective with the IC pool: what in the process itself slowed us down, what needs simplifying, is the severity scale calibrated?
- Re-baseline targets every six months; a process that hits all targets is under-ambitious.
- Year-two roadmap candidates: follow-the-sun coverage cell, automated mitigation and self-healing for the top three recurring causes, error-budget policy that gates releases, per-customer real-time impact reporting, and blast-radius reduction for the shared ledger cluster (the largest single structural risk).
- Move from lagging metrics (MTTR) to leading ones (error-budget burn, near-miss rate, drill performance).
- Make the annual policy review, certification renewal and drill calendar permanent calendar items owned by the Head of Reliability, independent of the audit cycle.
- Report to the board quarterly on availability, credits and incident trends so the process keeps executive attention after SOC 2 is signed.
Previous Proposal 2 (ID: 7ee53cee-6357-4ed7-bd99-535e92e8b1e6, Agent: gpt5.6-sol_initial_2, LLM: openai/gpt-5.6-sol):
Estimated Complexity: high
Success Metrics: - Within 7 days, every suspected SEV0–SEV2 has one incident record, one channel, and a named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named incident commander for more than 10 minutes.
- By day 30, 100% of Tier 0 and Tier 1 services have a named owner, primary escalation, secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and subject-matter coverage.
- By day 120, 100% of production services have sustainable ownership and tested escalation paths.
- At least 95% of SEV0 and SEV1 pages are acknowledged within 5 minutes by month 3.
- At least 95% of SEV2 pages are acknowledged within 10 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to 10 minutes by day 90 and 5 minutes by month 6.
- The proportion of incidents first detected by customers falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to below 90 minutes by day 120 and below 60 minutes by month 6.
- At least 95% of qualifying incidents meet their initial customer-communication deadline by month 3.
- At least 95% of published incidents meet their required update cadence by month 3.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without loss of Tier 0 or Tier 1 detection coverage.
- Monthly pages fall from 3,400 to no more than 1,500 by day 90, with alert actionability and missed-detection reviews used as countermeasures against unsafe suppression.
- 100% of new paging alerts satisfy the owner, runbook, dashboard, action, severity, and escalation quality rules by day 60.
- 100% of required SEV0 and SEV1 postmortems are drafted within 3 business days and reviewed within 5 business days by month 2.
- At least 90% of postmortem actions are completed by their approved due dates by month 6.
- All 53 currently open historical actions are triaged within 30 days; all unaccepted high-risk items are completed within 90 days.
- Repeat incidents with the same unaddressed contributing factor decline by at least 50% within 6 months.
- Every primary rotation has at least six trained responders or a documented, time-limited executive exception by day 120.
- No responder is routinely scheduled more frequently than one primary week in six by day 120.
- Two end-to-end cross-company exercises, including regional and ledger scenarios, are completed before the audit, with all critical findings assigned and tracked.
- Monthly availability meets or exceeds the 99.95% contractual target by month 6, with exceptions reviewed at the executive reliability meeting.
- SLA credits decline by at least 50% on an annualized trailing basis by month 8.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (19):
1. Establish ownership, authority, and funding
Launch the program within 48 hours under an executive sponsor. Give one program owner authority to standardize incident management across all 28 teams.
- Name the CTO or equivalent as executive sponsor and a Head of Incident Management or Reliability as directly accountable owner.
- Form a working group with Engineering, SRE or Platform, Product, Support, Customer Success, Communications, Security, Legal, Compliance, Risk, HR, Finance, and Internal Audit.
- Approve authority for an incident commander to stop deployments, roll back releases, disable features, shift traffic, invoke continuity plans, and pause payment processing when integrity is at risk.
- Preserve financial controls. The incident commander may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access requirements.
- Fund paging tools, compensation, training, observability work, exercises, and dedicated reliability capacity.
- Reserve engineering capacity for incident remediation. Start with 10% of capacity and adjust through quarterly risk reviews.
- Record the current baselines: 31 customer-impacting incidents, 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M in credits, 3,400 monthly alerts, 85% noise, and 11 of 64 actions closed.
- Maintain a risk register for staffing gaps, shared-ledger concentration, regional failover, alert coverage, third parties, and audit readiness.
2. Install immediate minimum controls (depends on: 1)
Put an interim process in place during the first seven days. Do not wait for tool consolidation, policy perfection, or the SOC 2 audit.
- Publish a one-page interim severity guide and incident declaration procedure.
- Establish one continuously monitored incident declaration path through chat, telephone, and the paging system.
- Create a standard incident channel, conference bridge, incident document, and event naming convention.
- Staff an interim primary and backup incident commander at all times. Compensate this duty retroactively under the final compensation policy.
- Give trained duty personnel access to the status page, paging system, dashboards, support queue, service catalog, and emergency contacts.
- Require an incident commander to be named within 10 minutes for every suspected major incident.
- Direct Support to escalate credible customer reports immediately rather than waiting for engineering confirmation.
- Triage all 53 open historical postmortem actions. Complete, re-plan, or formally risk-accept the items affecting ledger integrity, payment duplication, regional resilience, security, and detection first.
- Hold a daily 15-minute operational review until permanent controls are working.
3. Create the service and dependency catalog (depends on: 1)
Build a reliable ownership map for all production services and customer journeys. This is the basis for paging, escalation, impact assessment, and audit evidence.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, data stores, queues, external processors, banking partners, and customer-facing endpoints.
- Assign each component a single accountable team, primary responder group, secondary escalation group, engineering manager, and product owner.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 according to financial integrity, customer impact, dependency centrality, and contractual obligations.
- Treat the ledger, payment orchestration, authentication, settlement, reconciliation, and critical shared infrastructure as Tier 0 or Tier 1.
- Map every important customer journey to its service, database, cloud-region, and third-party dependencies.
- Record SLOs, RTOs, RPOs, data classification, dashboards, runbooks, deployment controls, feature flags, and failover methods.
- Assign separate but coordinated responders for the ledger application and the shared PostgreSQL platform.
- Document whether each service is active-active, active-passive, or region-bound. Identify dependencies that make nominal regional redundancy ineffective.
- Make missing ownership or missing runbooks a release-blocking risk for Tier 0 and Tier 1 services.
4. Adopt severity and incident lifecycle standards (depends on: 1)
Approve one impact-based severity model for operational, security, data, and third-party incidents. When evidence is incomplete, start at the higher credible severity and downgrade later.
- SEV0, crisis: Use for actual or credible unauthorized, lost, duplicated, or corrupted movement of money; ledger integrity loss; material security compromise; material data exposure; both-region failure; or an event likely to require crisis or regulatory management. Page all roles immediately, engage executives, Security, Legal, Compliance, and Risk, and consider pausing payment activity.
- SEV1, critical: Use for widespread inability to initiate, process, settle, or reconcile payments; a core journey failing without a viable workaround; material regional impact; fast SLA-budget exhaustion; or an imminent integrity risk. Staff all incident roles, notify the executive duty officer, and publish customer communications.
- SEV2, major: Use for a material customer subset, one or more critical customers, significant degradation with a workaround, partial transaction failure, or a likely contractual impact. Assign an incident commander and subject-matter responders; add communications and scribe roles whenever customers are affected.
- SEV3, minor: Use for localized, low-impact degradation with no financial-integrity, security, regulatory, or material contractual risk. The owning team leads the response and keeps an internal record; external communication is not normally required.
- Base severity on actual or credible impact, not the seniority of the reporter, number of alerts, or presumed complexity of the fix.
- Permit any employee to declare an incident. Only the incident commander may lower severity after recording the evidence and rationale.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of customer impact. Define resolution only after stability, backlog processing, transaction recovery, and required ledger reconciliation are complete.
- Start measurement from the earliest reliable indication of impact, including telemetry, customer reports, and partner notifications.
5. Define roles and sustainable 24x7 staffing (depends on: 3, 4)
Separate command from technical remediation. This allows trained commanders to coordinate any incident without asking engineers to debug code they do not own.
- Incident commander: Owns severity, priorities, role assignment, escalation, decision cadence, mitigation strategy, handoffs, and final closure. One person has command at a time.
- Communications lead: Owns internal notices, status-page updates, account-manager briefs, approved customer language, and coordination with Legal or regulators.
- Scribe: Maintains a timestamped timeline of observations, decisions, commands, owners, and status changes. Automation may assist but does not replace human validation for SEV0 and SEV1.
- Subject-matter responders: Diagnose and mitigate only services or domains for which they have accepted ownership, training, access, and runbooks.
- Executive duty officer: Removes organizational obstacles and approves exceptional business decisions. This role does not take command unless a formal transfer occurs.
- Security, Legal, Compliance, Support, Vendor Management, and Business Continuity join according to predefined triggers.
- Create a company-wide incident-command rotation with at least eight certified primary commanders and eight qualified backups. Use weekly rotations with explicit handoffs.
- Create similarly sustainable communications and scribe pools using Engineering Operations, Support, Customer Operations, and Communications personnel.
- Group service responders into approximately 8–12 coherent product or platform domains rather than creating 28 fragile rotations. Each domain rotation should normally contain at least six trained responders.
- Do not place an engineer into another team's responder pool without training, access, runbooks, shadow shifts, and explicit acceptance by both teams.
- Maintain dedicated database-platform and ledger-application escalation coverage for the shared PostgreSQL environment.
- Require distinct people for commander, communications, and primary technical lead during SEV0 and SEV1 incidents.
- Require a verbal and written handoff when any incident role changes. Record the exact time and the new role owner.
6. Implement compensation and fatigue safeguards (depends on: 5)
End unpaid on-call before expanding coverage. Treat availability, interrupted personal time, and overnight recovery as compensable work.
- Pay a fixed stipend for each primary on-call week and a secondary stipend equal to a defined percentage of the primary amount.
- Pay a higher holiday stipend. Apply overtime and call-out rules to non-exempt employees as required by law.
- Give exempt employees a minimum call-out credit or equivalent paid recovery time for material after-hours work.
- Provide a paid recovery day after prolonged overnight work, a SEV0, or a qualifying SEV1. Managers must arrange daytime coverage rather than expecting normal output.
- Have HR, Finance, and employment counsel publish dollar amounts, tax treatment, eligibility, and payroll procedures within 14 days. Apply the policy consistently across teams and locations.
- Target rotations no more frequent than one week in six. Exceptions require a time-limited staffing plan and executive risk acceptance.
- Avoid consecutive primary and secondary weeks. A person must not be primary for two simultaneous domain rotations.
- Track after-hours pages, sleep interruptions, swaps, missed acknowledgements, and reported burnout by rotation.
- Trigger a staffing or alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
- Permit responders to declare themselves temporarily unfit after overnight work without performance penalty.
7. Consolidate incident and paging tooling (depends on: 4)
Create one operational system of record while migrating safely from the six current alerting tools. Consolidation must reduce ambiguity without creating a monitoring gap.
- Select one enterprise paging and escalation platform and one integrated incident record.
- Initially ingest events from all six tools. Deduplicate, correlate, and route them through the new platform before retiring sources.
- Integrate paging with chat, conference bridges, ticket tracking, the service catalog, observability tools, and the customer status page.
- Automatically capture declaration time, acknowledgements, role assignments, severity changes, messages, decisions, mitigated time, and resolved time.
- Use role-based access, multifactor authentication, break-glass controls, immutable audit logs, and periodic access reviews.
- Provide mobile and telephone fallback paths if chat, identity, or the primary paging tool is unavailable.
- Test paging, escalation, status publication, and conference access every week.
- Retire a legacy alert path only after its signals have named owners, successful end-to-end tests, and at least two weeks of verified operation in the new platform.
8. Improve detection and enforce alert quality (depends on: 3, 7)
Shift detection toward customer journeys, payment outcomes, and ledger integrity. Infrastructure metrics alone will not solve the current customer-first detection problem.
- Instrument payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, and reporting with SLOs and business-level success metrics.
- Run external synthetic transactions and API checks from outside the production boundary and from both AWS regions.
- Monitor transaction failure rates, processing latency, queue age, unprocessed volume, reconciliation breaks, unexpected ledger balances, duplicate identifiers, regional asymmetry, and third-party response quality.
- Correlate application telemetry with Kubernetes, AWS, PostgreSQL, network, deployment, and feature-flag events.
- Route high-priority support cases and credible partner notifications into the same incident declaration path within five minutes.
- Define noise as a page that is duplicate, informational, unactionable, non-production, or requires no timely human action.
- Require every paging alert to name an owner, affected service, urgency, customer or SLO risk, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Send non-urgent conditions to a ticket queue rather than a pager.
- Run new alerts in shadow mode for at least seven days unless an emergency risk exception is approved. Test both firing and recovery behavior.
- Review any alert with less than 50% actionability or more than three firings in seven days within two business days.
- Never silently disable a noisy alert. Verify compensating detection, record the decision, and assign a correction owner first.
- Review alert actionability, false positives, missed detection, and page load with every responder group each month.
9. Codify acknowledgement and escalation paths (depends on: 3, 4, 5, 7, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts when a qualifying signal or customer report enters the system.
- For SEV0 and SEV1, page the owning primary immediately; page the secondary after five unacknowledged minutes; page the domain manager and company incident commander at 10 minutes; and engage the executive duty officer by 15 minutes.
- For SEV2, require primary acknowledgement within 10 minutes and incident-command assignment within 15 minutes. Escalate to the secondary and manager when either target is missed.
- For SEV3, require acknowledgement within 30 minutes when immediate production action is needed. Otherwise create a prioritized work item.
- Automatically page the company incident commander for any credible integrity or security concern, cross-team event, customer-visible Tier 0 failure, regional event, or unresolved ownership question.
- If impact remains unknown after 15 minutes, raise severity rather than waiting for certainty.
- Let the incident commander summon dependency owners, cloud support, database support, payment processors, banking partners, and vendors through maintained escalation contacts.
- Test vendor contacts and premium-support entitlements quarterly.
- Route alerts with no valid owner to the central command rotation, then treat the missing ownership record as a control defect.
- Require human acknowledgement. Delivery to a device or chat channel does not count.
- Record every missed acknowledgement, failed escalation, and manual contact workaround for review.
10. Standardize live incident execution (depends on: 4, 5, 7, 9)
Give responders one concise operating procedure for the first minutes through resolution. Prioritize limiting customer and financial harm before proving a root cause.
- Open a dedicated channel, bridge, incident record, and timeline immediately for SEV0 through SEV2.
- Have the incident commander state severity, known impact, current hypothesis, immediate objective, assigned roles, and next update time.
- Freeze unrelated production changes during SEV0 and SEV1 incidents. Record exceptions approved by the incident commander.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Use pre-approved runbooks for region failover, Kubernetes recovery, PostgreSQL failover, credential rotation, queue recovery, and payment suspension.
- Guard against split-brain, replay, duplication, and out-of-order processing during regional or database recovery.
- Require reconciliation and controlled backlog processing before declaring payment or ledger incidents resolved.
- Keep diagnosis and mitigation workstreams separate when enough responders are available.
- State decisions and owners aloud and in the incident record. Avoid unrecorded direct-message command paths.
- Require stability for a severity-specific observation period before closure. Reopen the incident if impact recurs during that period.
- Conduct an explicit operational handback to the owning team, Support, and Customer Success.
11. Standardize internal, customer, and regulatory communications (depends on: 4, 5, 7, 10)
Communicate known impact early without waiting for a root cause. Use approved facts, acknowledge uncertainty, and give the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish an initial customer status within 15 minutes when disclosure is operationally and legally appropriate, then update every 15 minutes.
- For SEV1, notify internal stakeholders within 15 minutes, publish an initial customer status within 15 minutes, and update at least every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send an initial notice within 30 minutes. Update at least every 60 minutes.
- Do not normally publish SEV3 events. Notify specifically affected customers if contracts or material impact require it.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. State affected capabilities, customer symptoms, workarounds, regions, and next update time.
- Do not speculate about root cause, blame, security scope, recovery time, or data integrity.
- Give account managers a single approved briefing and an affected-customer list. Prohibit contradictory or improvised incident explanations.
- Maintain templates for outages, delays, data-integrity investigation, third-party failure, regional failure, security events, and resolution.
- Issue a resolution notice only after operational recovery and required reconciliation. Provide a customer-facing incident summary within five business days for qualifying events.
- Have Legal and Compliance maintain a jurisdiction, regulator, sponsor-bank, network, cyber-insurer, partner, and contract notification matrix.
- Where applicable, explicitly track the current New York cybersecurity-event notification clock, including the 72-hour requirement, without assuming every incident is reportable.
- Have Legal record the reportability decision, decision time, evidence, approver, deadline, and submission confirmation.
- Allow Security or Legal to limit public detail during an active threat, but require the reason and an alternative stakeholder plan to be recorded.
- Coordinate service-credit calculations and contractual notices with Finance and Customer Success from the same incident record.
12. Make postmortems mandatory and actionable (depends on: 4, 7, 10)
Use postmortems to improve systems and controls, not to assign personal blame. Keep performance or misconduct processes separate from the learning review.
- Require a postmortem for every SEV0 and SEV1.
- Require one for a SEV2 that affected customers, incurred credits, breached an SLO or contract, involved financial or data integrity, repeated a prior failure, exposed a control gap, or lasted more than two hours.
- Permit incident command, Security, Compliance, or the service owner to require a review for a near miss.
- Produce a factual draft within three business days and hold the cross-functional review within five business days.
- Use one template covering executive summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, what worked, what failed, and lessons.
- Include why monitoring did or did not detect the event before customers.
- Avoid a single-root-cause assumption. Examine technical, organizational, process, dependency, testing, and incentive factors.
- Give every action one owner, due date, priority, expected risk reduction, verification method, and linked engineering item.
- Classify actions as containment due within 7 days, corrective work due within 30 days, or strategic work normally due within 90 days.
- Require director approval and documented residual-risk acceptance for overdue high-risk actions.
- Verify effectiveness after implementation. Closing a ticket without evidence does not close the action.
- Publish broadly useful reviews internally. Maintain access-restricted versions for security, privacy, personnel, or legally privileged details.
13. Measure performance and review it routinely (depends on: 7, 8, 11, 12)
Use outcome, process, quality, and human-sustainability measures together. Do not reward teams for suppressing declarations or hiding incidents.
- Measure detection time from first impact to first internal signal, declaration time, acknowledgement time, role-staffing time, mitigation time, resolution time, and recurrence.
- Report both median and 90th percentile. Break results down by severity, service tier, customer journey, region, detection source, and owning domain.
- Track customer-first detection, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, and role conflicts.
- Track availability, error-budget consumption, failed-payment volume, delayed value, reconciliation breaks, impacted customers, contractual breaches, and service credits.
- Track alert volume, actionability, duplicates, after-hours pages, missed pages, pages per responder, and tool-source distribution.
- Track required postmortems completed on time, actions completed by due date, action age, verified effectiveness, and repeat contributing factors.
- Track rotation size, on-call frequency, swaps, recovery days, attrition signals, and quarterly responder sentiment.
- Hold a weekly operational review for recent incidents, overdue actions, alert problems, and upcoming risk.
- Hold a monthly executive reliability review covering trends, investment decisions, accepted risks, and SLA exposure.
- Hold a quarterly resilience and control review with Security, Compliance, Risk, Internal Audit, and Product leadership.
- Use team scorecards to direct investment and assistance, not individual performance penalties.
- Reconcile dashboard data against a monthly sample of incident records and customer cases to detect metric gaming or missing incidents.
14. Train and certify participants (depends on: 4, 5, 9, 10, 11, 12)
Train people before assigning full independent duty. Use paid working time for training, shadowing, exercises, and certification.
- Train all employees to recognize impact, declare an incident, and find the incident channel and status page.
- Train engineers and Support on severity, escalation, customer-report handling, evidence preservation, and financial-integrity precautions.
- Certify incident commanders through instruction, tabletop exercises, shadow incidents, and observed command performance.
- Train communications leads in status writing, contractual communications, regulator escalation, and avoiding unsupported claims.
- Train scribes in timestamping, decision capture, evidence hygiene, and separating fact from hypothesis.
- Require subject-matter responders to demonstrate dashboard, runbook, rollback, failover, and access competence for their assigned domain.
- Add training to new-hire onboarding and repeat role-specific certification annually.
- Appoint an incident-management champion in each of the 28 teams to collect feedback and support local adoption.
- Conduct listening sessions focused on pager fairness, cross-team boundaries, tooling friction, and psychological safety.
- Publish that command duty is process coordination, not responsibility for understanding or repairing another team's code.
15. Pilot and expand the on-call model (depends on: 3, 5, 6, 8, 9, 14)
Pilot the model on the highest-risk customer journeys before expanding it. Correct staffing, alert, access, and compensation defects at each stage gate.
- Start with the ledger, payment orchestration, Kubernetes platform, PostgreSQL platform, authentication, settlement, and Support intake.
- Run the command, communications, and domain rotations in parallel with existing paths for two weeks.
- Verify primary and secondary coverage, handoffs, access, runbooks, paging, conference access, status publication, and compensation processing.
- Require at least two shadow shifts before independent primary duty.
- Review every pilot page within one business day for routing accuracy, actionability, responder load, and missing context.
- Expand by customer journey and dependency domain, not by arbitrary team order.
- Provide central first-line triage if useful, but keep technical remediation with the accepted service owner.
- Do not use contractors or a managed service as the sole incident commander or sole owner of payment and ledger remediation.
- Permit temporary shared domain rotations only after service owners document training, access, runbooks, and escalation boundaries.
- Set an executive-reviewed deadline and remediation plan for any production service that cannot provide sustainable 24x7 ownership.
16. Exercise regional, ledger, and communication failures (depends on: 10, 11, 14, 15)
Validate the process under realistic conditions before relying on it. Begin in tabletop and staging environments, then use controlled production tests where risk permits.
- Run a company-wide incident-command tabletop within 30 days of policy approval.
- Exercise loss of one AWS region, Kubernetes control-plane degradation, shared PostgreSQL failure, payment-processor failure, queue backlog, credential compromise, and suspected duplicate payments.
- Exercise a simultaneous operational and security event to test command boundaries and disclosure control.
- Exercise status-page failure and loss of the primary chat or paging provider.
- Exercise overnight staffing, role handoff, executive escalation, account-manager messaging, and a potential regulator-notification decision.
- Validate backups, restore procedures, RPO, RTO, failover prerequisites, and post-recovery reconciliation.
- Do not inject uncontrolled changes into the production ledger. Use replicas, staging, simulations, or tightly governed production tests.
- Record exercise observations as tracked actions under the same ownership and due-date rules as incident actions.
- Run at least one domain exercise per quarter and two cross-company exercises before the SOC 2 audit.
17. Execute a time-boxed enterprise rollout (depends on: 2, 6, 7, 11, 12, 13, 15)
Use fixed implementation waves so the audit deadline does not become the start date. Report progress weekly and escalate missed stage gates as business risks.
- Days 0–7: Establish governance, interim command coverage, one declaration path, provisional severity, and daily operational reviews.
- By day 14: Approve the core policy, role definitions, communications timings, compensation design, and historical-action triage.
- By day 30: Complete Tier 0 ownership, certify the first command roster, begin the on-call pilot, enable standard incident records, and run the first tabletop.
- By day 60: Provide 24x7 coverage for all Tier 0 and Tier 1 customer journeys, integrate the six alert sources, and enforce postmortem tracking.
- By day 90: Migrate critical paging, implement customer-journey detection, complete status and regulatory playbooks, and materially reduce alert noise.
- By day 120: Assign sustainable ownership and escalation for every production service and complete the first controlled regional or continuity exercise.
- By day 180: Complete tool retirement decisions, verify action closure, rerun weak scenarios, and demonstrate improving detection and mitigation trends.
- In month 7: Conduct a mock audit and executive readiness review, leaving at least one month to correct evidence or operating defects.
- Use exception records with owners, expiry dates, compensating controls, and executive approval. Do not allow indefinite verbal exceptions.
18. Build SOC 2 evidence as the process operates (depends on: 1)
Design evidence collection at the start rather than reconstructing it before the audit. Demonstrate both control design and sustained operation.
- Map the incident process to applicable SOC 2 criteria with Compliance and the auditor, including detection, response, communication, change management, access, availability, and corrective action.
- Maintain approved, version-controlled policies, procedures, severity definitions, role descriptions, and exception records.
- Preserve rotation schedules, compensation activation, training attendance, certification, paging tests, access reviews, and exercise results.
- Preserve incident declarations, timestamps, role assignments, communications, decisions, status updates, postmortems, and corrective-action evidence.
- Record regulatory and contractual notification assessments, including decisions that no notification was required.
- Define retention, confidentiality, legal-hold, and access requirements for operational and security records.
- Sample evidence monthly and trace incidents from initial signal through action verification.
- Have Internal Audit or an independent control owner test the process in months 4 and 6.
- Correct control failures through tracked actions rather than editing historical records.
- Conduct the formal mock audit in month 7 using the same evidence populations expected for the external audit.
19. Sustain accountability and continuous improvement (depends on: 13, 17, 18)
Make incident management an operating discipline rather than an audit project. Keep policy, staffing, tools, and investment aligned with changing customer and system risk.
- Assign permanent owners for the incident policy, paging platform, status page, service catalog, training program, and metrics.
- Review severity thresholds, communication timings, compensation, and staffing at least annually and after material incidents.
- Use incident trends to prioritize architectural work on the shared ledger, regional independence, deployment safety, dependency isolation, and graceful degradation.
- Review repeat incidents and repeat contributing factors quarterly. Require executive action when remediation repeatedly loses priority.
- Survey responders quarterly and publish actions addressing fatigue, fairness, psychological safety, and tool friction.
- Recognize effective incident leadership, early declaration, useful postmortems, and preventive work.
- Prohibit retaliation for good-faith incident declaration or escalation.
- Provide the board or risk committee a quarterly summary of severe incidents, SLA exposure, regulatory events, overdue high-risk actions, and resilience investment.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
Previous Proposal 3 (ID: c051663c-fdf6-4324-bfbe-a5f42ed62eed, Agent: qwen3.8-max_initial_3, LLM: alibaba/qwen3.8-max):
Estimated Complexity: high
Success Metrics: - **MTTD reduced from 22 minutes to <5 minutes** within 6 months of full rollout.
- **Customer-first detection rate reduced from 40% to <10%** within 6 months.
- **MTTM for SEV1 incidents reduced from 3 h 10 min to <60 minutes** within 9 months.
- **Monthly SLA credit payouts reduced from ~$325K to <$100K per quarter** within 12 months.
- **Alert volume reduced from 3,400/month to <600 actionable alerts/month** within 6 months; signal-to-noise ratio >80%.
- **Postmortem completion rate: 100% of SEV1/SEV2 incidents** have a blameless postmortem within 5 business days.
- **Postmortem action-item completion rate >90% within 30 days** of the postmortem (up from ~17%).
- **Zero incidents with >15 minutes of unowned command** (down from 2 incidents with >1 hour).
- **100% on-call coverage**: all 28 teams staffed with primary + secondary on-call 24×7 within 14 weeks.
- **On-call compensation adopted**: 100% of on-call engineers receiving stipends and page pay; on-call satisfaction score ≥4/5 in quarterly survey.
- **Status-page first-update within 15 minutes for SEV1 and 30 minutes for SEV2**, 100% compliance.
- **SOC 2 Type II audit passed** at month 8 with zero incident-response findings.
- **All 260 engineers trained**; 56+ certified ICs and 28+ certified CLs active within 14 weeks.
- **Six alerting tools consolidated to one** within 6 months; legacy tools decommissioned.
- **Regulator notification process tested**: at least one tabletop exercise includes a NY DFS / FinCEN notification drill, and the legal/compliance playbook is documented and approved.
- **Quarterly IMC reviews held consistently** with published KPI dashboards and action-item tracking.
- **On-call participation resistance resolved**: <10% of engineers report 'unwilling to participate' in the 6-month pulse survey (baseline to be measured in S1).
Steps (13):
1. Assess Current State and Baseline Metrics
Build an evidence-based picture of the current incident management reality before designing anything new.
Collect and catalog the last 12 months of incident data: all 31 customer-impacting incidents, 3,400 monthly alerts, on-call coverage gaps across the 28 teams, and the 11 of 64 closed postmortem action items. Interview one lead from each of the 28 teams to surface pain points, political concerns (the 'carrying a pager for other teams' pushback), and tool sprawl.
Deliverables to produce:
- **Alert inventory**: which of the six alert tools feed which teams, alert volume per team, noise rate per tool, and overlap between tools.
- **Incident timeline analysis**: median detection-to-notification-to-mitigation-to-resolution times, who detected first (internal vs. customer), who mitigated, and where handoff gaps occurred.
- **On-call coverage map**: which 12 teams have on-call, which 16 do not, rotation length, compensation status, and escalation paths (or lack thereof).
- **Postmortem audit**: format variance, action-item tracking gaps, and the two incidents with no clear owner for over an hour.
- **Tooling and integration audit**: Kubernetes observability stack, six alerting tools, status-page provider, communication channels (Slack, email, phone), and any existing runbooks.
- **Compliance gap analysis**: SOC 2 Type II CC7.3/CC7.4 requirements vs. current practice, with a risk register for the eight-month window.
- **Peer benchmarking**: incident management practices at 3–4 comparable B2B fintech platforms (e.g., Plaid, Stripe, Adyen) for severity scales, on-call comp, and MTTR targets.
2. Secure Executive Sponsorship and Form the IM Governance Body (depends on: 1)
Anchor the program with visible, top-down authority so that 28 teams adopt changes they did not individually request.
The CEO email about 'outages we hear about from clients' is a ready-made mandate. Convert it into a formal sponsorship structure.
- Appoint an **executive sponsor** (CTO or VP Engineering) who owns the program end-to-end and reports to the CEO monthly.
- Create an **Incident Management Office (IMO)**: one dedicated senior incident management lead, one tooling/platform engineer, and one part-time data analyst.
- Establish an **Incident Management Council (IMC)**: one engineering manager from each of the 28 teams, plus the VP of Customer Success, a compliance lead, and a security lead. The IMC meets bi-weekly during rollout, monthly thereafter.
- Draft and circulate an **executive mandate memo** that states: incident response is a shared operational obligation, not a per-team favor; participation in on-call rotations is a condition of employment for production-facing roles; and the program is not optional pending the SOC 2 audit.
- Allocate a dedicated budget line for on-call compensation, tooling consolidation, status-page licensing, training, and external facilitation.
3. Define Severity Levels and Automatic Triggers (depends on: 1)
Replace the current ad-hoc triage with a five-level severity taxonomy that every engineer, support agent, and account manager can apply in under 60 seconds.
- **SEV1 – Critical**: Ledger data corruption or loss, complete payment processing halt, confirmed data breach affecting customer PII or funds, or regulatory reporting breach. Triggers: automatic all-hands page to on-call, CTO + CEO paged within 10 minutes, dedicated bridge call within 5 minutes, status-page update within 15 minutes, regulator notification assessment within 1 hour, customer comms within 30 minutes.
- **SEV2 – High**: Payment processing degraded >30% throughput or >5% error rate, single-region failover failure, ledger read-only mode, or any condition likely to breach the 99.95% SLA within the current window. Triggers: primary + secondary on-call paged, incident commander assigned within 10 minutes, bridge call within 15 minutes, status-page update within 30 minutes, VP Engineering notified within 20 minutes.
- **SEV3 – Medium**: Non-critical service degradation affecting <30% of customers, single non-ledger microservice outage with working fallback, or elevated latency above SLA threshold but below halt. Triggers: primary on-call paged, IM notified within 30 minutes, status-page update within 1 hour if customer-visible, daily standup update.
- **SEV4 – Low**: Degraded internal tooling, minor UX bug with workaround, non-customer-facing alert. Triggers: next-business-day response, ticket created, no page unless on-call agrees.
- **SEV5 – Informational / Noise**: Cosmetic issues, planned-maintenance notifications, alert misfires logged for tuning. Triggers: no page, logged for weekly alert-quality review.
Define **escalation rules**: any SEV3 unresolved after 4 hours auto-escalates to SEV2; any SEV2 unresolved after 2 hours auto-escalates to SEV1. Severity can be **downgraded** only by the incident commander with IMC notification.
Publish the taxonomy as a one-page decision tree, a Slack slash-command (`/sev`), and an integration into the alerting tool so that every alert carries a suggested severity.
4. Define Incident Roles and Staffing Model (depends on: 3)
Codify four mandatory roles for every SEV1/SEV2 incident and optional roles for SEV3, then solve the 24×7 staffing problem across 28 teams.
**Roles**
- **Incident Commander (IC)**: owns the incident end-to-end, declares severity, assigns tasks, authorizes mitigations, decides when to escalate or stand down. Never writes code during the incident.
- **Communications Lead (CL)**: owns status-page updates, internal Slack channels, account-manager briefings, and regulator notifications. Separate from the IC so the IC can focus on mitigation.
- **Scribe / Timeline Keeper**: logs every decision, action, and timestamp in the incident channel and the incident-management tool. Produces the raw timeline for the postmortem.
- **Subject-Matter Responders (SMRs)**: 1–3 engineers from the owning team(s) who diagnose and fix. For the shared PostgreSQL ledger, a dedicated DBA responder is always required.
**24×7 Staffing via a Three-Tier Follow-the-Sun Model**
- **Tier 1 – Front-line on-call**: Primary + secondary responder per team, paged first. Covers the team's own services.
- **Tier 2 – Platform / SRE on-call**: A dedicated 6-person SRE rotation covering cross-cutting infrastructure: Kubernetes, the shared PostgreSQL ledger, networking, and the two AWS regions. This tier directly addresses the 'carrying a pager for other teams' concern by absorbing infrastructure incidents.
- **Tier 3 – IMC escalation**: Engineering managers and the IMO on-call for multi-team or SEV1 incidents. Provides the IC and CL when no team-level IC is available.
**Follow-the-Sun**: If any engineering hub exists in a second timezone, use it for overnight Tier 1 coverage. If not, partner with a managed on-call service for overnight first-response triage (severity declaration + paging the correct team), reducing 3 a.m. pages for NY-based engineers.
**IC and CL pools**: Nominate at least 2 ICs and 1 CL per team (56 ICs, 28 CLs minimum). ICs are trained and certified before they rotate. For SEV1 incidents, the IC must be a certified IC from the IMC pool, not just 'whoever is around'.
**Ledger-specific rule**: Because the PostgreSQL ledger is shared, a **Ledger Duty Officer** from the SRE Tier 2 is always on the bridge for any incident touching ledger services, regardless of which team owns the failing microservice.
5. Design On-Call Rotations, Compensation, and Alert-Quality Rules (depends on: 4)
Make on-call sustainable, fairly compensated, and free of alert noise so engineers stop resisting participation.
**Rotation Design**
- 7-day rotations, one primary + one secondary per team per week. No engineer is on-call more than one week in four.
- Minimum 48-hour rest between rotations. No on-call during approved PTO.
- All 28 teams participate. Teams without current on-call get a 90-day ramp with a shadow rotation before going live.
- Tier 2 SRE rotation: 6 engineers, one week on / five weeks off, with a dedicated backup.
**Compensation Package**
- **Base on-call stipend**: $500 per week of primary on-call, $250 for secondary, paid regardless of whether pages fire.
- **Page pay**: $75 per acknowledged page outside business hours; $150 if the page leads to active incident work.
- **Time-off-in-lieu (TOIL)**: Any engineer who works >4 hours overnight (00:00–06:00 local) gets a full TOIL day. >8 hours in a single incident gets 1.5 TOIL days.
- **SEV1 bonus**: $300 flat bonus for every engineer who actively works a SEV1 incident, paid within the next pay cycle.
- **Annual on-call cap**: No engineer exceeds 13 weeks of on-call per year. Exceeding the cap triggers a mandatory team-staffing review.
- Budget estimate: ~$420K/year for stipends and page pay across 28 teams; present this to the CFO as a fraction of the $1.3M annual SLA credit cost.
**Alert-Quality Rules (the '85% noise' problem)**
- Every alert must carry: owning team, suggested severity, runbook link, and a 30-day noise score.
- **Alert budget**: each team gets a maximum of 100 actionable alerts per month. Exceeding the budget triggers a mandatory alert-tuning session with the IMO.
- **Noise threshold**: any alert that fires >10 times in 7 days with no human action is auto-flagged for suppression or tuning within 14 days.
- **Alert review cadence**: weekly 30-minute alert-quality review per team; monthly cross-team alert review in the IMC.
- **Sunset rule**: alerts with no runbook are demoted to SEV5 after 30 days and suppressed after 60 days unless a runbook is written.
- Target: reduce monthly alert volume from 3,400 to <600 actionable alerts within 6 months.
6. Build Detection, Escalation, and Communication Paths (depends on: 3, 4)
Eliminate the 22-minute median detection gap and the 40% customer-first-detection rate with layered monitoring and a single escalation spine.
**Detection Layers**
- **Synthetic transactions**: run a payment end-to-end through the full stack (API → service → ledger → confirmation) every 60 seconds from both AWS regions. Alert if latency >2× baseline or any step fails. This catches what per-service metrics miss.
- **Customer-traffic anomaly detection**: monitor API error rates, payment success rates, and latency percentiles per customer cohort. Alert on >2σ deviation.
- **SLO-based alerting**: define SLIs for the 99.95% SLA (availability, latency p99, ledger consistency). Alert when error budget burn rate exceeds threshold, before the SLA actually breaches.
- **Infrastructure health**: Kubernetes node/pod health, PostgreSQL replication lag, disk I/O, and cross-region latency.
- **Support-ticket spike detection**: if >5 customers open tickets about the same symptom within 10 minutes, auto-create a SEV3 candidate.
**Escalation Path**
- Alert fires → PagerDuty routes to Tier 1 primary → 5-min no-ack → Tier 1 secondary → 10-min no-ack → Tier 2 SRE → 15-min no-ack → IMC on-call manager → 20-min no-ack → VP Engineering auto-page.
- Any SEV1 declaration auto-pages the CTO, opens a dedicated Slack channel + Zoom bridge, and notifies the CL.
- **No incident goes unowned for >15 minutes.** If no IC is assigned by minute 15, the IMC on-call manager assumes IC role by default.
**Internal Communications**
- Dedicated Slack channels: `#inc-sev1`, `#inc-sev2`, `#inc-sev3` (auto-created per incident), plus `#inc-updates` for broadcast.
- IC posts a structured update every 15 minutes (SEV1), 30 minutes (SEV2), 1 hour (SEV3) into the incident channel.
- CL posts a summary to `#inc-updates` and notifies relevant engineering managers.
**Customer Communications**
- **Status page**: auto-updated via API. SEV1: first update within 15 minutes, then every 30 minutes until resolved. SEV2: first update within 30 minutes, then every hour. SEV3: within 1 hour if customer-visible.
- **Account managers**: CL briefs AMs via a dedicated Slack channel within 30 minutes (SEV1) or 1 hour (SEV2). AMs contact their top-20 revenue accounts directly.
- **Customer email/SMS**: for SEV1 and SEV2, automated notification to all 2,100 customers via the status-page subscription system within 30 minutes.
- **Regulator notification**: Legal/Compliance assesses within 1 hour whether NY DFS, FinCEN, or card-network notification is required. If yes, file within the regulatory deadline (typically 72 hours for NY DFS cybersecurity events). Log the decision and filing in the incident record.
**Post-resolution**: CL publishes a 'resolved' update within 30 minutes of mitigation. For SEV1/SEV2, a preliminary customer-facing RCA summary is published within 5 business days.
7. Standardize Postmortems with Tracking and Accountability (depends on: 3, 6)
Fix the '11 of 64 action items closed' problem with a mandatory, uniform, blameless postmortem process backed by engineering-manager accountability.
**When Mandatory**
- All SEV1 and SEV2 incidents: postmortem required within 5 business days.
- SEV3 incidents: postmortem required if >50 customers affected, if the incident lasted >4 hours, or if it was customer-detected.
- SEV4/SEV5: optional, but any recurring SEV4 (≥3 times in 30 days) triggers a mandatory review.
**Format (single template, enforced by tooling)**
- Incident summary (severity, duration, customers affected, revenue impact, SLA credit exposure).
- Timeline (auto-generated from scribe notes + alert timestamps).
- Detection analysis: how was it detected, why did it take X minutes, could it have been faster.
- Root cause analysis using 5-Whys or fault-tree, not blame.
- Contributing factors (process, tooling, staffing, knowledge gaps).
- Impact quantification: customers affected, transactions failed, SLA credits triggered.
- Action items: each with a **named owner**, **due date**, **priority**, and **ticket in Jira**.
- Lessons learned and what went well.
**Blameless Review Meeting**
- Held within 5 business days, facilitated by the IMO or a trained facilitator (never the IC of that incident).
- All responders, the CL, relevant engineering managers, and an IMC representative attend.
- Ground rules: focus on system and process failures, not individual mistakes. The facilitator enforces this.
- Meeting recorded; notes published to the engineering-wide wiki within 48 hours.
**Action-Item Tracking and Accountability**
- Every action item is created as a Jira ticket with a due date and a named owner.
- **Engineering managers are accountable**: action-item completion is a standing agenda item in the bi-weekly IMC meeting. Any action item >7 days overdue is escalated to the VP Engineering.
- **Completion gate**: no team may close its postmortem until 100% of its action items have Jira tickets. Postmortem is 'closed' only when all tickets are resolved.
- **Quarterly audit**: the IMO audits action-item completion rates and reports to the IMC and the executive sponsor. Target: >90% completion within 30 days of the postmortem.
- Link postmortem quality and action-item completion to team health scores and engineering-manager performance reviews.
8. Define KPIs, Dashboards, and Governance Reviews (depends on: 7)
Create a measurable feedback loop so leadership can see whether the program is working and where to intervene.
**Primary KPIs (tracked weekly, reported monthly)**
- **MTTD** (Median Time to Detect): target <5 minutes (from 22).
- **MTTA** (Median Time to Acknowledge): target <5 minutes.
- **MTTM** (Median Time to Mitigate): target <60 minutes for SEV1 (from 3 h 10 min), <4 hours for SEV2.
- **Customer-first detection rate**: target <10% (from 40%).
- **SLA compliance**: maintain 99.95%; track monthly SLA credit payouts, target <$100K/quarter (from ~$325K/quarter).
- **Alert signal-to-noise ratio**: target >80% actionable (from ~15%).
- **Alert volume**: target <600/month (from 3,400).
- **Postmortem completion rate**: 100% for SEV1/SEV2 within 5 business days.
- **Action-item completion rate**: >90% within 30 days (from ~17%).
- **On-call health**: pages per engineer per week (target <5), TOIL usage, on-call satisfaction survey score.
- **Unowned incident duration**: target 0 incidents with >15 minutes without an IC.
**Dashboards**
- Real-time operational dashboard (Grafana): current incidents, active alerts, on-call roster, SLA error-budget burn.
- Weekly leadership dashboard (auto-generated): KPI trends, open action items, alert-noise report, on-call load distribution.
- Quarterly IMC scorecard per team.
**Review Cadence**
- **Weekly**: IMO publishes KPI snapshot to `#inc-updates`.
- **Bi-weekly IMC**: review open incidents, overdue action items, alert-quality exceptions, and on-call load.
- **Monthly executive review**: CTO presents KPI trends, SLA credit cost, and risk register to the CEO.
- **Quarterly incident-management review**: deep-dive into trends, training gaps, tooling needs, and process improvements. Output fed into the next quarter's roadmap.
9. Consolidate Tooling and Build the Incident Management Platform (depends on: 2, 3)
Replace six alerting tools and ad-hoc status-page updates with a single, integrated incident management stack.
**Target Tool Architecture**
- **Single alerting and on-call platform** (e.g., PagerDuty or Opsgenie): ingest all alerts, apply severity routing, manage on-call schedules, handle escalations, and send pages. Retire the other five tools within 6 months.
- **Observability consolidation**: standardize on one APM/metrics stack (e.g., Datadog or Grafana Cloud) for all 180 Kubernetes services across both AWS regions. Ensure the shared PostgreSQL ledger has dedicated dashboards.
- **Status page**: a dedicated, branded status page (e.g., Statuspage.io or Instatus) with API integration for auto-updates. Subscribe all 2,100 customers.
- **Incident coordination tool**: integrate incident-management workflows into Slack (auto-create channels, invite responders, post templates) and a dedicated incident record system (e.g., Jira Service Management, incident.io, or Rootly) for timelines, postmortems, and action-item tracking.
- **Runbook repository**: a central wiki (Confluence or Notion) with mandatory runbooks for every alert. No alert goes live without a linked runbook.
**Implementation Tasks**
- Migrate all 28 teams' alert rules into the single platform in three waves (highest-volume teams first).
- Build the severity-based routing rules and escalation policies per S3 and S6.
- Automate status-page updates triggered by severity declaration.
- Build the synthetic-transaction monitor and SLO-based alerting per S6.
- Integrate Jira for automatic action-item ticket creation from postmortems.
- Decommission legacy tools only after all teams have completed training on the new stack.
- Budget: allocate $150K–$250K/year for licensing, plus engineering time for migration.
10. Prepare for the SOC 2 Type II Audit (depends on: 7, 8, 9)
Ensure the incident management process produces the evidence the auditor will need, well before the audit window opens in eight months.
**SOC 2 Requirements to Address (CC7.3, CC7.4, CC7.5)**
- Documented incident response procedures (the severity taxonomy, role definitions, communication templates).
- Evidence of incident detection, response, and recovery for every SEV1/SEV2 incident during the audit period.
- Postmortem records with action-item tracking.
- On-call schedules, training records, and escalation evidence.
- Status-page update logs and customer notification records.
- Regulator notification logs (if any).
**Preparation Tasks**
- The IMO maintains a **SOC 2 evidence folder**: every incident record, postmortem, action-item ticket, status-page update, and training completion certificate is stored and indexed.
- Conduct a **mock SOC 2 audit** at month 5: an internal or external auditor reviews the incident management process end-to-end and identifies gaps.
- Remediate mock-audit findings before month 7.
- Ensure the incident management tool retains all records for at least 12 months (the SOC 2 Type II observation window).
- Document the **chain of custody** for incident records: who accessed, modified, or closed each record.
- Prepare a **narrative document** describing the incident management process, roles, and controls for the auditor.
- Coordinate with the compliance lead to align incident management evidence with the broader SOC 2 scope (access controls, change management, etc.).
11. Design and Deliver Training, Runbooks, and Change Management (depends on: 4, 5, 9)
Equip all 260 engineers, 28 team leads, account managers, and support staff with the knowledge and muscle memory to execute the new process.
**Training Tracks**
- **All 260 engineers** (2-hour session): severity taxonomy, how to acknowledge a page, how to join an incident bridge, how to hand off to an IC, and how to write a postmortem contribution. Delivered in team-level sessions over 4 weeks.
- **IC pool (56+ engineers)** (8-hour certification): incident command techniques, severity declaration, escalation decision-making, bridge facilitation, and blameless postmortem facilitation. Includes two tabletop exercises. Certification valid for 12 months, renewed annually.
- **CL pool (28+ staff)** (4-hour session): status-page writing, customer communication templates, regulator notification triggers, and AM briefing protocol.
- **Account managers and support staff** (1-hour session): how to read the status page, how to escalate a customer report into an incident, and what information to collect.
- **SRE Tier 2** (16-hour onboarding): Kubernetes and PostgreSQL ledger deep-dive, cross-region failover runbooks, and escalation authority.
**Runbooks**
- Every alert must have a runbook before it is routed to on-call. The IMO provides a runbook template and audits compliance weekly.
- Priority runbooks to write first: shared PostgreSQL ledger failover, Kubernetes cluster degradation, payment-processing pipeline failure, cross-region failover, and ledger data-integrity check.
- Runbooks are peer-reviewed and version-controlled.
**Change Management for Adoption**
- Address the 'carrying a pager for other teams' concern directly: publish an FAQ explaining the three-tier model, the SRE Tier 2 absorbing cross-team infrastructure, the compensation package, and the TOIL policy.
- Run **office hours** weekly for the first 8 weeks where any engineer can ask questions or raise concerns.
- Identify **team champions**: one engineer per team who volunteers as an early adopter and peer mentor.
- Publish a **weekly 'incident management newsletter'** during rollout: what changed, what improved, KPI trends, and success stories.
- Make on-call participation a documented expectation in job descriptions and performance reviews for production-facing roles.
12. Execute Phased Rollout, Tabletop Exercises, and Continuous Improvement (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11)
Introduce the process in three waves so teams are not overwhelmed, then validate with exercises and iterate continuously.
**Phase 1 – Weeks 1–6: Foundation**
- Publish the severity taxonomy, role definitions, and communication protocols (S3, S4, S6).
- Launch the single alerting platform for the 12 teams already on-call; begin migration for the other 16.
- Activate the SRE Tier 2 rotation for the shared PostgreSQL ledger and cross-cutting infrastructure.
- Deploy the status page and test the API integration.
- Begin IC and CL training (first cohort of 20 ICs, 10 CLs).
- Publish the on-call compensation package; HR integrates stipends into payroll.
- Write the top 10 priority runbooks.
**Phase 2 – Weeks 7–14: Expansion**
- All 28 teams live on the single alerting platform; legacy tools in read-only mode.
- All 28 teams on the on-call rotation schedule (the 16 new teams in shadow mode for the first 4 weeks).
- Second and third IC/CL training cohorts completed.
- First **tabletop exercise**: simulate a SEV1 ledger corruption scenario with all roles, test the escalation path, status-page updates, and AM briefings. Debrief and fix gaps.
- Postmortem template and Jira integration live; all new incidents use the standard process.
- Alert-tuning sprint: each team reduces its alert volume by 50%.
**Phase 3 – Weeks 15–24: Optimization**
- All teams fully live; legacy alerting tools decommissioned.
- Second **tabletop exercise**: simulate a SEV1 cross-region failure with regulator notification.
- First quarterly IMC review with full KPI dashboard.
- Mock SOC 2 audit (month 5) and remediation.
- Retrospective on the rollout: survey all 260 engineers for feedback, adjust compensation or rotation rules if needed.
- Establish the **continuous improvement cadence**: quarterly process review, annual severity-taxonomy review, and annual on-call compensation benchmarking.
**Ongoing Governance**
- The IMC owns the process document and approves changes.
- The IMO tracks all KPIs and reports to the CTO monthly.
- Any process change requires IMC approval and a 2-week notice period before enforcement.
- Annual external benchmarking against peer B2B payments platforms.
13. Establish Ongoing Governance, Annual Review, and Audit Readiness Cycle (depends on: 12)
Embed incident management as a permanent organizational capability, not a one-time project.
- **Annual process review**: the IMC reviews the severity taxonomy, role definitions, on-call structure, and compensation against industry benchmarks and internal KPIs. Update as needed.
- **Bi-annual tabletop exercises**: one SEV1 infrastructure scenario, one SEV1 data-breach/regulator scenario. Rotate the IC and CL assignments so everyone gets practice.
- **Quarterly alert-quality audit**: the IMO reviews alert volumes, noise rates, and runbook coverage across all 28 teams.
- **On-call health survey**: quarterly anonymous survey measuring burnout, fairness, and compensation satisfaction. Results reviewed by the IMC.
- **SOC 2 readiness cycle**: begin evidence collection immediately after each audit ends. The IMO maintains a rolling evidence folder. Mock audit at month 5 of every 12-month audit cycle.
- **Postmortem maturity tracking**: track the action-item completion rate monthly. If it drops below 80%, the VP Engineering intervenes.
- **Incident management maturity model**: adopt a 5-level maturity model (ad-hoc → defined → managed → optimized → predictive). Assess annually. Target: Level 3 within 12 months, Level 4 within 24 months.
- **Budget review**: annually review on-call compensation, tooling costs, and training budget against the reduction in SLA credits and incident frequency.
Previous Proposal 4 (ID: 31fbbc6c-8c88-4abb-9c6d-b67706d9da75, Agent: grok4.6_initial_4, LLM: xai/grok-4.6):
Estimated Complexity: high
Success Metrics: - Median time to detect customer-impacting incidents ≤ 5 minutes within 6 months of go-live.
- Share of SEV-1/SEV-2 incidents first detected by customers ≤ 5% (from 40%).
- Median time to mitigate SEV-1/SEV-2 ≤ 45 minutes (from 3 h 10 min).
- Named Incident Commander assigned within 5 minutes for ≥ 95% of SEV-1/SEV-2.
- First status-page update within policy time for ≥ 95% of SEV-1/SEV-2.
- SLA credits down ≥ 80% versus the trailing $1.3M within 12 months.
- Paging volume ≤ 500 per month and noise ≤ 15% (from 3,400 and 85%).
- 100% of production services have a named owning team and a paging policy.
- Postmortems filed within 5 business days for 100% of SEV-1/SEV-2; action-item close rate ≥ 80% within 30 days.
- 24×7 IC and critical-path coverage with zero unfilled shifts per quarter.
- Paid on-call live for every rotation before that rotation pages humans.
- SOC 2 Type II incident-response controls evidenced for ≥ 5 months before the auditor's report.
- On-call pulse: ≥ 70% of engineers agree rotations are fair and limited to their services.
Steps (30):
1. Secure executive mandate and budget
Get a written CEO/CTO mandate that incident command is a company process, not a team hobby.
The mandate must state that **paid on-call** is required for production ownership. "No pager for other teams' code" is solved by named ownership, not by refusing coverage.
- Approve budget for tooling, stipends, training, and a dedicated program lead for six months.
- Name an executive sponsor (CTO or VP Engineering) who will chair the weekly incident review.
- Tie the clock to SOC 2 Type II: the process must be live in about 10 weeks so ~6 months of evidence remain.
- Commit that the CEO will hear about outages from this process, not from customers.
2. Form the working group and decision rights (depends on: 1)
Stand up a small group that can decide. Do not form a 28-team committee.
**Core seats:** SRE/platform lead, payments/ledger engineering manager, support lead, legal/compliance, HR, one rotating team EM, and a program manager.
- Meet twice a week for 10 weeks, then weekly.
- RACI: the group proposes; the sponsor decides in 48 hours; teams implement.
- Publish one Slack channel and one source-of-truth doc on day one.
- Time-box design to four weeks. Ship v1 rather than wait for consensus.
3. Inventory services, owners, and on-call gaps (depends on: 1)
Build a living catalog of all ~180 services: owning team, criticality, current on-call, alert sources, and runbook link.
Walk the last 31 customer-impacting incidents and the two events where **nobody was in charge**. Record who detected, who led, time to mitigate, and which alerts fired.
- Tag each service as critical-path, customer-visible, or internal.
- List the 16 teams with no on-call and every orphan service with no owner.
- Map all six alerting tools and the 3,400 monthly alerts onto services.
- Flag the shared PostgreSQL ledger and two-region failover as named special cases.
4. Map regulatory and contractual notification duties (depends on: 1)
Legal and compliance list every duty an incident can trigger. Do not invent clocks that violate a contract.
Cover SOC 2 CC7, customer MSA/SLA credit terms, money-transmitter rules, NYDFS 23 NYCRR 500 if applicable, PCI if in scope, and breach clocks.
- Extract **notification timings** from the largest customer contracts (status page, named AM, written notice).
- Define when legal, regulators, insurers, or the board must be told.
- Feed these clocks into severity triggers and the communications playbook.
5. Approve paid on-call and incident pay (depends on: 1, 4)
Unpaid on-call is why 16 teams refuse the pager and why nights are uncovered. Fix the money before asking for coverage.
HR, legal, and finance design a New York–compliant package: weekly stipend for primary and secondary, extra stipend for company IC and comms, and **after-hours incident pay** or comp time.
- Treat exempt vs non-exempt staff explicitly under NY wage-hour rules.
- Put stipend in the next pay cycle after policy publish, not "later."
- Cap consecutive night weeks. Fund hiring if a team cannot rotate fairly (minimum six people for 24x7 primary plus secondary).
- Publish the package before any new rotation starts. This is the main answer to pager pushback.
6. Ratify four severity levels and their triggers (depends on: 3, 4)
Adopt a business-impact scale. Engineers do not invent severity in the moment.
**SEV-1:** material payments failure; ledger down or inconsistent; security or customer-data incident; both regions impaired; or many customers already in SLA-credit territory.
**SEV-2:** degraded payments or a contracted feature down for multiple customers; SLA at risk.
**SEV-3:** narrow or single-customer impact with a workaround; no fleet-wide SLA risk.
**SEV-4:** no customer impact; ticket only.
- SEV-1 pages IC, comms, scribe, owning SMEs, and an exec; war room in 5 minutes; status page in 10; AM outreach in 20.
- SEV-2 pages IC and owning SMEs; comms may be the IC; status page in 15 minutes; updates every 30 minutes.
- SEV-3 pages the owning team only; customer notice only if that customer is affected.
- Anyone may declare. Only the IC may downgrade. When unsure, start high.
7. Define incident roles and their authority (depends on: 6)
Four roles. Separate coordination from debugging so "who is in charge" cannot stall for an hour again.
- **Incident Commander:** owns severity, the room, the clock, and the next action. Does not write code. May page anyone, freeze deploys, and invoke failover. Staffed from a company-wide trained pool, not from the failing team.
- **Communications Lead:** status page, customers, AMs, execs, regulators. Speaks only from IC-approved facts.
- **Scribe:** timeline in the incident tool. Required for SEV-1 and SEV-2.
- **SME responders:** the owning team's on-call. They mitigate. They do not run the room.
Publish a one-page authority card. The IC stays in charge even if a VP joins.
8. Design 24x7 coverage without 28 night rotations (depends on: 3, 5, 7)
Do not put 28 teams on 24x7. That is what engineers are rejecting.
Use **three layers** so people page for their own code, plus a trained commander.
- Layer A — company IC and SEV-1 comms: 24x7; about 24 trained people; week-long primary and secondary.
- Layer B — critical-path team on-call (ledger, payments processing, auth, API edge, platform/Kubernetes, data stores): 24x7 primary plus secondary.
- Layer C — all other teams: business-hours on-call; after hours the IC pages the team EM, who has a written escalation list.
Platform on-call is the safety net for unknown-owner pages, never the permanent owner. Every service must have a named team within 60 days or be scheduled to shut off.
9. Set rotation, handoff, and load rules (depends on: 8)
Write mechanical rules so rotations are fair and load is visible.
Primary week, then secondary week, then at least two weeks off. No one holds primary on two rotations at once.
- Handoff is a 30-minute overlap covering open incidents, silenced alerts, and upcoming changes.
- Page-load SLO: p50 ≤ 4 pages per 12-hour night shift; p95 ≤ 10. A breach opens an alert-quality action.
- Require a shadow week before a first IC shift or a first critical-path rotation.
- Swaps live in the paging tool. Managers own coverage gaps, not the last person on the roster.
10. Write detection and escalation paths (depends on: 6, 9)
Customers currently detect 40% of incidents and median time to detect is 22 minutes. That is the first failure mode.
Detection path: synthetic full-payment probes in both regions, SLO burn-rate alerts, support-to-incident intake, and one customer callback path that can create a SEV.
- A page must be acked in **5 minutes** or it auto-escalates to secondary, then IC, then the EM, then the VP.
- Support may declare SEV-2 or higher without engineering permission.
- If ownership is unclear for 10 minutes, the IC keeps the incident and assigns a temporary owner. Never wait.
- An exec bridge auto-opens for every SEV-1 at T+15 minutes.
11. Write internal, customer, and regulator communications (depends on: 4, 6, 7)
Stop "whoever is around" from writing the status page. Comms follow the clock, not convenience.
Timings from declaration:
- Internal war room: immediate. Exec summary for SEV-1/2 at 15 minutes, then every 30 minutes.
- **Public status page:** SEV-1 in 10 minutes, SEV-2 in 15. Updates at least every 30 minutes until resolve. Templates only. No speculation.
- Account managers get an affected-customer list and a script at T+20 minutes for SEV-1/2.
- Resolve notice and credit assessment within one business day.
- Comms pages legal on SEV-1 security, ledger integrity, or any outage that will breach contractual notice. Legal owns outbound regulatory letters; the IC owns facts.
12. Standardize blameless postmortems and action tracking (depends on: 6)
A written postmortem is mandatory for every SEV-1 and SEV-2 within 5 business days. SEV-3 if the IC or EM requests it.
Use one template: timeline, customer impact (volume, duration, credits), detection gap, what went well, what did not, process-focused five whys, and numbered actions with owner and due date.
- Review is **blameless** and scheduled. The IC attends. The exec sponsor reads every SEV-1.
- Actions live in one tracker, not in the doc. No action without an owner and a date. Default due date 14 days; 30 days max unless architecture work with a milestone.
- Close rate is a published metric. The old 11-of-64 pattern is a process failure.
13. Set alert quality rules that make paging acceptable (depends on: 3, 6)
3,400 alerts a month and 85% noise is why on-call feels like punishment. Pages are a product with a quality bar.
A page (not a ticket) must map to a customer-facing SLO or a hard dependency of one. It must have an owner team, a runbook link, and a default severity. It must be actionable at 3 a.m. by the person who is paged.
- Ban parallel paging from six tools. One paging policy: symptom-based; burn-rate preferred over raw thresholds.
- Every team gets a monthly noise budget. Exceeding it is a sprint task, not heroics.
- A human may silence a flapping alert only with a linked ticket.
14. Publish Incident Management Policy v1 (depends on: 5, 6, 7, 8, 9, 10, 11, 12, 13)
Collapse the design into a short policy people will open during an outage.
Ten pages or fewer, plus one-page cards for severity, roles, and comms timings. Host it where the incident tool can link it.
- Include the compensation summary and the rule: you are **not on-call for other teams' services**.
- Version it. v1 is mandatory from the pilot start date.
- Legal, HR, and the exec sponsor sign. Announce in all-hands, not only in Slack.
15. Implement a single incident command tool (depends on: 7, 10, 11)
Put one tool in the path that creates the room, pages roles from severity, records the timeline, and prompts status-page updates.
Requirements: Slack (or equivalent) incident bot, severity in one click, role assignment, stakeholder groups, and timeline export for postmortems and auditors.
- Integrate with the pager so IC, comms, and SME pages are automatic.
- Retain artifacts at least one year for SOC 2.
- Ad-hoc Zoom/Slack threads are no longer the system of record.
16. Consolidate six alerting tools onto one pager (depends on: 9, 13)
Pick one paging product. Connect existing monitors to it. Migrate **pages** first, tickets second.
- Inventory every page-producing rule. Delete or downgrade the noisy majority in S23.
- Route by service label → owning team schedule → the escalation policy from S10.
- IC and comms schedules live in the same product.
- Set a hard date after which pages outside the chosen tool are not valid on-call obligations.
17. Operationalize the status page and AM path (depends on: 11, 15)
Put the status page behind the Comms Lead role. Use templates for investigating, identified, mitigating, and resolved.
Subscribe AMs and customers whose contracts require it. Generate the affected-customer list from the incident (tenant, region, payment method).
- Dry-run a SEV-2 update before the pilot goes live.
- Record every public update in the incident timeline for the audit.
- Define partial vs full-outage wording so impact cannot be understated.
18. Stand up one postmortem repo and action board (depends on: 12)
Create the template, the filing location, and one Jira/Linear board with states: open, in progress, blocked, done, and won't-do (with exec reason).
Wire the incident tool so a SEV-1/2 automatically opens a draft postmortem and action tickets.
- Each week the program manager reports actions past due to the exec sponsor.
- If it is not in the board, it does not exist.
19. Assign owners and write critical-path runbooks (depends on: 3, 9)
Close the ownership gaps that cause "pager for other people's code."
Every production service gets a team in the catalog. Unowned services get an owner in 30 days or a decommission date.
- Write runbooks for ledger Postgres, regional failover, payments API, auth, and the Kubernetes control plane: symptoms, dashboards, mitigate vs escalate, customer impact.
- Link runbooks from alerts. If there is no runbook, the alert cannot page at night unless the EM accepts the gap in writing.
20. Detect payments failures before customers do (depends on: 3, 13)
Build or finish end-to-end synthetics: create payment, ledger write, webhook, both regions, both critical payment methods.
Alert on SLO burn, not on a single 500. Page SEV-2 or SEV-1 from these probes. This is the fastest lever on the 22-minute MTTD and the 40% customer-detected rate.
- Add ledger lag, replication, disk, and failover-readiness as first-class pages to the ledger team.
- Every postmortem asks first: **why did a customer see this first?**
21. Train the first cadre of ICs, comms leads, and scribes (depends on: 7, 14, 15)
Train about 24 ICs and 12 comms leads before the pilot. Classroom plus a recorded shadow of a simulated SEV-1.
Curriculum: severity, authority card, tool, comms timings, when to call legal, how to run a room of 20, how to hand off at 2 a.m.
- Certification: pass a tabletop. No certificate, no rotation.
- Recertify yearly and after any SEV-1 where process failed.
- Managers of ICs protect calendar time. This is part of the job.
22. Pilot on the payments critical path for six weeks (depends on: 5, 14, 15, 16, 17, 18, 19, 21)
Go live with policy, tool, paid rotations, and IC coverage for ledger, payments, API, platform, and support intake.
Keep old paths as backup for one week, then cut over. Real incidents use the new process only.
- Staff the program lead in every SEV-2+ as coach, not as secret IC.
- Collect friction daily. Fix tooling and wording in 48 hours.
- Expansion gate: named IC in under 5 minutes, first status update on time, no unpaid pages, postmortem filed.
23. Cut alert noise with a forced burn-down (depends on: 13, 16)
Give every team a numbered list of their noisiest alerts. Move noise from 85% to **under 15%**, and monthly pages from 3,400 toward 500.
Each sprint, critical-path teams must delete, debounce, or convert to ticket a fixed quota. Platform provides burn-rate and grouping libraries.
- Publish a weekly noise leaderboard. Shame systems, not people.
- After eight weeks, any page without a runbook or with >30% false pages in 14 days is auto-downgraded until fixed.
24. Roll remaining teams onto the model by risk (depends on: 22)
After the pilot gate, add teams in waves of four to five every two weeks. Highest customer-impact first.
Layer C teams get business-hours schedules and the EM night list. Do not surprise anyone with a pager.
- Each wave: ownership confirmed, alerts routed, runbooks for paging alerts, paid rotation in HR, one tabletop.
- Finish all 28 teams at least five months before the SOC 2 report date so the observation window covers the company.
- Orphan services still unowned at wave end are escalated to the sponsor for shutdown or reassignment.
25. Run tabletops and multi-region game days (depends on: 21, 22)
Schedule a monthly tabletop: SEV-1 ledger, SEV-1 region loss, SEV-2 degraded payments, customer-detected incident, and a "who is in charge" chaos drill.
Quarterly game day: fail a region or a ledger replica in staging or a controlled production drill.
- Include support, AMs, legal, and an exec. Process fails if only engineers show up.
- Capture actions on the same board as real postmortems.
- Use results as SOC 2 evidence that IR is tested.
26. Resolve ownership fights and pager culture (depends on: 5, 8, 22)
Treat pushback as design input, not defiance. Repeat the contract in office hours: you carry a pager for **your** services; IC is coordination; nights are paid; noise is a defect.
- EMs who cannot staff a fair rotation get headcount or have services reassigned. Do not run two-person 24x7.
- Publicly close the two historical "nobody in charge" incidents with what would be different now.
- Pulse-survey on-call at 60 and 120 days. If load or fairness is red, stop expansion until fixed.
27. Launch metrics, weekly review, and error budgets (depends on: 15, 22)
The CEO email exists because there was no operating rhythm. Stand up a dashboard and use it.
Track incident count by SEV, MTTD, MTTA, MTTM, customer-detected %, time-to-IC, time-to-status, pages per shift, noise %, SLA credits, postmortem SLA, and action close rate.
- **Weekly 30-minute review** chaired by the exec sponsor. Last week's incidents, open actions, alert SLO breaches. Mandatory for EMs of teams that had SEV-2+.
- Error budgets on critical SLOs: burn too fast and the team stops features and pays down reliability.
- Monthly report to the CEO so customer-first news does not recur.
28. Collect SOC 2 Type II evidence as a side effect (depends on: 4, 14, 22, 27)
Do not create a parallel "audit process." The real process is the evidence.
From pilot start, retain policy versions, on-call schedules, incident records, status-page history, postmortems, action tickets, training attendance, tabletop notes, and access reviews of the incident and pager tools.
- Map each artifact to SOC 2 CC7.2–CC7.5. Compliance owns the matrix; engineering owns completeness.
- Internal audit or a consultant runs readiness reviews at week 12 and week 20, not the week before the auditor.
29. Inspect and adapt the process after 90 days (depends on: 22, 27)
After 90 days live, revise v1 using data, not opinions.
Look at MTTD/MTTM, uncovered shifts, IC burnout, missed status updates, action close rate, and survey results.
- Change severity definitions if teams inflate or deflate them.
- Adjust Layer B vs C membership from real pages.
- Cut steps nobody follows. Add only what the last 90 days proved missing.
- Freeze v2 as the SOC 2 described process unless a defect is found.
30. Rehearse the audit and freeze the observation window (depends on: 24, 28, 29)
Six to eight weeks before the auditor, run a mock interview with an IC, a random engineer, support, and compliance.
Show one SEV-1 and one SEV-2 end-to-end: detect, page, command, comms, postmortem, closed actions.
- Fix evidence gaps immediately.
- Freeze process wording for the rest of the observation window; log exceptions.
- Brief the CEO and customer success with the metrics so the "outages we hear about from clients" story is retired.
Previous Proposal 5 (ID: 53260d57-1d36-433f-83b4-7205e06869b2, Agent: deepseek-v4-pro_initial_5, LLM: deepseek/deepseek-v4-pro):
Estimated Complexity: high
Success Metrics: - Customer-detected incidents decrease from 40% to less than 15% within six months.
- Median time to detect (MTTD) is under 5 minutes for SEV1 and SEV2 incidents.
- Median time to mitigate (MTTM) is under 60 minutes for SEV1 and under 2 hours for SEV2.
- Alert noise decreases from 85% to below 10% within three months.
- 100% of SEV1 and SEV2 incidents have a completed blameless postmortem within 5 business days.
- 100% of postmortem action items are tracked with an owner and due date; 90% are completed on time.
- 24x7 on-call coverage achieved across all 28 teams with no unpaid on-call.
- 99% of pages are acknowledged within 5 minutes.
- SLA credits paid reduce by at least 50% over the next 12 months.
- SOC 2 readiness: all incident response controls are documented, tested, and evidence is produced by month 7.
Steps (14):
1. Baseline current incident response and align stakeholders
Collect data from the last 12 months of incidents, all six alert tools, on-call practices, and team interviews. Identify gaps against the target incident management process and secure executive sponsorship.
- Gather incident timeline, detection source, mitigation time, customer impact, SLA credit, and postmortem status for all 31 incidents.
- Survey 28 teams on on-call burden, alert quality, and operational pain.
- Map current tools, escalation paths, and communication workflows.
- Create baseline metrics and a stakeholder map with executive sponsor and audit owner.
2. Define severity levels and response triggers (depends on: 1)
Define a four-level severity scale with objective business-impact criteria so any engineer can classify an incident consistently.
- SEV1: widespread transaction processing outage, data breach, security incident, or severe SLA breach; triggers full incident command, executive notification, and a 5-minute status page update.
- SEV2: major feature outage, significant degradation without workaround, or customer financial risk; triggers incident commander, full communications role, and status page updates.
- SEV3: partial impairment with workaround or limited customer impact; triggers on-call response, internal communication, and optional status page update.
- SEV4: minor or internal issue, no customer impact; handled during business hours through ticketing.
- Include an escalation matrix showing who can declare, downgrade, and invoke regulatory or legal involvement.
3. Define incident roles and decision authority (depends on: 1, 2)
Define incident roles, responsibilities, and decision authority using RACI to remove ambiguity about who is in charge.
- Incident Commander: owns the incident, declares severity, and coordinates resolution.
- Communications Lead: owns internal and external messaging, status page updates, and account manager notifications.
- Scribe: maintains timeline, incident log, and postmortem notes.
- Subject-matter responders: diagnose and fix the incident; may come from multiple teams.
- Executive sponsor: optional for SEV1; customer liaison: handles account managers.
- Define decision rights for severity declaration, escalation, rollback, customer communications, and incident closure.
4. Design 24x7 staffing model across 28 teams (depends on: 3)
Design 24x7 coverage across 28 teams without overloading engineers. Use service-based on-call plus a central incident command pool.
- Each service or domain team assigns primary and secondary on-call for its own services.
- Create central incident commander, communications, and scribe rotations staffed from a trained incident response guild across all teams; use follow-the-sun between the two AWS regions and time zones.
- Define escalation layers: service on-call to team lead or manager to service owner to executive.
- Define handoff times, shadow shifts, and load balancing; target at most one week of on-call per engineer per month.
- Bridge the current 12-team paid on-call to 28-team paid coverage; no team remains uncovered.
5. Define on-call rotations, compensation, and alert quality rules (depends on: 4)
Define sustainable rotations, pay, and rules that eliminate noisy pages.
- Rotations: weekly or biweekly, at least one primary and one secondary, with 12-hour shifts where possible or 24-hour for low-volume services.
- Compensation: monthly on-call stipend for all on-call engineers, additional incident response bonus for after-hours work, and time off in lieu; align with market rates.
- Alert quality rules: every page must be actionable, have a runbook link, specify a service owner, include severity, and be based on SLO burn or known failure signals; no dashboard-only alerts.
- Noise budget: reject or downgrade non-actionable alerts; all pages must go to on-call only after suppression and deduplication.
- Weekly alert review removes the top noisy alerts.
6. Design detection, escalation, and alert routing (depends on: 2, 4, 5)
Define how incidents are detected, routed, and escalated so nothing waits on a human to notice.
- Consolidate the six alert tools into one alerting and paging platform with routing by service, severity, and tags.
- Detection sources: infrastructure metrics, application synthetic transactions, log-based anomalies, business transaction SLI monitoring, and customer-reported issues through support or account managers.
- Routing: alert is paged to service on-call within 30 seconds; primary must acknowledge within 5 minutes; if no ack, page secondary then on-call manager.
- Escalation timeouts: unresolved SEV1 escalates to service owner at 15 minutes and to leadership at 30 minutes; any engineer can escalate to the incident commander.
- Define customer-reported incident intake and classification in the same tool.
7. Define internal and external communication protocols (depends on: 2, 3)
Define communication channels, templates, and timing for internal, customer, and regulator audiences.
- Internal: dedicated incident Slack channel, internal status page mirror, and war room bridge for SEV1; incident commander and communications lead own these channels.
- Status page: SEV1 post within 5 minutes, updates every 30 minutes or on material change, resolution within 60 minutes of mitigation; SEV2 post within 15 minutes, updates hourly; SEV3 optional.
- Account managers: SEV1 and SEV2 notify account managers within 15 minutes with an approved customer-facing description and expected impact.
- Regulators: legal or compliance determines notification for data breaches, security incidents, funds availability issues, or regulatory reportable events; criteria and timing follow legal and regulatory requirements; communications lead coordinates.
- Use pre-approved message templates and an approval chain; no ad-hoc wording.
8. Define postmortem policy and action tracking (depends on: 3)
Define mandatory blameless postmortems and action tracking.
- Mandatory for all SEV1 and SEV2 incidents, and any SEV3 that breaches SLA or is customer-detected.
- Format: impact, timeline, root causes, contributing factors, detection and response gaps, what worked well, and action items.
- Blameless: focus on system and process causes, not individual blame; use trained facilitators.
- Ownership: each action has an owner, due date, and tracking ID in a single backlog.
- Review postmortems at the weekly incident review; track action closure; expect 100% completion.
- Complete postmortems within 5 business days for SEV1 and SEV2 incidents.
9. Define metrics, dashboards, and review cadence (depends on: 2, 3, 8)
Define metrics and review cadence to measure process health.
- Metrics: MTTD, MTTM, customer detected percentage, alert noise percentage, on-call response time, on-call load, SLA credits paid, and postmortem action completion.
- Dashboards: real-time operational dashboard for on-call engineers and management.
- Weekly incident review: review all SEV1 and SEV2 incidents, action items, and noisy alerts.
- Monthly trends with leadership; quarterly review against SLOs and audit controls.
- Success thresholds: MTTD under 5 minutes, MTTM under 60 minutes for SEV1, customer detected under 15%, and alert noise under 10%.
10. Configure incident tooling and integrations (depends on: 5, 6, 7, 8, 9)
Implement and integrate the tools that automate the defined process.
- Aggregate alerts from the existing six tools into PagerDuty, Opsgenie, or a similar platform.
- Configure on-call schedules, escalation policies, and paging targeted at service owners.
- Integrate status page API for automated or one-click updates.
- Add Slack commands to declare incidents, start war rooms, assign roles, and post status updates.
- Integrate runbook and service catalog access; create postmortem templates in Jira or Notion with action item tracking.
- Ensure audit trails and role assignments are logged for SOC 2.
11. Pilot with 2-3 volunteer teams and iterate (depends on: 10)
Run a controlled pilot before full rollout to validate and refine the process.
- Select 2-3 volunteer teams with representative services and on-call patterns.
- Run the new severity, roles, on-call, alerting, and communication process for 2 weeks.
- Track metrics and gather feedback from on-call engineers, incident commanders, and communications leads.
- Iterate severity thresholds, alert rules, templates, and runbooks based on findings.
- Exit criteria: no SEV1 without a declared incident commander, alert noise below target, and positive on-call survey results.
12. Train and certify all 28 teams (depends on: 11)
Prepare all 28 teams so the process is not a paper exercise.
- Create role-specific training: all engineers, incident commanders, communications leads, scribes, subject-matter responders, and managers.
- Deliver interactive workshops and online modules.
- Run at least two game-day drills per quarter simulating SEV1 and SEV2 incidents, testing escalation, communication, and postmortem.
- Certify incident commanders and communications leads; only certified staff can hold those roles in rotations.
- Include scenario-specific training for account managers and legal or regulatory handling.
13. Full rollout and enforce process (depends on: 12)
Move every team onto the new process and retire old habits.
- Migrate all 28 teams into consolidated on-call schedules and alert routing.
- Retire old alerting tools or methods that bypass the process.
- Activate status page templates and postmortem action tracking globally.
- Monitor the first 30 days with a daily operational review and quick fixes.
- Communicate rollout status to leadership and all engineers; provide office hours and support.
14. SOC 2 readiness and continuous improvement (depends on: 13)
Map the process to SOC 2 and keep improving after go-live.
- Map each process component, including severity, roles, detection, communications, postmortems, and metrics, to SOC 2 control requirements.
- Produce runbooks, evidence of training, incident logs, postmortems, and action item tracking.
- Conduct an internal readiness assessment at month 6 and remediate gaps before the audit at month 8.
- Maintain weekly and monthly metrics reviews to drive continuous improvement.
- Refresh on-call compensation and alert quality rules based on operational data.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Create the mandate, ownership, and funding", "description": "Launch incident management as a company operating program within 48 hours. Give one leader authority to set standards across all 28 teams.\n\n- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable program owner.\n- Form a small steering group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.\n- Fund two to three implementation staff, paging and incident tooling, observability work, training, exercises, and on-call compensation.\n- Reserve 10%–15% of engineering capacity for alert remediation, runbooks, and incident actions.\n- Authorize incident commanders to freeze deployments, order rollback, disable features, shift traffic, and invoke continuity plans.\n- Preserve financial controls. Incident commanders may coordinate ledger recovery but may not bypass dual approval, privileged-access controls, or reconciliation.\n- Set the delivery target: critical controls operational within 60 days, enterprise rollout within 120 days, and a mock audit in month 7.", "dependencies": []}, {"step_id": "S2", "title": "Install an interim process in seven days", "description": "Do not wait for new tools or the final policy. Put a minimum viable incident process into operation immediately and start collecting evidence.\n\n- Publish a one-page provisional severity guide, declaration procedure, role card, and communications schedule.\n- Establish one monitored declaration path through chat, telephone, and the existing paging tools.\n- Create a standard incident channel, bridge, incident document, and naming convention.\n- Staff temporary primary and backup incident commanders 24×7 from the existing on-call teams and engineering leadership.\n- Compensate interim duties retroactively under the final compensation policy.\n- Require a named incident commander within 10 minutes for every suspected major incident.\n- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.\n- Hold a daily 15-minute control review until the permanent process is live.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Build the baseline, ownership catalog, and risk map", "description": "Establish the facts behind the current failures and assign every production component an owner. Use the resulting catalog as the source for routing, escalation, and audit evidence.\n\n- Reconstruct all 31 customer-impacting incidents, including first impact, detection source, declaration, command assignment, mitigation, resolution, customer communications, and credits.\n- Analyze the two incidents with no clear leader and every case detected first by customers.\n- Inventory all 180 services, Kubernetes clusters, AWS accounts, databases, queues, regional dependencies, payment processors, and banking partners.\n- Assign one accountable team, engineering manager, product owner, primary escalation, secondary escalation, dashboard, and runbook to each service.\n- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 based on financial integrity, customer impact, contractual exposure, and dependency centrality.\n- Map critical customer journeys to their application, PostgreSQL, Kubernetes, regional, and third-party dependencies.\n- Inventory the six alert sources and all 3,400 monthly alerts by owner, volume, actionability, and duplication.\n- Interview representatives from all 28 teams and baseline on-call sentiment, fatigue, and objections.\n- Give orphan services an owner or decommissioning decision within 30 days.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Design compliance and evidence controls from day one", "description": "Map the operating process to audit, legal, contractual, and record-retention requirements before finalizing it. Confirm the expected SOC 2 observation period with the auditor immediately.\n\n- Map controls to applicable SOC 2 criteria for monitoring, incident identification, response, recovery, communications, corrective action, access, and availability.\n- Define evidence required for declarations, pages, acknowledgements, role assignments, decisions, status updates, postmortems, actions, training, drills, and exceptions.\n- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.\n- Version and approve the Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, and Postmortem Standard.\n- Record control exceptions with an owner, compensating control, approval, and expiry date.\n- Start monthly evidence sampling immediately rather than reconstructing evidence before the audit.", "dependencies": ["S1"]}, {"step_id": "S5", "title": "Adopt one severity and lifecycle standard", "description": "Use four impact-based severity levels across operational, security, data, and third-party incidents. Start at the highest credible severity when facts are uncertain, then downgrade with recorded evidence.\n\n- **SEV0 — financial or security crisis:** suspected ledger corruption, unauthorized or duplicated funds movement, material data compromise, both-region loss, or a decision to suspend payment processing. Page all roles immediately; engage executives, Security, Legal, Compliance, and Risk; assess regulatory duties within one hour; use distinct role holders; require a postmortem.\n- **SEV1 — critical availability event:** a core payment journey is unavailable, payment failures exceed an initial 10% guardrail for five minutes, regional loss has impaired failover, a settlement deadline is at imminent risk, or rapid error-budget burn makes a material SLA breach likely. Assign all roles; notify internal stakeholders within 10 minutes; publish customer status within 15 minutes; update every 30 minutes; require a postmortem.\n- **SEV2 — major bounded event:** approximately 1%–10% of payment attempts fail, a material customer subset or critical customer is down, degradation has a workaround, or contractual impact is likely. Assign an incident commander and responders; add communications and scribe roles for customer impact; publish status within 30 minutes; update every 60 minutes; require a postmortem for customer-visible events.\n- **SEV3 — limited event:** localized impact, a safe workaround, and no financial-integrity, security, regulatory, or material contractual risk. The owning team leads; page only if immediate action is necessary; use a ticket otherwise.\n- Treat the percentage thresholds as declaration guardrails, not reasons to under-classify integrity, settlement, security, or strategic-customer risk.\n- Permit any employee to declare an incident. Only the incident commander may lower severity, with the rationale logged.\n- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.\n- Define mitigation as the end of active customer harm. Declare resolution only after stability, backlog recovery, and required ledger reconciliation.", "dependencies": ["S2", "S3", "S4"]}, {"step_id": "S6", "title": "Define roles, authority, and handoffs", "description": "Separate command, communication, recordkeeping, and technical repair. One named person must hold command at every moment of a major incident.\n\n- **Incident Commander:** owns severity, priorities, role assignment, escalation, decision cadence, mitigation coordination, and closure. The commander does not act as the primary technical operator.\n- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language. Legal or Compliance retains ownership of regulatory submissions.\n- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, and status changes.\n- **Subject-Matter Responders:** diagnose and mitigate services for which they have accepted ownership, access, training, and runbooks.\n- **Executive Duty Officer:** removes organizational obstacles and approves exceptional business decisions without displacing the incident commander.\n- Require separate commander, communications, scribe, and primary technical lead for SEV0 and SEV1.\n- For customer-visible SEV2, keep the commander separate from the primary technical responder; communications and scribe may be combined if workload permits.\n- Announce every role assignment and transfer in the incident channel. Require verbal and written handoff with current impact, decisions, risks, and next actions.\n- Keep executives and account managers out of the technical command path; questions flow through the Executive Duty Officer or Communications Lead.", "dependencies": ["S5"]}, {"step_id": "S7", "title": "Create sustainable 24×7 coverage across the service estate", "description": "Use central command coverage and service-domain responder coverage rather than creating 28 fragile night rotations. Engineers remain responsible only for services they own or have formally trained to support.\n\n- Build a company incident-command pool of approximately 18–24 certified senior engineers and managers, with a primary and backup scheduled at all times.\n- Build 12–16-person communications and scribe pools using Support, Customer Operations, Engineering Operations, and qualified engineering managers.\n- Schedule an Engineering Director or equivalent as the 24×7 Executive Duty Officer.\n- Group related service owners into approximately 8–12 coherent responder domains only where members have training, access, and explicit acceptance.\n- Require Tier 0 and Tier 1 domains to provide 24×7 primary and secondary responders, normally with at least six trained people in each sustainable rotation.\n- Give Tier 2 services business-hours coverage plus a maintained manager escalation path. Treat Tier 3 conditions as tickets unless their impact changes.\n- Maintain distinct but coordinated coverage for the ledger application and the PostgreSQL platform.\n- Route unknown-owner events to the duty commander and platform triage temporarily. Treat every such event as an ownership control defect.\n- If a small team cannot staff a fair rotation, merge coverage only after training or provide headcount, service reassignment, or decommissioning.", "dependencies": ["S3", "S6"]}, {"step_id": "S8", "title": "Approve compensation and fatigue protections", "description": "End unpaid on-call before expanding mandatory coverage. Publish the policy through HR, Legal, Finance, and Payroll within 14 days.\n\n- Pay fixed weekly stipends for primary and secondary service rotations.\n- Pay separate stipends for duty commander, communications, and scribe assignments.\n- Provide additional call-out compensation or equivalent paid recovery time for material after-hours work.\n- Apply overtime and reporting rules correctly for non-exempt employees under federal and New York requirements.\n- Pay higher rates for company holidays and provide a protected recovery day after qualifying overnight work, SEV0 events, or prolonged SEV1 response.\n- Target no more than one primary week in six and prohibit simultaneous primary assignments.\n- Avoid consecutive primary weeks and make all swaps visible in the paging system.\n- Reduce sprint commitments for people carrying primary duty rather than expecting normal delivery capacity.\n- Provide a documented accommodation path for health, disability, or caregiving constraints without career penalty.\n- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.", "dependencies": ["S7"]}, {"step_id": "S9", "title": "Set the service readiness and runbook standard", "description": "A team cannot respond effectively at night without ownership, access, telemetry, and rehearsed recovery procedures. Apply a formal readiness gate to every Tier 0 and Tier 1 service.\n\n- Require a current architecture diagram, dependency map, dashboards, SLOs, runbooks, rollback method, feature-control method, contacts, and tested escalation path.\n- Record RTO, RPO, data-integrity requirements, regional mode, and customer-facing capabilities in the service catalog.\n- Link every paging alert to the exact runbook step expected from the responder.\n- Prevent new paging alerts for services that fail readiness review. Preserve existing critical detection through a documented exception until a safe replacement exists.\n- Write priority playbooks for PostgreSQL failure, ledger-integrity investigation, regional failover, Kubernetes control-plane degradation, payment-processor failure, queue backlog, credential compromise, and payment suspension.\n- For the ledger, document read-only or stop-processing modes, failover controls, replay and duplicate protections, backlog handling, and post-recovery reconciliation.\n- Require runbook review after material incidents and at least twice per year through exercises.", "dependencies": ["S3", "S5", "S7"]}, {"step_id": "S10", "title": "Establish one paging and incident system of record", "description": "Consolidate paging and incident coordination without requiring an unsafe big-bang replacement of every monitoring system. Monitoring sources may remain specialized, but all human pages must enter one controlled platform.\n\n- Select one enterprise paging platform and one integrated incident record, with chat, telephone, SMS, conference, status-page, ticketing, and service-catalog integrations.\n- Initially ingest events from all six current tools, then deduplicate, correlate, and route by service ownership.\n- Automate incident channel and bridge creation, role paging, timeline capture, severity changes, communications reminders, and postmortem creation.\n- Preserve immutable records of declarations, acknowledgements, role assignments, decisions, and messages.\n- Use role-based access, multifactor authentication, break-glass controls, and periodic access reviews.\n- Provide telephone and offline fallback procedures for loss of chat, identity, the paging vendor, or an AWS region.\n- Test paging and fallback paths weekly.\n- Retire a legacy paging route only after its signals have owners, quality review, successful end-to-end tests, and at least two weeks of verified operation in the new path.", "dependencies": ["S4", "S5", "S6", "S7"]}, {"step_id": "S11", "title": "Enforce alert quality and burn down noise safely", "description": "Treat paging alerts as production products with owners and quality requirements. Do not reduce noise by silently disabling detection.\n\n- Require every page to identify the service, owner, customer or SLO risk, urgency, dashboard, runbook, expected action, deduplication key, and escalation policy.\n- Define an actionable page as one requiring prompt human judgment or intervention that materially reduces customer, financial, security, or contractual risk.\n- Route informational, capacity-planning, and non-urgent conditions to dashboards or ticket queues.\n- Prefer symptom and error-budget burn alerts over raw CPU, memory, pod, or log-volume thresholds.\n- Run new alerts in shadow mode for at least seven days unless an emergency exception is approved.\n- Review alerts with less than 50% actionability, more than three firings in seven days, or repeated no-action acknowledgements within two business days.\n- Require compensating detection and an owner before suppressing or removing an alert.\n- Set a responder load target of no more than two after-hours pages per person per week. A breach creates a mandatory alert-remediation plan.\n- Review alert actionability, duplication, missed detection, and page load monthly by domain.\n- Prioritize the small number of rules producing most of the current 85% noise.", "dependencies": ["S3", "S10"]}, {"step_id": "S12", "title": "Detect payment and ledger failures before customers", "description": "Shift detection from infrastructure symptoms to customer journeys and financial outcomes. Set internal objectives stricter than the contractual 99.95% SLA.\n\n- Define SLIs and SLOs for payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, webhooks, and reporting freshness.\n- Run external synthetic transactions from outside the production boundary and through both regions at least every minute for critical paths.\n- Monitor payment failure rates, latency, queue age, delayed value, settlement-window risk, regional asymmetry, and third-party response quality.\n- Add continuous ledger controls for reconciliation breaks, unexpected balances, duplicate identifiers, replication lag, backup health, and failover readiness.\n- Add tenant or cohort anomaly detection for high-value customers and common payment methods.\n- Convert high-priority Support, account-manager, bank, and processor reports into incident candidates within five minutes.\n- Review every customer-first incident as a missed-detection defect and create a corrective action.\n- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.", "dependencies": ["S3", "S11"]}, {"step_id": "S13", "title": "Codify escalation and live incident execution", "description": "Create one time-bound path from signal to ownership and mitigation. Delivery of a notification does not count as acknowledgement.\n\n- Page the owning Tier 0 or Tier 1 primary immediately; page the secondary after five unacknowledged minutes; page the domain manager at 10 minutes; escalate to the engineering director at 15 minutes.\n- For SEV0–SEV2, page the duty commander immediately. Page the backup after five minutes and require the Engineering Director to assume command if no certified commander owns the event by 10 minutes.\n- Automatically involve command for integrity concerns, security concerns, regional events, cross-team impact, critical customer journeys, or unresolved ownership.\n- If impact remains materially unknown after 15 minutes, increase response posture rather than waiting for certainty.\n- Open one incident channel, bridge, and system record. State severity, known impact, assigned roles, current objective, and next update time.\n- Freeze unrelated changes during SEV0 and SEV1 unless the commander records an exception.\n- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.\n- Separate mitigation and diagnosis workstreams when staffing permits.\n- Require controlled backlog processing and reconciliation before resolving payment or ledger incidents.\n- Require a formal command handoff for incidents extending beyond four hours or when fatigue impairs a role holder.", "dependencies": ["S6", "S10"]}, {"step_id": "S14", "title": "Standardize internal and customer communications", "description": "Communicate known impact early without waiting for root cause. The Communications Lead uses approved facts and always states the next update time.\n\n- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish customer status within 15 minutes when customer-visible and legally safe, then update at least every 30 minutes.\n- For SEV1, brief internal stakeholders and publish status within 15 minutes, then update every 30 minutes.\n- For customer-visible SEV2, brief Support and account managers and publish or directly send notice within 30 minutes, then update every 60 minutes.\n- For SEV3, communicate directly to affected customers only when impact or contract terms require it.\n- Give account managers an approved statement and affected-customer list within 30 minutes for SEV0 or SEV1 and within 60 minutes for SEV2.\n- Use status states such as Investigating, Identified, Monitoring, and Resolved. Describe affected capabilities, symptoms, workarounds, and the next update.\n- Do not speculate about root cause, blame, data integrity, security scope, or recovery time.\n- Post a Monitoring update within 30 minutes of mitigation. Post Resolved only after stability and required reconciliation.\n- Provide a customer-facing incident summary within five business days for qualifying incidents.\n- Record and approve any delay or restriction of public details during an active security threat.", "dependencies": ["S5", "S6", "S10"]}, {"step_id": "S15", "title": "Operationalize regulatory, partner, contract, and credit decisions", "description": "Give Legal and Compliance a timed decision process while keeping technical command with the incident commander. Record both reportable and non-reportable determinations.\n\n- Build a jurisdiction and obligation matrix covering applicable NYDFS requirements, state breach laws, GLBA or FTC requirements, PCI obligations, money-transmitter rules, sponsor banks, payment networks, cyber insurance, and customer contracts.\n- Validate applicability and deadlines with counsel rather than assuming every operational incident is reportable.\n- Begin reportability assessment immediately for every SEV0, security event, suspected ledger-integrity event, and relevant SEV1.\n- Target a documented initial legal assessment within one hour for SEV0 and within two hours for other potentially reportable events.\n- Record the decision, evidence, approver, legal deadline, submission owner, and confirmation of delivery.\n- Maintain tested 24×7 contacts for regulators, banks, networks, insurers, outside counsel, and critical vendors.\n- Encode customer-specific notification clocks and channels in the customer record used by the Communications Lead.\n- Have Finance calculate affected minutes, likely credits, and contractual exposure from the incident record within five business days.\n- Review whether proactive credits or claims-based handling applies by customer segment and contract.", "dependencies": ["S4", "S5", "S14"]}, {"step_id": "S16", "title": "Make postmortems mandatory, consistent, and blameless", "description": "Use one review standard to learn from incidents and test whether controls worked. Keep learning reviews separate from performance or misconduct processes.\n\n- Require a postmortem for every SEV0 and SEV1.\n- Require one for customer-visible SEV2, customer-first detection, incidents lasting more than two hours, contractual breaches, repeat failures, control gaps, and ledger-integrity near misses.\n- Produce the factual draft within three business days, conduct the review within five, and publish the approved version within 10.\n- Use one template covering summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, lessons, and actions.\n- Analyze why detection was not earlier and why mitigation took as long as it did.\n- Examine technical, organizational, process, testing, dependency, and incentive factors rather than forcing a single root cause.\n- Use a trained facilitator and written blameless rules. Describe decisions in the context and information available at the time.\n- Publish broadly useful findings internally while maintaining restricted versions for security, privacy, personnel, or privileged material.\n- Review postmortem quality and recurring factors in a weekly Incident Review Board.", "dependencies": ["S5", "S6", "S10"]}, {"step_id": "S17", "title": "Enforce action ownership and effectiveness tracking", "description": "Treat corrective actions as risk commitments rather than suggestions. Closing a ticket is insufficient without evidence that the control or system behavior improved.\n\n- Give every action one named individual owner, manager, priority, due date, expected risk reduction, verification method, and linked work item.\n- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with intermediate milestones.\n- Reserve 10%–15% of team capacity for approved reliability and incident work.\n- Place recurrence-prevention actions for SEV0 and SEV1 ahead of discretionary feature work unless an executive accepts the residual risk.\n- Escalate overdue high-risk actions to the manager after seven days, director after 14 days, and CTO after 30 days.\n- Require written residual-risk acceptance, compensating controls, and a new review date when high-risk work is deferred.\n- Verify completed actions through tests, telemetry, drills, or production evidence.\n- Triage the 53 currently open historical actions within 30 days. Complete, re-plan, or formally accept the risk, prioritizing ledger, regional, security, and detection items.", "dependencies": ["S16"]}, {"step_id": "S18", "title": "Measure outcomes, controls, business impact, and human load", "description": "Use a balanced scorecard so teams are not rewarded for suppressing alerts or avoiding incident declarations. Report medians and 90th percentiles, not averages alone.\n\n- Measure time from first impact to internal detection, declaration, acknowledgement, command assignment, mitigation, and resolution.\n- Track customer-first detection, missed escalations, status-page timeliness, update-cadence compliance, and role conflicts.\n- Track incident count, recurrence, availability, error-budget burn, affected payment value, delayed transactions, reconciliation breaks, and SLA credits.\n- Track page volume, actionability, duplicates, after-hours pages, missed acknowledgements, and load per responder.\n- Track postmortem timeliness, action completion, action age, verified effectiveness, and repeated contributing factors.\n- Track rotation size, duty frequency, recovery days, swaps, attrition signals, and quarterly responder sentiment.\n- Hold a weekly Incident Review Board for incidents, actions, missed controls, and noisy alerts.\n- Hold a monthly executive reliability review for trends, investment decisions, contractual exposure, and accepted risks.\n- Hold a quarterly resilience and controls review with Product, Security, Compliance, Risk, and Internal Audit.\n- Reconcile dashboard data against customer cases and sampled incident records monthly to identify missing incidents or metric gaming.", "dependencies": ["S10", "S11", "S14", "S17"]}, {"step_id": "S19", "title": "Train participants and address pager resistance", "description": "Introduce the operating model as a fair exchange, not an audit mandate. The message is that people are paid, paged only for accepted services, supported by trained command, and given capacity to fix defects.\n\n- Train all employees to recognize impact, declare an incident, and find the incident channel and customer status page.\n- Give all 260 engineers role-based training in severity, escalation, evidence preservation, handoff, and financial-integrity precautions.\n- Certify incident commanders through instruction, simulation, two shadowed events or exercises, and observed performance.\n- Train Communications Leads in status writing, account-manager briefing, contractual clocks, and legal escalation.\n- Train scribes in timeline quality, decision capture, fact-versus-hypothesis labeling, and evidence handling.\n- Require responders to demonstrate access, dashboards, rollback, runbooks, and escalation competence before primary duty.\n- Require two shadow shifts before independent primary on-call.\n- Appoint one adoption champion in each team and hold weekly office hours during rollout.\n- Publish compensation, fatigue protections, service boundaries, and the accommodation process before assigning shifts.\n- Use paid working time for training, exercises, shadowing, runbook work, and certification.\n- Survey engineers at baseline, day 60, day 120, and quarterly thereafter.", "dependencies": ["S6", "S7", "S8", "S13", "S14", "S16"]}, {"step_id": "S20", "title": "Pilot on the payment critical path", "description": "Run a four-to-six-week pilot across the highest-risk customer journey. Use real incidents and exercises to correct the process before wider rollout.\n\n- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, authentication or API edge, settlement, and Support intake.\n- Activate paid primary and secondary rotations, central command, communications, the incident record, status templates, postmortems, and action tracking.\n- Run old paging paths in parallel for no more than one week, then make the new system authoritative.\n- Review every pilot page within one business day for routing, actionability, responder load, and missing context.\n- Have the program team coach incidents without silently taking command from assigned role holders.\n- Hold a weekly pilot retrospective and correct critical process or tool defects within 48 hours.\n- Require 95% command assignment within five minutes, 95% timely communications, no unpaid pages, complete postmortems, and at least a 50% page-noise reduction before expansion.", "dependencies": ["S8", "S9", "S11", "S12", "S13", "S14", "S17", "S19"]}, {"step_id": "S21", "title": "Roll out by customer journey and risk", "description": "Expand in fixed waves rather than waiting for every team to become perfect. Apply explicit readiness gates and time-limited exceptions.\n\n- Days 0–7: operate the interim declaration, command, and communications process.\n- By day 30: complete Tier 0 ownership, activate the first certified command roster, approve compensation, and begin the pilot.\n- By day 60: provide compensated 24×7 coverage for every Tier 0 and Tier 1 customer journey and route all critical pages through the new platform.\n- By day 90: complete critical detection upgrades, customer communications, regulatory playbooks, and the first major alert-noise reduction.\n- By day 120: assign every production service a response tier, owner, tested escalation, and appropriate coverage model.\n- Roll out remaining teams in two-to-three-week waves ordered by customer and dependency risk.\n- Require each wave to pass ownership, alert, runbook, access, training, compensation, and tabletop gates.\n- Disable legacy paging paths after verified cutover rather than leaving ambiguous parallel obligations.\n- Publish a weekly adoption dashboard by team and escalate failed gates as business risks.\n- Never start mandatory night coverage before compensation, staffing, training, and access are ready.", "dependencies": ["S18", "S20"]}, {"step_id": "S22", "title": "Exercise command, regional resilience, and ledger recovery", "description": "Validate the process under realistic conditions before depending on it during a crisis. Use the same action-tracking rules for exercises and real incidents.\n\n- Run the first company-wide command and communications tabletop within 30 days.\n- Run monthly domain tabletops covering payment failure, customer-first detection, third-party failure, and ambiguous ownership.\n- Exercise loss of one AWS region, Kubernetes degradation, PostgreSQL failure, suspected duplicate payments, queue backlog, and settlement risk.\n- Exercise simultaneous operational and security events to test command and disclosure boundaries.\n- Exercise loss of chat, status-page, identity, or paging providers using telephone and offline fallbacks.\n- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors.\n- Validate backups, restoration, RPO, RTO, failover prerequisites, financial controls, and post-recovery reconciliation.\n- Avoid uncontrolled production-ledger experiments; use staging, replicas, simulations, or tightly governed production tests.\n- Complete at least two cross-company exercises before the audit, including one overnight or unannounced paging test.", "dependencies": ["S9", "S13", "S14", "S15", "S19"]}, {"step_id": "S23", "title": "Test SOC 2 operating effectiveness before the auditor", "description": "Demonstrate that the controls operate consistently, not merely that policies exist. Correct failures through tracked remediation rather than rewriting historical records.\n\n- Preserve policy approvals, service ownership, schedules, compensation activation, access reviews, training, certifications, incidents, communications, postmortems, actions, and exercises.\n- Sample evidence monthly from initial signal through verified action closure.\n- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.\n- Use exceptions to document missed acknowledgements, late communications, incomplete records, and compensating controls.\n- Conduct a formal mock audit in month 7 using the populations and interview roles expected by the external auditor.\n- Trace at least one SEV0 or SEV1, one SEV2, a customer-reported event, and an exercise end to end.\n- Remediate evidence and operating gaps before external fieldwork.\n- Brief commanders, responders, Support, and Compliance for auditor interviews without scripting inaccurate answers.", "dependencies": ["S4", "S18", "S21", "S22"]}, {"step_id": "S24", "title": "Institutionalize continuous improvement", "description": "Keep incident management active after the audit by assigning permanent owners, budgets, and review cycles. Use incident trends to drive architectural investment.\n\n- Assign permanent owners for the policy, service catalog, paging platform, status page, training program, metrics, and evidence repository.\n- Review severity thresholds, communications timing, staffing, and compensation annually and after material process failures.\n- Recertify commanders and communications leads annually through observed exercises.\n- Review recurring failure families quarterly and require executive action when remediation repeatedly loses priority.\n- Maintain quarterly responder surveys and publish actions addressing fatigue, fairness, tool friction, and psychological safety.\n- Report severe incidents, SLA exposure, overdue risks, resilience investments, and customer-first detection to the board or risk committee quarterly.\n- Prioritize reduction of the shared-ledger concentration risk, stronger regional independence, deployment safety, graceful degradation, and automated mitigation.\n- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.", "dependencies": ["S18", "S21", "S23"]}], "estimated_complexity": "high", "success_metrics": "- By day 7, every suspected SEV0–SEV2 uses one incident record, one coordination channel, and one named incident commander.\n- After day 30, no SEV0 or SEV1 remains without a named commander for more than 10 minutes; at least 95% are assigned within 5 minutes.\n- By day 30, 100% of Tier 0 services have a named owner, primary and secondary escalation, dashboard, and runbook.\n- By day 60, every Tier 0 and Tier 1 customer journey has compensated 24×7 command and subject-matter coverage.\n- By day 120, all 180 production services have an owner, response tier, tested escalation path, and appropriate coverage model.\n- No mandatory night rotation starts before its compensation, training, access, and staffing controls are active.\n- Every critical primary rotation has at least six qualified responders or a documented, expiring executive exception by day 120.\n- No responder is routinely scheduled for primary duty more often than one week in six by day 120.\n- At least 95% of critical pages are acknowledged within 5 minutes by month 3.\n- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.\n- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.\n- Median mitigation time falls from 190 minutes to under 90 minutes by day 120 and under 60 minutes by month 8.\n- At least 95% of SEV0 and SEV1 customer notices are issued within 15 minutes of declaration by month 3.\n- At least 95% of customer-visible SEV2 notices are issued within 30 minutes by month 3.\n- At least 95% of incidents meet their required update cadence by month 3.\n- Monthly paging volume falls from 3,400 to no more than 1,500 by day 90 and no more than 700 by month 6.\n- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6 without loss of critical detection coverage.\n- All new paging alerts satisfy the owner, impact, action, dashboard, runbook, deduplication, and escalation standard by day 60.\n- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.\n- All 53 currently open historical actions are triaged within 30 days.\n- At least 90% of high-priority corrective actions are completed by their approved due date by month 6, with effectiveness evidence.\n- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.\n- Two end-to-end cross-company exercises, including regional failure and ledger recovery, are completed before the audit.\n- Monthly availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.\n- SLA credits decline by at least 50% on an annualized trailing basis within 12 months.\n- Quarterly on-call surveys show improving fairness and sustainability, with at least 75% favorable responses by month 6.\n- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence."}Doubled from 13 to 27 steps and adopted most of P1's R0 architecture: baseline, listening tour, layered staffing, escalation ladder, regulatory playbook, credits, runbooks, exercises, gated pilot and waves, risk register, SOC 2 dry-run. Far more complete and better sequenced, but largely derivative and it lost its own concrete compensation numbers.
- S2 forensic baseline with a specific case review of the two unowned incidents and a freeze of the reference metrics.
- S8 replaces the R0 'all 28 teams on rotation' stance with a layered model over 8–12 domains and a six-person minimum.
- S15 regulatory playbook now requires the reportability decision to be recorded even when the answer is 'not reportable'.
- S18 readiness bar blocks night paging for services without sign-off unless the EM accepts the gap in writing.
- S26 risk register with pre-committed contingencies (commander shortfall, comp delay, tool slip, pruning-caused misses, SEV1 during rollout).
- Metrics are now staged (day 90 / month 6 / month 9) rather than single 6–12 month endpoints.
- Lost the R0 dollar table ($500/$250 stipends, $75/$150 page pay, $300 SEV1 bonus, ~$420k/yr budget) — S7 now only says 'weekly stipend differentiated by tier', weakening the CFO case that R0 made well.
- No interim command step: S1 lists 'interim controls week 1' in the timeline but nothing defines them, while four rivals staff an interim commander in 48 hours to 7 days.
- S7 contradicts itself on load: 'no more than one primary week in four or one week in six depending on staffing'.
- S21 pilot depends on S7, S10, S11, S14, S18, S19 but not on the escalation path (S12) or postmortem standard (S16) it claims to exercise.
- Dropped R0's peer benchmarking against Stripe/Adyen-class platforms and the explicit tooling budget line.
- Proposal 1 : Forensic incident re-coding, resistance mapping with a co-design group, and a machine-readable ownership catalog with tiers.
- Proposal 1 : Page budget, SLA credit workflow, action escalation ladder with reserved capacity, and the readiness bar.
- Proposal 1 : Critical-path pilot with published exit gates, gated waves with coaches, and a program risk register.
- Proposal 2 : Triage of the 53 open historical actions and shadow-mode testing for new alerts.
- Proposal 4 : Hard date after which pages outside the platform are not valid obligations, plus a weekly noise leaderboard.
- Proposal 2 : A staffed interim incident-command process in the first seven days.
- Proposal 2 : SEV0 crisis tier with distinct crisis-management triggers.
- Proposal 5 : Status page within 5 minutes of SEV1.
+ Executive mandate, program funding, and governance+ Evidence baseline from incidents, alerts, and coverage gaps+ Paid on-call, fatigue safeguards, and HR compliance+ On-call staffing model and rotation rules+ Detection uplift across payments, ledger, and customer signals+ Single incident platform and alert-tool consolidation+ Customer status page, account-manager outreach, and SLA credit workflow+ Simulation program and game days+ Critical-path pilot and gate reviewProgram charter, executive mandate and fundingForensic baseline of the 31 incidents and the alert estate24x7 incident command coverage modelTeam on-call structure, rotations and routing rulesOn-call compensation, labour compliance and fairness policyDetection uplift: SLOs, synthetic journeys and ledger assuranceTool consolidation and incident platform implementationCustomer communications and status page policySLA credit and financial impact workflowSimulation programme: game days, drills and wheel of misfortunePilot with wave 0 teams
The plan produced
1. Executive mandate, program funding, and governance new
Convert the CEO email into a written mandate within 48 hours. Name one accountable program owner and a small decision group. Time-box design to five weeks so rollout starts well before the audit.
- Appoint the CTO as executive sponsor and a Head of Reliability or Incident Management as program owner with full-time authority.
- Create an 8–10 person design working group: SRE/platform lead, payments and ledger engineering managers, support lead, compliance, legal, HR, finance, and two rotating engineering managers.
- Approve budget lines: tooling consolidation, on-call compensation, training, coaching, and 2–3 dedicated program staff.
- State the non-negotiables: paid on-call, named service ownership, one severity scale, one paging platform, mandatory postmortems, and protected engineering capacity for reliability actions.
- Fix the master timeline: interim controls week 1, design weeks 1–5, pilot weeks 6–11, full rollout weeks 12–22, internal audit rehearsal month 6, audit-ready month 7.
- Publish a one-page charter to the company that says incident response is a company operating process, not an optional team practice.
2. Evidence baseline from incidents, alerts, and coverage gaps (after 1) new
Before changing anything, create an auditable before picture from the last 12 months. This baseline drives the severity design, staffing model, and executive reporting.
- Reconstruct all 31 customer-impacting incidents: detection source, first owner, severity, mitigation time, credits paid, and whether command was clear.
- Write a specific case review of the two incidents with unclear ownership for more than one hour.
- Inventory the six alert tools, alert volume by team, noise rate, rules with no owner, and rules with no runbook.
- Map the 16 teams without on-call and the 12 teams with unpaid on-call.
- Review the 53 open postmortem actions and triage the highest-risk items first.
- Freeze baseline metrics: 22-minute detection, 40% customer-first detection, 3h10 mitigation, $1.3M credits, 3,400 alerts, 85% noise, and 17% action closure.
3. Stakeholder listening, resistance mapping, and co-design group (after 1) from P1 step 3
Treat engineer pushback as design input, not an attitude problem. The objection to carrying a pager for other teams' code must shape ownership, routing, compensation, and command staffing.
- Interview leads from all 28 teams plus support, customer success, sales, compliance, and HR within two weeks.
- Separate the real causes of resistance: unpaid night work, unfamiliar systems, poor runbooks, unfair load, fear of blame, or unclear authority.
- Recruit 10–15 respected engineers and managers as co-designers so the model is built with the teams.
- Survey baseline on-call sentiment, trust in alerts, and psychological safety; repeat at 90, 180, and 365 days.
- Document the central promise: responders are paged for services they own, trained incident commanders coordinate, and all on-call work is compensated.
4. Service ownership catalog, criticality tiers, and dependency map (after 2) from P1 step 4
No alert can page the right team until every service has a named owner. Build the catalog as the routing foundation for on-call, severity impact mapping, status-page components, and audit evidence.
- Assign one accountable team to each of the 180 services, with an engineering manager, Slack channel, escalation policy, dashboard, and runbook link.
- Tier services: Tier 0 for money movement, ledger integrity, authentication, settlement, and shared PostgreSQL; Tier 1 for customer-facing degradable services; Tier 2 for internal or batch services; Tier 3 for non-critical services.
- Map customer journeys to services, databases, regions, third parties, and contractual SLA components.
- Mark orphan and shared services; require ownership, reassignment, or decommissioning within 30 days.
- Document cross-region dependencies, failover constraints, and services that defeat nominal regional redundancy.
- Treat missing ownership for Tier 0 or Tier 1 services as an executive escalation and a release-blocking risk.
5. Severity scale, declaration rights, and automatic triggers (after 2, 4)
Adopt one severity scale so declaration is a lookup, not a debate. Anyone may declare; only the incident commander may downgrade. When uncertain, start higher.
- SEV1: money movement stopped or incorrect, ledger integrity in doubt, confirmed security or data event, both regions impaired, or broad SLA-credit exposure.
- SEV2: major degradation, settlement window at risk, one or more critical customers fully down, or likely contractual breach.
- SEV3: limited impact with a workaround, no financial-integrity or security risk.
- SEV4: no customer impact; handled by ticket during business hours.
- Define fixed triggers for each level: pages, command staffing, bridge call, status-page timing, account-manager outreach, executive notification, and postmortem obligation.
- Add automatic escalation: SEV3 open more than 2 hours becomes SEV2; any incident touching the shared ledger or unclear ownership becomes SEV2 unless the commander documents otherwise.
- Publish a decision tree and 10–12 worked examples from the actual 31 incidents.
6. Incident roles, authority, and handover rules (after 5) from P1 step 6
Solve the nobody-was-in-charge failure by making command explicit, trained, and transferable. Separate coordination from technical remediation so responders are not asked to debug unfamiliar code.
- Incident Commander: owns severity, priorities, escalation, mitigation strategy, role assignment, handoffs, and closure. Does not write code during the incident. May freeze deploys, invoke pre-approved failover, and pull any on-call responder.
- Communications Lead: owns status page, internal updates, account-manager briefs, and coordination with legal or compliance.
- Scribe: maintains the timestamped record of decisions, actions, role changes, and customer communications.
- Subject-matter responders: engineers from owning teams who diagnose and remediate within their accepted domain.
- Executive liaison: required for SEV1; shields the commander from executive questions and owns board or regulator escalation.
- Require the commander to claim the role within 5 minutes of SEV1 or SEV2 declaration, state it in the incident channel, and record every handover.
- Allow role combination only for SEV3 or SEV4; require separate people for command, communications, and primary technical work at SEV1.
7. Paid on-call, fatigue safeguards, and HR compliance (after 3) new
Unpaid on-call is a retention, fairness, and legal risk in New York. Pay must be approved before any new mandatory rotation starts. This is the fastest way to reduce pager resistance.
- Create a weekly stipend for primary and secondary on-call, differentiated by tier and night responsibility.
- Add per-page or per-incident pay for out-of-hours activation, plus guaranteed recovery time after night work or major incidents.
- Create a separate stipend for duty incident commanders and communications leads because command is a heavier burden.
- Verify FLSA, New York wage-hour, overtime, holiday, and payroll treatment with HR, legal, and finance.
- Cap rotation frequency: no more than one primary week in four or one week in six depending on staffing; no simultaneous primary assignments.
- Require at least six trained responders for any 24x7 rotation; fund hiring or service reassignment where teams are too small.
- Publish the compensation package and payroll start date before asking engineers to join rotations.
8. On-call staffing model and rotation rules (after 4, 6, 7) new
Do not force 28 identical night rotations. Use a layered model: central command coverage, critical-path team coverage, and business-hours coverage for lower-tier services.
- Company incident-command and communications rotation: 24x7 pool of 25–35 trained volunteers and designated senior staff, giving primary plus secondary coverage at all times.
- Critical-path teams: 24x7 primary and secondary on-call for Tier 0 and Tier 1 domains, including payments, ledger, authentication, API edge, Kubernetes platform, and shared PostgreSQL.
- Other teams: business-hours on-call with a documented night escalation list owned by the engineering manager.
- Group services into 8–12 coherent domains so rotations are sustainable; no rotation with fewer than six people is approved without an executive exception.
- Require handoff overlap, shadow shifts before first primary duty, and no on-call during approved PTO.
- Define cross-team pull rules: the commander may page another team's on-call with a 10-minute acknowledgement obligation; this is coordinated by command, not pushed onto responders.
- Publish schedules, swap rules, and load limits in the paging tool.
9. Alert quality standard, page budget, and noise controls (after 2, 4) from P1 step 10
With 3,400 alerts and 85% noise, detection fails because people stop trusting pages. Make alert quality a condition of paging anyone.
- Every page must have a named owning team, customer or SLO impact, severity, dashboard, runbook, expected action, and escalation policy.
- Page on customer-visible symptoms: payment success rate, latency, settlement deadlines, error-budget burn, ledger integrity, and replication health.
- Demote cause-based CPU, memory, or infrastructure-only alerts to dashboards or tickets unless they map to a customer journey.
- Set a page budget: maximum 2 out-of-hours pages per person per week; breach triggers a mandatory alert-tuning sprint.
- Auto-flag alerts that fire repeatedly without action or have high no-action acknowledgement rates.
- Require shadow mode for new alerts before they page humans, except for documented emergencies.
- Review alert quality monthly by team and publish a noise leaderboard.
- Target fewer than 500 actionable pages per month and above 80% actionability within six months.
10. Detection uplift across payments, ledger, and customer signals (after 4, 9) new
Customers detected 40% of incidents first. Detection must shift to payment outcomes, ledger integrity, and inbound customer signals, not host metrics alone.
- Define SLIs and SLOs for payment initiation, authorization, settlement, reconciliation, refunds, API availability, and reporting freshness.
- Run synthetic end-to-end payment tests from outside the platform in both AWS regions every 60 seconds.
- Add continuous ledger assurance: double-entry reconciliation, replication lag, failover readiness, disk pressure, and settlement-window countdown alerts.
- Monitor per-customer anomalies for top accounts so a single-tenant outage is detected before the account manager calls.
- Convert high-priority support tickets and account-manager reports into incident candidates within 5 minutes.
- Track detection source for every incident; make customer-first detection a reviewed defect with a corrective action.
- Add partner, banking, and card-network notification intake into the same declaration path.
11. Single incident platform and alert-tool consolidation (after 5, 8, 9) new
Collapse six tools into one paging and incident-management system with one queue, one timeline, and one audit trail. Do not create a parallel audit process.
- Select an integrated stack: paging and on-call scheduling, incident workflow, Slack or chat integration, bridge calling, status-page API, and ticketing integration.
- Implement one-command declaration that creates the incident record, channel, bridge, severity label, role prompts, and clock.
- Migrate alert sources by wave; retire a legacy paging path only after routing, ownership, and acknowledgement tests pass.
- Automate evidence capture: timestamps, acknowledgements, role assignments, severity changes, communications, and postmortem links.
- Integrate with the service catalog, on-call schedules, Jira or equivalent action tracker, customer-success tooling, and the status page.
- Verify the platform works during a single-region failure, including SMS, phone, and offline fallback paths.
- Set a hard date after which pages outside the chosen platform are not valid on-call obligations.
12. Escalation paths and five-minute command rule (after 6, 8, 11) from P1 step 13
Create one path from signal to named commander in under five minutes, any hour of the day. Make escalation automatic and time-bound.
- Accept declarations from automated alerts, engineers, support, account managers, partners, and customers through the same command.
- For SEV1 and SEV2, page the duty incident commander and owning-team primary immediately.
- Acknowledgement ladder: primary 5 minutes, secondary 10 minutes, manager 15 minutes, director or executive 20 minutes.
- If no commander claims the incident within 5 minutes, the platform assigns and announces one; the assignee may hand over but cannot leave the incident unowned.
- For SEV3, require team acknowledgement within 30 minutes; otherwise create a tracked work item.
- Give the commander pre-approved authority to invoke regional failover, ledger read-only mode, feature kill switches, and partner notifications without waiting for executive sign-off.
- Record every missed acknowledgement and escalation failure for weekly review.
13. Internal communications protocol (after 6, 11) from P1 step 14
Separate the working incident channel from the audience channel so responders can work and executives, support, and sales stay informed without interrupting the commander.
- Create one incident channel and one bridge per SEV1–SEV3 incident; use a read-only broadcast channel for executives and support.
- Set update cadence: every 15 minutes for SEV1, 30 minutes for SEV2, and at state changes for SEV3.
- Use a fixed update template: impact, customer-visible symptoms, current action, ETA or next update, commander, and communications lead.
- Brief support and customer success with a live affected-customer list and approved holding statements within 15 minutes of SEV1 or SEV2.
- Require executives to route questions through the executive liaison; the commander is not interrupted.
- Define fatigue rules and formal handover for incidents lasting more than 4 hours.
14. Customer status page, account-manager outreach, and SLA credit workflow (after 5, 13) new
Replace ad-hoc status updates with a timed, owned, template-driven customer communication process. Link incidents to SLA credits so finance and customer success are not surprised.
- Publish status-page updates within 15 minutes of SEV1 declaration and 30 minutes for SEV2; update every 30 or 60 minutes until resolution.
- Assign the communications lead as the single author; use pre-approved templates reviewed by legal and communications.
- Map status-page components to customer capabilities: payments, settlement, reporting, API, onboarding, and regional availability.
- For top accounts, require direct account-manager outreach within 30 minutes of SEV1 with an approved briefing pack.
- Send proactive email or webhook notifications to subscribed customers for SEV1 and SEV2.
- Publish a resolution notice within 30 minutes of mitigation and a customer-facing incident summary within 5 business days for SEV1.
- Automate SLA credit calculation from incident duration and affected capability; review with finance and legal within 5 business days.
- Track credits by incident and root-cause family to guide reliability investment.
15. Regulatory, legal, and partner notification playbook (after 5, 14) from P1 step 16
In payments, some incidents are reportable and the clock starts at detection. Make regulatory assessment a mandatory step in the incident process, not an afterthought.
- Map obligations: NYDFS cybersecurity-event rules, state breach laws, GLBA safeguards, PCI DSS if in scope, money-transmitter duties, card-network and sponsor-bank contracts, and cyber-insurer notice.
- Require a legal or compliance reportability assessment within 2 hours for every SEV1 and every security-related SEV2, even when the answer is not reportable.
- Maintain a 24x7 contact matrix for regulators, sponsor banks, card networks, outside counsel, insurers, and law enforcement.
- Pre-draft notification templates and preserve legal review.
- Encode customer-contract notification deadlines into account tiers and the communications workflow.
- Record every reportability decision, approver, deadline, and submission confirmation in the incident record.
16. Postmortem standard and blameless review (after 5, 6) from P1 step 18
Replace inconsistent postmortems with a mandatory, blameless, single-format process. The discipline comes from deadlines, facilitation, and action tracking.
- Require postmortems for all SEV1 and SEV2 incidents, customer-detected incidents, incidents over 2 hours, repeat failures, and ledger near-misses.
- Draft within 3 business days, peer review within 5, publish within 10 for SEV1 and SEV2.
- Use one template: summary, impact, timeline, detection analysis, response analysis, contributing factors, what worked, what failed, and action items.
- Make blamelessness explicit: focus on systems and decisions, not individual fault; never use postmortems in performance discipline.
- Hold a weekly incident review board to review postmortems, ratify severity, and challenge weak actions.
- Maintain a searchable postmortem library and quarterly recurring-cause analysis.
- Require a trained facilitator for major reviews; the incident commander attends but does not facilitate.
17. Action-item tracking, ownership, and delivery gates (after 11, 16) from P1 step 19
Only 11 of 64 actions were closed. Give postmortem actions the same status as customer commitments, with named owners and visible escalation.
- Create every action as a ticket with one named individual owner, priority, due date, and verification method.
- Use delivery classes: containment within 7 days, corrective work within 30 days, strategic work within 90 days.
- Reserve 15–20% of team sprint capacity for reliability and incident actions.
- Escalate overdue items: manager at 7 days, director at 14 days, CTO dashboard at 30 days.
- Block related feature releases when overdue P0 actions prevent recurrence of a severe incident.
- Require director approval and documented residual risk acceptance for overdue high-risk items.
- Verify effectiveness after completion; closing a ticket without evidence does not close the action.
- Target 90% of high-priority actions completed on time within two quarters.
18. Runbooks, critical-incident playbooks, and readiness bar (after 4, 8) from P1 step 20
Poor runbooks are a real cause of pager resistance and slow mitigation. Define a minimum readiness bar before a service is allowed to page anyone at night.
- Require for every Tier 0 and Tier 1 service: architecture summary, dependencies, dashboards, alert-to-runbook map, rollback procedure, feature flags, escalation contacts, and customer-impact statement.
- Write major playbooks for shared PostgreSQL ledger failure, regional failover, Kubernetes control-plane loss, payment-processor outage, settlement-window breach, duplicate-payment suspicion, and security compromise.
- Define ledger recovery rules: failover procedure, read-only degraded mode, reconciliation, RPO/RTO, and data-loss tolerance approved by executives.
- Test runbooks in drills at least twice a year; mark untested runbooks stale.
- Prevent paging alerts for services without readiness sign-off unless the engineering manager accepts the gap in writing.
- Keep runbooks linked from every alert and incident template.
19. Training, certification, and role readiness (after 6, 12, 13, 16) from P1 step 21
Command and communications are skills. Build a tiered certification path so rotations are staffed by people who have practiced, not by whoever is around.
- All employees: 1-hour module on declaring incidents, finding the incident channel, and reading the status page.
- All responders: half-day training on severity, acknowledgement, escalation, runbooks, and evidence hygiene.
- Incident commanders: 2-day course plus two shadowed incidents and one simulation before certification.
- Communications leads: training on status-page writing, customer language, account-manager briefs, and regulatory triggers.
- Scribes: training on timeline discipline and audit evidence.
- Certify for 12 months; renew through a simulation.
- Require shadow shifts before independent primary duty; no new hire holds primary within 90 days.
- Publish the certification register as an audit artifact.
20. Simulation program and game days (after 11, 18, 19) new
Rehearse the process before it meets a real SEV1. Simulations build commander confidence, expose runbook gaps, and produce audit evidence.
- Run monthly 60-minute tabletops using real incidents from the 31-incident baseline.
- Run quarterly game days covering regional failover, ledger replica promotion, dependency failure, partner outage, and security event.
- Run twice-yearly unannounced paging drills to measure night acknowledgement times.
- Include support, account managers, legal, compliance, and executives in at least one exercise per quarter.
- Produce tracked action items from every exercise using the same board as real incidents.
- Measure time to commander, time to first status update, and time to mitigation decision.
21. Critical-path pilot and gate review (after 7, 10, 11, 14, 18, 19) new
Prove the model on the highest-risk services with willing teams before full rollout. Run a six-week pilot with daily feedback and public exit criteria.
- Pilot with payments, ledger/database, platform/Kubernetes, API edge, authentication, and support intake.
- Activate severity scale, duty commanders, paid rotations, single paging platform, alert budget, status-page policy, and postmortem process.
- Hold a weekly pilot retrospective and fix process defects quickly.
- Validate night acknowledgement, cross-team pull response, severity clarity, and compensation payroll.
- Exit gate: commander assigned within 5 minutes in 95% of incidents, status page on time, page noise down at least 50%, postmortems on time, and positive on-call sentiment.
- Publish pilot results to the whole company as the main adoption argument.
22. Wave rollout across all 28 teams (after 21) from P1 step 25
Roll out by criticality and dependency, not by calendar alone. Use readiness gates so teams are not forced live without coverage.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1 teams. Wave 3: Tier 2 teams. Wave 4: Tier 3 and internal platforms.
- Gate per team: catalog entry complete, alerts migrated, runbooks ready, six-person rotation staffed, one commander candidate nominated, compensation in payroll, and one tabletop passed.
- Assign a program coach to each wave for three weeks.
- Freeze legacy paging paths for each team after successful onboarding.
- Publish a live adoption scoreboard by team.
- Complete all teams by week 22, leaving months of operating evidence before the audit.
23. Metrics, dashboards, and review cadence (after 11, 16, 21) from P1 step 24
Instrument the process itself. Leadership must see whether the program is working, and auditors must see operating evidence, not retrospective paperwork.
- Track response metrics: MTTD, time to declare, time to commander, acknowledgement time, mitigation time, resolution time, and customer-first detection rate.
- Track quality metrics: page volume, alert actionability, missed pages, postmortem timeliness, action closure rate, and status-page compliance.
- Track business metrics: availability against 99.95%, SLA credits, repeat incidents, and error-budget consumption.
- Track people metrics: on-call load, night pages per person, recovery time usage, sentiment, and attrition signals.
- Hold weekly incident review, monthly reliability review, quarterly executive review, and annual policy review.
- Publish dashboards internally and give every metric a target and owner.
24. Change management, incentives, and culture (after 3, 7, 21) from P1 step 27
The process will be judged on fairness. Communicate the deal repeatedly and make participation recognized, compensated, and safe.
- Core message: you are paid, paged only for what you own, supported by a trained commander, and given real capacity for actions.
- Run CTO all-hands, team roadshows, office hours, FAQ, and an internal incident-management hub.
- Add incident response and reliability work to promotion criteria and manager objectives.
- Recognize good postmortems, alert-noise reduction, and calm incident leadership.
- Provide a written path for engineers who cannot do nights; cover those shifts with paid volunteers or adjusted staffing.
- Prohibit retaliation for good-faith declaration or escalation.
- Publish sentiment survey results, including bad news, to maintain credibility.
25. SOC 2 evidence design and internal dry-run audit (after 16, 17, 22, 23) from P1 step 26
Design the process so audit evidence is a by-product of normal operations. Test it internally before the external auditor does.
- Map the process to SOC 2 criteria: incident identification, response, recovery, monitoring, communication, control activities, availability, and corrective action.
- Approve versioned policies: incident response policy, severity standard, on-call policy, communications policy, postmortem standard, and exception process.
- Retain incident records, paging logs, status-page history, postmortems, action tickets, training records, drill records, and access reviews for the audit period.
- Record regulatory reportability decisions even when no notification is required.
- Run an internal dry-run audit at month 6: sample at least 15 incidents and trace evidence end to end.
- Fix gaps with at least four weeks before the audit.
- Maintain an exception and remediation log instead of claiming perfection.
26. Program risk register and contingency planning (after 1) from P1 step 28
Name the likely failure modes now and pre-commit responses. Review the register monthly with the executive sponsor.
- Commander volunteer shortfall: roster command duty for senior engineers and managers until the pool reaches 25 trained people.
- Compensation delay: use immediate time-off-in-lieu plus phased stipend, but do not start mandatory night rotations without pay.
- Tool migration slip: preserve single paging and routing scope; defer advanced automation if needed.
- Alert pruning causing missed incidents: move noisy alerts to ticket first, observe 30 days, then delete.
- Burnout in experienced teams: monitor load weekly and cap pages per person.
- Major SEV1 during rollout: program owner shifts to incident support, wave schedule slips one wave, sponsor is informed same day.
- Shared ledger concentration: track blast-radius reduction and failover improvements as top strategic actions.
27. Continuous improvement, maturity roadmap, and post-audit sustainability (after 23, 25) from P1 step 29
Prevent the classic failure where the process decays after the audit. Build the second-year operating plan before the first year ends.
- Hold quarterly process retrospectives with commanders, responders, support, and communications leads.
- Re-baseline targets every six months; tighten goals once current targets are met.
- Move from lagging metrics to leading indicators: error-budget burn, near-miss rate, drill performance, and action ageing.
- Plan year-two improvements: follow-the-sun coverage, automated mitigation, ledger blast-radius reduction, error-budget release gates, and per-customer real-time impact reporting.
- Keep annual policy review, certification renewal, drills, and board reporting on a permanent calendar independent of the audit cycle.
- Report quarterly to the board or risk committee on availability, credits, severe incidents, overdue actions, and resilience investment.
- Median time to detect customer-impacting incidents falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median mitigation time falls from 3h10 to under 90 minutes by day 120 and under 60 minutes by month 9.
- A named incident commander is assigned within 5 minutes in at least 95% of SEV1 and SEV2 incidents; zero incidents remain unowned for more than 10 minutes.
- First status-page update occurs within 15 minutes of SEV1 declaration and 30 minutes of SEV2 declaration in at least 95% of qualifying incidents.
- Monthly paging volume falls from 3,400 to under 500 actionable pages within six months, with alert actionability above 80%.
- Out-of-hours pages average no more than 2 per responder per week; sustained breaches trigger mandatory alert remediation.
- Six legacy alerting tools are consolidated into one paging and incident platform, with legacy paging paths disabled by week 16.
- 100% of Tier 0 and Tier 1 services have a named owning team, escalation path, dashboard, and runbook by day 60.
- All 28 teams are onboarded with readiness gates by week 22; every 24x7 critical-path rotation has at least six trained responders.
- 30 or more certified incident commanders and 20 or more certified communications leads provide continuous primary and secondary coverage.
- Paid on-call is approved by HR, legal, and finance and active in payroll before any new mandatory night rotation begins.
- 100% of required SEV1 and SEV2 postmortems are drafted within 3 business days and published within 10 business days in the standard template.
- Postmortem action closure rises from 17% to at least 90% of high-priority actions completed by their due date within two quarters.
- All 53 open historical actions are triaged within 30 days; high-risk unaccepted items are completed or formally risk-accepted within 90 days.
- Repeat incidents from a known unaddressed contributing factor decline by at least 50% within six months.
- Annualized SLA credits fall from $1.3M to under $400k within 12 months.
- Monthly availability meets or exceeds 99.95% by month 6, with exceptions reviewed at the executive reliability meeting.
- At least two cross-company exercises, including regional and ledger scenarios, are completed before the SOC 2 audit, with critical findings tracked.
- The month-6 internal dry-run audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete evidence.
- SOC 2 Type II incident-response controls pass with zero exceptions at the month-8 audit.
- On-call sentiment improves quarter over quarter; fewer than 10% of engineers report unwillingness to participate in their owned-service rotation by month 6.
- No increase in attrition among engineers on active rotations compared with baseline.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 223f0a8a-068a-4034-ad20-8e2ff8553040, Agent: opus5_initial_1, LLM: anthropic/claude-opus-5):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 9 months.
- Customer-first detection falls from 40% of incidents to under 10% within 6 months and under 5% within 12.
- Median time to mitigate falls from 3 h 10 min to under 60 minutes within 12 months.
- An Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes.
- Status page updated within 15 minutes of SEV1 declaration and 30 minutes of SEV2 in 95% of cases.
- Monthly alert volume falls from 3,400 to under 500 pages, with actionability above 75%; out-of-hours pages under 2 per person per week.
- All six legacy alerting tools consolidated into one paging platform, legacy paging paths disabled, by week 16.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by their due date.
- SLA credits fall from $1.3M to under $400k in the first 12 months.
- Customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar.
- All 28 teams onboarded by week 24, with 24x7 rotations of 6+ certified responders for every Tier 0/1 team.
- 30+ certified Incident Commanders and 20+ certified Communications Leads, giving 24x7 primary and secondary command cover.
- Paid on-call policy approved by HR, Legal and Finance and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among engineers on rotation.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked action items.
Steps (29):
1. Program charter, executive mandate and funding
Convert the CEO's frustration into a named program with one accountable owner, a budget and a deadline that is earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, the 28 teams will opt out.
- Appoint a single Incident Management Program Lead (Head of Reliability/SRE) with direct exec sponsorship from CTO and CEO.
- Form a steering group: CTO, VP Eng, Head of Support/CS, CISO/Compliance, Legal, Finance (SLA credits), HR (on-call pay).
- Set the non-negotiables: one severity scale, one paging tool, one postmortem format, mandatory action tracking, paid on-call.
- Fix the timeline: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready by week 28 (four weeks of buffer before the audit).
- Approve budget lines: tooling (~$150–250k/yr), on-call compensation (~$600k–1M/yr), 2–3 dedicated program FTEs. Anchor it against $1.3M of credits plus incident cost.
2. Forensic baseline of the 31 incidents and the alert estate (depends on: 1)
Before designing anything, rebuild the facts. Re-open all 31 incidents and profile the 3,400 monthly alerts so every later design decision is evidence-based.
This also creates the "before" picture the exec and the auditor will compare against.
- Re-code each incident: trigger, service, detection source (customer vs monitor), timestamps for detect/acknowledge/declare/mitigate/resolve, who led, credits paid, root cause family.
- Quantify the 40% customer-first detections: which signal was missing in each case.
- Classify the two "nobody in charge" incidents minute by minute; use them as the burning-platform story.
- Audit the six alerting tools: volume per tool, per team, per alert rule; identify the top 50 rules that produce most of the 85% noise; find rules with no owner and no runbook.
- Baseline the numbers formally: MTTD 22 min, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed. Freeze them as the reference line.
3. Stakeholder listening tour and resistance map (depends on: 1)
Engineer pushback against "carrying a pager for other teams' code" is the main delivery risk. Treat it as a design input, not an attitude problem.
Run structured interviews across all 28 teams, plus Support, CS and Sales, in two weeks.
- Test the real objection: is it unpaid work, night sleep, unfamiliar code, poor runbooks, or fear of blame? Each has a different fix.
- Collect current informal practices — the 12 teams already on-call are the pilot candidates and the source of veterans.
- Document the promise that answers the objection: **you are only paged for services your team owns**, plus a trained commander who runs the incident and pulls in others.
- Map influencers and blockers by team; recruit 10–15 credible engineers as a design working group so the process is co-authored, not imposed.
- Survey baseline sentiment (trust in alerts, willingness to be on-call, burnout) to re-measure at 6 and 12 months.
4. Service ownership catalog and criticality tiering (depends on: 2)
You cannot page the right person across 180 services until each service has a named owning team. This is the foundation of both on-call fairness and severity mapping.
Build a machine-readable catalog (Backstage or equivalent) that is the single source of truth for routing.
- One owning team per service, a named engineering manager, a Slack channel, a paging escalation policy, a dependency list.
- Tier services by business impact: Tier 0 (money movement, ledger, auth, shared PostgreSQL cluster), Tier 1 (customer-facing but degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map each Tier 0/1 service to the customer-visible capability it supports (payment initiation, settlement, reporting, onboarding).
- Flag orphan services and cross-team shared components; force an ownership decision for each within 30 days, or schedule decommissioning.
- Publish coverage gaps to the steering group: any Tier 0 service without an owner is an executive escalation.
5. Severity scale and declaration criteria (depends on: 2, 4)
Define a five-level scale with objective, payments-specific triggers so declaration is a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade.
Each level triggers a fixed bundle of response, comms and postmortem obligations.
- **SEV1**: money movement stopped or incorrect, ledger integrity in doubt, data breach, full region loss, >10% of customers impacted. Triggers: immediate 24x7 page of IC + comms + exec, bridge within 5 min, status page within 15 min, mandatory postmortem, regulator assessment.
- **SEV2**: severe degradation, settlement at risk of missing a window, single large/strategic customer fully down, SLA breach likely. Triggers: IC paged, status page within 30 min, mandatory postmortem.
- **SEV3**: partial or workaround-available degradation, no credit exposure. Team-led, business-hours comms, postmortem optional but encouraged.
- **SEV4/5**: minor or internal-only; ticket-tracked, no paging.
- Add auto-escalation rules: any SEV3 open >2 h, or any incident touching the shared ledger cluster, becomes SEV2 automatically. Include a severity decision tree and 12 worked examples drawn from the 31 real incidents.
6. Incident roles, decision authority and handover rules (depends on: 5)
Solve the "nobody in charge for an hour" failure by making command explicit, transferable and logged.
Define five roles with written responsibilities, entry criteria and explicit authority.
- **Incident Commander**: owns the incident, not the fix. Authority to declare severity, pull any engineer, approve customer-impacting mitigations, invoke failover and authorise spend. The IC never types in the terminal.
- **Communications Lead**: owns status page, internal updates, account-manager briefings and the exec summary. Single voice to customers.
- **Scribe**: maintains the timeline, decisions and open questions; feeds the postmortem and the audit evidence trail.
- **Subject-Matter Responders**: engineers from owning teams; they investigate and remediate, and report to the IC.
- **Executive Liaison** (SEV1 only): shields the IC from exec questions and owns regulator/board escalation.
- Rules: the IC role is assumed within 5 minutes of declaration, stated explicitly in the channel ("I am IC"), and any handover is announced and logged. Roles may be combined below SEV2; never at SEV1.
7. 24x7 incident command coverage model (depends on: 6, 4)
Command is staffed by a small, trained, cross-team pool — not by 28 teams individually. This is what makes 24x7 realistic in one New York time zone.
Design a central rotation that scales with a growing certified pool.
- Create a **Duty Incident Commander** rotation of 25–35 certified volunteers (target ~1 per team, plus managers and senior engineers), giving each person roughly one week per 6–8 months.
- Pair with a Duty Comms Lead rotation (Support/CS leads plus engineering managers, ~15–20 people) and a Scribe pool (rotating, lowest barrier, used as the training entry point).
- Coverage: primary + secondary IC at all times; hard 5-minute acknowledgement SLA with automatic failover to secondary, then to the on-call engineering director.
- Night coverage options to evaluate in writing: US-only rotation with paid night stipend now; a Lisbon/Dublin or APAC follow-the-sun cell as a 12-month option; a 24x7 NOC-style triage desk for first-line detection.
- Eligibility: certification required (S21); commanders are volunteers with manager approval and can step out with 30 days' notice.
8. Team on-call structure, rotations and routing rules (depends on: 4, 6)
Rebuild team on-call around the principle that answers the pushback: **you are only paged for code your team owns**.
Apply a tiered obligation so 28 teams are not treated identically.
- Tier 0/1 owning teams (expected 14–18 teams): 24x7 primary + secondary, minimum 6 people per rotation, one-week shifts, handover Wednesday mornings.
- Tier 2/3 teams: business-hours on-call with a best-effort out-of-hours escalation path, no night paging.
- Platform/Infrastructure and Database teams: 24x7, since they own the shared PostgreSQL ledger cluster and the Kubernetes/regional layer.
- Rotations under 6 people are merged across teams or backfilled by hiring; no rotation of fewer than 4 is approved.
- Routing: every page resolves through the service catalog to the owning team's escalation policy; cross-team pages are made by the IC, never by an alert.
- Guardrails: maximum one week in four, no on-call in the week after a SEV1 you led, protected recovery time after any night page, and a per-person page budget (see S10).
9. On-call compensation, labour compliance and fairness policy (depends on: 8, 3)
Unpaid on-call is both a retention risk and a legal exposure in New York. Paying for it is the fastest way to convert resistance into participation.
Design the scheme with HR, Legal, Finance and Payroll, and publish it before asking anyone to sign up.
- Base stipend per week on rotation, differentiated by tier: e.g. $800–1,200 for 24x7 Tier 0/1, $300–500 for business-hours rotations, with premiums for holidays and weekends.
- Per-incident payment for out-of-hours activation (e.g. $150 per night page plus hourly beyond one hour) and guaranteed time-off-in-lieu after night work.
- Separate Duty IC stipend, since command is a distinct and heavier burden.
- Verify FLSA exempt/non-exempt treatment, NY State wage rules and overtime exposure for non-exempt staff; document the legal review.
- Budget and model the annual cost; get board/CFO approval as a line item, benchmarked against $1.3M of credits.
- Add non-cash elements: on-call time counted as delivery load (teams reduce sprint commitment by ~15%), incident leadership recognised in promotion criteria, and a public quarterly report of on-call load per team.
10. Alert quality standard and page budget (depends on: 2, 4)
3,400 alerts a month at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a contractual condition of paging someone.
Publish a standard, then enforce it mechanically.
- Every paging alert must have: a named owning team, a documented customer impact, a runbook link, a tested threshold, and a severity mapping. Alerts failing this are demoted to ticket or deleted.
- Page only on symptoms that affect customers (SLO burn rate, error budget, queue depth against settlement deadlines); cause-based CPU/memory alerts become dashboards or tickets.
- Set a **page budget**: maximum 2 out-of-hours pages per person per week. Breach triggers a mandatory alert-tuning sprint for the owning team and blocks new alert creation.
- Auto-quarantine: any alert that fires more than 5 times a month without action, or has >70% no-action acknowledgements, is silenced automatically and returned to its owner.
- Monthly alert review per team: kill, tune, or keep, with the numbers on screen.
- Target: 3,400 → under 500 pages a month, with actionability above 75% within six months.
11. Detection uplift: SLOs, synthetic journeys and ledger assurance (depends on: 10, 4)
The goal is to stop customers telling you first. Detection must be driven by customer-visible outcomes, not host metrics.
Instrument the money path end to end and alert on it.
- Define SLOs for each Tier 0/1 customer capability: payment initiation success rate, authorisation latency, settlement file timeliness, API availability and reporting freshness. Tie them to the 99.95% contractual SLA with a stricter internal target.
- Deploy synthetic transactions from outside the platform, in both regions, every 60 seconds, covering the full payment lifecycle including a small real-value canary flow where feasible.
- Add ledger assurance checks: continuous double-entry balance reconciliation, replication lag and failover-readiness alarms on the shared PostgreSQL cluster, and settlement-window countdown alerts.
- Build per-customer anomaly detection for the top 100 accounts (volume drop, error spike) so a single-tenant outage is detected before the account manager calls.
- Create an "inbound signal" bridge: any support ticket or account-manager report matching impact keywords auto-creates a triage incident within 5 minutes.
- Track every incident's detection source; make "customer detected first" a reviewed defect with its own follow-up action.
12. Tool consolidation and incident platform implementation (depends on: 5, 7, 8, 10)
Collapse six alerting tools into one paging and incident platform so there is a single queue, a single timeline and a single audit record.
Run a short, time-boxed selection and migrate within the pilot window.
- Select an integrated stack: paging/on-call scheduling plus an incident management layer (e.g. PagerDuty + incident.io/FireHydrant, or a single vendor) and a hosted status page.
- Implement one-command declaration in Slack (`/incident declare`) that creates the channel and bridge, pages the Duty IC, sets severity, opens the timeline and starts the clock.
- Migrate all monitoring sources to route into the one platform; decommission direct paging from the legacy six and block new integrations that bypass it.
- Automate the evidence trail: timestamps, role assignments, severity changes, comms sent, and postmortem linkage exported for SOC 2.
- Integrate with the service catalog for routing, Jira for actions, Salesforce/CS tooling for affected-customer lists, and Zoom/Slack Huddle for the bridge.
- Hard requirement: the platform must work when AWS in one region is down — verify out-of-band paging (SMS/phone) and a printed/offline fallback runbook.
13. Detection-to-escalation path and the five-minute command rule (depends on: 6, 7, 12)
Write the single path from "something looks wrong" to "someone is in charge", and make it impossible to skip.
The design target is detection to commander in under five minutes, any hour.
- Entry points: automated alert, engineer observation, support ticket, account manager, partner bank, customer-facing SEV hotline. All converge on the same declaration command.
- Anyone in the company may declare up to SEV2; nobody is punished for over-declaring. Publish that rule in writing and repeat it.
- Auto-page ladder: Duty IC (5 min) → secondary IC (5 min) → on-call Director (10 min) → CTO. Same ladder for the owning team's responder.
- Cross-team pull: the IC can page any team's on-call directly, with a 10-minute acknowledgement obligation. This is the reciprocal commitment that makes single-team ownership viable.
- Explicit takeover protocol: if no one claims IC within 5 minutes, the platform assigns it and announces it; the assignee cannot decline, only hand over.
- Define standing severity triggers for immediate regional failover, ledger read-only mode and partner-bank notification, with pre-authorised decision rights so the IC does not wait for an executive.
14. Internal communications protocol (depends on: 6, 12)
Standardise the internal channel so responders, executives and support see the same picture without interrupting the IC.
Separate the working channel from the audience channel.
- One incident channel per incident (auto-created), one bridge, and a read-only broadcast channel for executives, Support and Sales.
- Update cadence by severity: SEV1 every 30 minutes even if nothing has changed; SEV2 every 60 minutes; SEV3 at state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, ETA or next update time, current IC and Comms Lead.
- Exec briefing rule: executives ask questions only to the Executive Liaison; the IC is not interrupted. Publish this as a behavioural expectation signed by the exec team.
- Support/CS enablement: a live affected-customer list and a holding statement within 15 minutes of SEV1/SEV2 so the front line is never guessing.
- Handover protocol for incidents beyond 4 hours: formal IC handover checklist, fatigue rule, and staffing of a second shift.
15. Customer communications and status page policy (depends on: 5, 14)
Customers currently learn of outages from their own monitoring and hear from whoever happens to be around. Replace that with a timed, owned, pre-approved process.
The Comms Lead is the single author; templates remove the need to write under pressure.
- Timing commitments: status page posted within 15 minutes of SEV1 declaration and 30 minutes for SEV2; updates every 30/60 minutes; resolution notice within 30 minutes of mitigation; customer-facing summary within 5 business days for SEV1.
- Pre-approve 12–15 templates with Legal and Comms (degradation, delay in settlement, API errors, security event, third-party failure) so nothing needs legal review mid-incident.
- Subscription-based status page with per-component granularity mapped to the customer capabilities from S11, plus an email/webhook/RSS feed.
- Tiered outreach: top 100 accounts get a direct named call or email from their account manager within 30 minutes of SEV1, with a briefing pack from the Comms Lead; long tail gets the status page and a proactive email.
- Rules of language: state impact and next update time, never speculate on cause, never assign blame to a vendor before facts are confirmed.
- Run a quarterly customer-perception check with the top accounts on whether comms were timely and useful.
16. Regulatory, partner and legal notification playbook (depends on: 5, 15)
In payments, some incidents are reportable and the clock starts at detection. Build the assessment into the process so it is never an afterthought.
Work with Legal, Compliance and the CISO to produce a decision tree and contact matrix.
- Map obligations: NYDFS Part 500 (72-hour cybersecurity event notification), state breach laws, GLBA/FTC Safeguards, PCI DSS if card data is in scope, sponsor-bank and card-network contractual notice windows, and any FinCEN/OFAC implications.
- Add a mandatory regulatory-assessment checkpoint to every SEV1 and every security-related SEV2, owned by the Executive Liaison, completed within 2 hours of declaration and recorded even when the answer is "not reportable".
- Build the contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel, with 24x7 numbers and named backups.
- Pre-draft notification letters and hold them under legal privilege review.
- Check customer contracts for bespoke notification SLAs (often 1–4 hours for enterprise accounts) and encode them in the customer tiering.
- Test the playbook once per quarter as part of the simulation programme.
17. SLA credit and financial impact workflow (depends on: 5, 15)
Link incidents to money so severity, credits and prioritisation stay consistent — and so Finance stops being surprised.
Make credit calculation an automated output of the incident record, not a negotiation.
- Define the availability measurement method per contract, per component, and agree it with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and the per-capability telemetry; generate a proposed credit schedule within 5 business days of resolution.
- Decide the posture: proactive credits for the top tier (reputation upside) versus claims-based for the rest; document the approval chain.
- Track credits per incident and per root-cause family; feed a quarterly report showing which reliability investments would have prevented which credits.
- Set a target: reduce credits from $1.3M to under $400k in year one, and use that delta as the ongoing business case.
18. Postmortem standard and blameless review forum (depends on: 5, 6)
Replace "some incidents, various formats" with a mandatory, single-format, blameless process with fixed deadlines.
The discipline is in the deadlines and the forum, not in the template.
- Mandatory for: every SEV1 and SEV2, every incident where a customer detected it first, every incident over 2 hours, every repeat of a known cause, and every near-miss involving the ledger. Optional but templated for SEV3.
- Fixed timeline: draft within 3 business days, peer review within 5, published company-wide within 10. The IC owns delivery; the owning team's manager is accountable.
- One template: timeline, customer and financial impact, detection analysis (why not sooner), response analysis (why mitigation took as long as it did), contributing factors, what went well, action items with owner and due date.
- Blameless rules in writing: describe systems and decisions in the context available at the time; no individual named as a cause; HR and management commit that postmortems are never used in performance reviews.
- Weekly 60-minute Incident Review Board: reviews all postmortems from the prior week, challenges quality, ratifies severity, and approves or rejects action items. Attendance by engineering directors is mandatory.
- Publish a searchable postmortem library and a quarterly "top five recurring causes" analysis.
19. Action item ownership and tracking system (depends on: 18, 12)
11 of 64 actions closed is the single clearest symptom of a process nobody enforces. Give actions the same status as customer commitments.
Track them where engineering work already lives, with visible escalation.
- Every action gets: a named individual owner (not a team), a priority class, a due date and a Jira ticket auto-created from the postmortem.
- Priority classes with hard SLAs: P0 prevents recurrence of a SEV1, due in 30 days; P1 in 60 days; P2 in 90 days. P0s are committed into the next sprint before any roadmap work.
- Capacity rule: teams reserve a standing 15–20% of sprint capacity for reliability and incident actions. Without reserved capacity, the actions will not land.
- Escalation ladder for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Monthly reporting of closure rate by team in the engineering leadership review; include it in manager performance objectives.
- Target: 90% of P0/P1 actions closed on time within two quarters.
20. Runbooks, major-incident playbooks and the on-call readiness bar (depends on: 4, 8)
Nobody can respond well to unfamiliar systems at 3 a.m. without runbooks — and poor runbooks are a real part of the pager resistance.
Define a minimum readiness bar that a service must meet before it is allowed to page anyone.
- Readiness checklist per Tier 0/1 service: current architecture diagram, dependency map, dashboard link, alert-to-runbook mapping, rollback procedure, feature-flag kill switches, escalation contacts, and a data-loss/latency impact statement.
- Write major-incident playbooks for the top failure modes derived from S2: shared PostgreSQL ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank/third-party outage, settlement-window breach, and suspected security compromise.
- Prioritise the shared ledger: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and a clearly stated data-loss tolerance (RPO/RTO) signed off by the exec.
- Runbooks must be tested at least twice a year in a drill; untested runbooks are marked stale in the catalog.
- Enforcement: a service without readiness sign-off cannot create paging alerts, and the gap is reported to its director.
21. Training, certification and the commander academy (depends on: 6, 13, 14, 18)
Command is a skill, not a title. Build a certification path so 24x7 coverage is staffed by people who have practised.
Use a tiered curriculum with real assessment.
- **Scribe** (2 hours): timeline discipline and tooling. The entry point for everyone.
- **Responder** (half a day): severity scale, declaration, escalation, runbook use, comms hygiene. Mandatory for every engineer joining an on-call rotation.
- **Incident Commander** (two days plus shadowing): command presence, delegation, decision-making under uncertainty, severity calls, handover, exec management. Requires two shadowed incidents and one simulated SEV1 before certification.
- **Communications Lead** (one day): status-page writing, customer tiering, legal boundaries, regulator triggers.
- Certification is valid 12 months and renewed via a simulation; the register of certified people is an audit artefact.
- Add on-call onboarding per team: a new joiner shadows two shifts before holding primary, and never holds primary in their first 90 days.
22. Simulation programme: game days, drills and wheel of misfortune (depends on: 21, 12, 20)
The process must be rehearsed before it meets a real SEV1. Simulations also build the commander pool and expose runbook gaps cheaply.
Run a standing calendar rather than one-off exercises.
- Monthly 60-minute tabletop ("wheel of misfortune") per engineering group, using a real past incident from the 31.
- Quarterly full-scale game day in production or a production-like environment: regional failover, ledger replica promotion, dependency failure, with the whole role structure activated and timed.
- Twice-yearly unannounced paging drill to measure real acknowledgement times at night.
- One security-incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Every exercise produces a lightweight postmortem and action items in the same system as real incidents.
- Measure and publish drill metrics: time to IC, time to first status update, time to correct mitigation decision.
23. Pilot with wave 0 teams (depends on: 22, 9, 11)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Run a six-week pilot with tight measurement and a public verdict.
- Select 5–6 teams: core payments, ledger/database, platform/Kubernetes, API gateway, plus two of the 12 teams already on-call.
- Activate the full stack for them: new severity scale, Duty IC rotation, single paging tool, alert budget, status-page policy, mandatory postmortems, paid on-call.
- Hold a weekly pilot retro; expect and document 20–40 process defects, and fix them in the standard before rollout.
- Validate the hard questions: does the 5-minute IC rule hold at 3 a.m.? Do cross-team pulls get answered? Is the severity tree unambiguous?
- Exit criteria: MTTD under 10 minutes for pilot services, IC assigned within 5 minutes in 95% of incidents, page volume down 50%, all postmortems on time, positive on-call sentiment.
- Publish a one-page pilot result to the whole company — this is the main adoption argument for the remaining teams.
24. Metrics, dashboards and the review cadence (depends on: 12, 18, 23)
Instrument the process itself, so improvement is visible and the audit has evidence of monitoring and review.
Define a small set of metrics with owners and a fixed meeting rhythm.
- Response metrics: MTTD, time to declare, time to IC assigned, MTTA, MTTM, MTTR, incidents per month by severity, % detected by customers first.
- Quality metrics: page volume per person per week, alert actionability rate, page budget breaches, postmortem on-time rate, action closure rate and ageing.
- Business metrics: SLA credits paid, availability against 99.95% per capability, error-budget consumption, repeat-incident rate.
- People metrics: on-call load distribution across teams, out-of-hours pages per person, on-call sentiment and attrition among on-call staff.
- Cadence: weekly Incident Review Board (postmortems and actions), monthly Reliability Review (metrics per team, alert hygiene, on-call load), quarterly Executive/Board review (credits, trends, investment asks), annual policy review.
- Every metric gets a target and a named owner; dashboards are self-serve and public inside the company.
25. Wave rollout across all 28 teams with readiness gates (depends on: 23, 24)
Roll out in four waves of six to eight teams, every three weeks, ordered by criticality. Each wave passes an explicit gate rather than a deadline.
Gates keep quality high and make the standard credible.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with 6+ certified responders, one IC candidate nominated, one drill passed.
- Assign each wave a named coach from the program team for three weeks of hands-on support.
- Gate criteria are checked and signed by the director; teams that fail are re-scheduled, not waived.
- Freeze legacy tooling per wave: after onboarding, the old alerting paths are disabled, not left as a fallback.
- Publish a live adoption scoreboard by team so progress is social, not administrative.
26. SOC 2 control mapping, evidence automation and internal dry-run audit (depends on: 18, 19, 25)
Design the process so audit evidence is a by-product of doing the work, then test it before the auditors do.
Engage the auditor early to confirm the interpretation of controls.
- Map the process to the Trust Services Criteria: CC7.3 and CC7.4 (incident identification, response, recovery), CC7.2 (monitoring), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), plus availability criteria A1.2.
- Produce and approve formal policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard — versioned, signed, annually reviewed.
- Automate evidence: incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, postmortem library, action-item closure reports, training and certification register, drill records.
- Confirm the observation window with the auditor and ensure the process is operating for a minimum of three months before fieldwork.
- Run an internal dry-run audit at month six: sample 15 incidents and walk the full evidence chain; fix gaps with 8 weeks to spare.
- Keep a remediation log for any incident where the process was not followed, with the corrective action — auditors respond better to documented exceptions than to a claim of perfection.
27. Change management, incentives and communications campaign (depends on: 3, 9, 23)
Run this in parallel from day one. The process will be judged by engineers on fairness, and by executives on visible results.
Communicate the deal explicitly and repeatedly.
- The deal in one sentence: **you are paid for on-call, you are paged only for what you own, a trained commander runs the incident, and your postmortem actions get real sprint capacity**.
- Launch communications: CTO all-hands, per-team roadshows, a one-page process card for laptops, an internal wiki hub and a Slack support channel with a 4-hour answer SLA.
- Recognition: incident-response contribution in promotion criteria and performance frameworks, quarterly awards for best postmortem and biggest alert-noise reduction, public thanks after every SEV1.
- Manager accountability: adoption, alert hygiene, action closure and on-call load in each engineering manager's quarterly objectives.
- Handle the exceptions: a written path for engineers who cannot do nights (caring responsibilities, health), covered by stipended volunteers elsewhere.
- Track sentiment quarterly and publish the results, including bad news, to keep credibility.
28. Program risk register and contingency planning (depends on: 1)
Name the ways this program fails and pre-commit the response. Review it monthly in the steering group.
The main risks are predictable.
- **Volunteer shortfall for the IC pool**: contingency is to make command a rostered duty for engineering managers and senior engineers until the pool reaches 25.
- **Compensation not approved in time**: fall back to time-off-in-lieu plus a phased stipend, but do not launch mandatory night on-call without some compensation.
- **Tool migration slipping**: keep the single-queue requirement and cut scope on the incident-management layer, not on paging consolidation.
- **Alert pruning causing a missed incident**: prune from paging to ticket first, observe for 30 days, then delete; keep a recovery path.
- **Burnout or attrition among the 12 experienced on-call teams**: monitor load weekly and cap individual page counts.
- **A major SEV1 mid-rollout**: pre-agree that the program lead becomes a full-time responder and the wave schedule slips by one wave, with the steering group informed the same day.
29. Continuous improvement, maturity roadmap and post-audit sustainability (depends on: 24, 25, 26)
Protect against the classic failure: the process decays once the audit is passed. Build the second-year plan before the first year ends.
Set a maturity model and a forward roadmap with owners.
- Quarterly process retrospective with the IC pool: what in the process itself slowed us down, what needs simplifying, is the severity scale calibrated?
- Re-baseline targets every six months; a process that hits all targets is under-ambitious.
- Year-two roadmap candidates: follow-the-sun coverage cell, automated mitigation and self-healing for the top three recurring causes, error-budget policy that gates releases, per-customer real-time impact reporting, and blast-radius reduction for the shared ledger cluster (the largest single structural risk).
- Move from lagging metrics (MTTR) to leading ones (error-budget burn, near-miss rate, drill performance).
- Make the annual policy review, certification renewal and drill calendar permanent calendar items owned by the Head of Reliability, independent of the audit cycle.
- Report to the board quarterly on availability, credits and incident trends so the process keeps executive attention after SOC 2 is signed.
Previous Proposal 2 (ID: 7ee53cee-6357-4ed7-bd99-535e92e8b1e6, Agent: gpt5.6-sol_initial_2, LLM: openai/gpt-5.6-sol):
Estimated Complexity: high
Success Metrics: - Within 7 days, every suspected SEV0–SEV2 has one incident record, one channel, and a named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named incident commander for more than 10 minutes.
- By day 30, 100% of Tier 0 and Tier 1 services have a named owner, primary escalation, secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and subject-matter coverage.
- By day 120, 100% of production services have sustainable ownership and tested escalation paths.
- At least 95% of SEV0 and SEV1 pages are acknowledged within 5 minutes by month 3.
- At least 95% of SEV2 pages are acknowledged within 10 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to 10 minutes by day 90 and 5 minutes by month 6.
- The proportion of incidents first detected by customers falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to below 90 minutes by day 120 and below 60 minutes by month 6.
- At least 95% of qualifying incidents meet their initial customer-communication deadline by month 3.
- At least 95% of published incidents meet their required update cadence by month 3.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without loss of Tier 0 or Tier 1 detection coverage.
- Monthly pages fall from 3,400 to no more than 1,500 by day 90, with alert actionability and missed-detection reviews used as countermeasures against unsafe suppression.
- 100% of new paging alerts satisfy the owner, runbook, dashboard, action, severity, and escalation quality rules by day 60.
- 100% of required SEV0 and SEV1 postmortems are drafted within 3 business days and reviewed within 5 business days by month 2.
- At least 90% of postmortem actions are completed by their approved due dates by month 6.
- All 53 currently open historical actions are triaged within 30 days; all unaccepted high-risk items are completed within 90 days.
- Repeat incidents with the same unaddressed contributing factor decline by at least 50% within 6 months.
- Every primary rotation has at least six trained responders or a documented, time-limited executive exception by day 120.
- No responder is routinely scheduled more frequently than one primary week in six by day 120.
- Two end-to-end cross-company exercises, including regional and ledger scenarios, are completed before the audit, with all critical findings assigned and tracked.
- Monthly availability meets or exceeds the 99.95% contractual target by month 6, with exceptions reviewed at the executive reliability meeting.
- SLA credits decline by at least 50% on an annualized trailing basis by month 8.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (19):
1. Establish ownership, authority, and funding
Launch the program within 48 hours under an executive sponsor. Give one program owner authority to standardize incident management across all 28 teams.
- Name the CTO or equivalent as executive sponsor and a Head of Incident Management or Reliability as directly accountable owner.
- Form a working group with Engineering, SRE or Platform, Product, Support, Customer Success, Communications, Security, Legal, Compliance, Risk, HR, Finance, and Internal Audit.
- Approve authority for an incident commander to stop deployments, roll back releases, disable features, shift traffic, invoke continuity plans, and pause payment processing when integrity is at risk.
- Preserve financial controls. The incident commander may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access requirements.
- Fund paging tools, compensation, training, observability work, exercises, and dedicated reliability capacity.
- Reserve engineering capacity for incident remediation. Start with 10% of capacity and adjust through quarterly risk reviews.
- Record the current baselines: 31 customer-impacting incidents, 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M in credits, 3,400 monthly alerts, 85% noise, and 11 of 64 actions closed.
- Maintain a risk register for staffing gaps, shared-ledger concentration, regional failover, alert coverage, third parties, and audit readiness.
2. Install immediate minimum controls (depends on: 1)
Put an interim process in place during the first seven days. Do not wait for tool consolidation, policy perfection, or the SOC 2 audit.
- Publish a one-page interim severity guide and incident declaration procedure.
- Establish one continuously monitored incident declaration path through chat, telephone, and the paging system.
- Create a standard incident channel, conference bridge, incident document, and event naming convention.
- Staff an interim primary and backup incident commander at all times. Compensate this duty retroactively under the final compensation policy.
- Give trained duty personnel access to the status page, paging system, dashboards, support queue, service catalog, and emergency contacts.
- Require an incident commander to be named within 10 minutes for every suspected major incident.
- Direct Support to escalate credible customer reports immediately rather than waiting for engineering confirmation.
- Triage all 53 open historical postmortem actions. Complete, re-plan, or formally risk-accept the items affecting ledger integrity, payment duplication, regional resilience, security, and detection first.
- Hold a daily 15-minute operational review until permanent controls are working.
3. Create the service and dependency catalog (depends on: 1)
Build a reliable ownership map for all production services and customer journeys. This is the basis for paging, escalation, impact assessment, and audit evidence.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, data stores, queues, external processors, banking partners, and customer-facing endpoints.
- Assign each component a single accountable team, primary responder group, secondary escalation group, engineering manager, and product owner.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 according to financial integrity, customer impact, dependency centrality, and contractual obligations.
- Treat the ledger, payment orchestration, authentication, settlement, reconciliation, and critical shared infrastructure as Tier 0 or Tier 1.
- Map every important customer journey to its service, database, cloud-region, and third-party dependencies.
- Record SLOs, RTOs, RPOs, data classification, dashboards, runbooks, deployment controls, feature flags, and failover methods.
- Assign separate but coordinated responders for the ledger application and the shared PostgreSQL platform.
- Document whether each service is active-active, active-passive, or region-bound. Identify dependencies that make nominal regional redundancy ineffective.
- Make missing ownership or missing runbooks a release-blocking risk for Tier 0 and Tier 1 services.
4. Adopt severity and incident lifecycle standards (depends on: 1)
Approve one impact-based severity model for operational, security, data, and third-party incidents. When evidence is incomplete, start at the higher credible severity and downgrade later.
- SEV0, crisis: Use for actual or credible unauthorized, lost, duplicated, or corrupted movement of money; ledger integrity loss; material security compromise; material data exposure; both-region failure; or an event likely to require crisis or regulatory management. Page all roles immediately, engage executives, Security, Legal, Compliance, and Risk, and consider pausing payment activity.
- SEV1, critical: Use for widespread inability to initiate, process, settle, or reconcile payments; a core journey failing without a viable workaround; material regional impact; fast SLA-budget exhaustion; or an imminent integrity risk. Staff all incident roles, notify the executive duty officer, and publish customer communications.
- SEV2, major: Use for a material customer subset, one or more critical customers, significant degradation with a workaround, partial transaction failure, or a likely contractual impact. Assign an incident commander and subject-matter responders; add communications and scribe roles whenever customers are affected.
- SEV3, minor: Use for localized, low-impact degradation with no financial-integrity, security, regulatory, or material contractual risk. The owning team leads the response and keeps an internal record; external communication is not normally required.
- Base severity on actual or credible impact, not the seniority of the reporter, number of alerts, or presumed complexity of the fix.
- Permit any employee to declare an incident. Only the incident commander may lower severity after recording the evidence and rationale.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of customer impact. Define resolution only after stability, backlog processing, transaction recovery, and required ledger reconciliation are complete.
- Start measurement from the earliest reliable indication of impact, including telemetry, customer reports, and partner notifications.
5. Define roles and sustainable 24x7 staffing (depends on: 3, 4)
Separate command from technical remediation. This allows trained commanders to coordinate any incident without asking engineers to debug code they do not own.
- Incident commander: Owns severity, priorities, role assignment, escalation, decision cadence, mitigation strategy, handoffs, and final closure. One person has command at a time.
- Communications lead: Owns internal notices, status-page updates, account-manager briefs, approved customer language, and coordination with Legal or regulators.
- Scribe: Maintains a timestamped timeline of observations, decisions, commands, owners, and status changes. Automation may assist but does not replace human validation for SEV0 and SEV1.
- Subject-matter responders: Diagnose and mitigate only services or domains for which they have accepted ownership, training, access, and runbooks.
- Executive duty officer: Removes organizational obstacles and approves exceptional business decisions. This role does not take command unless a formal transfer occurs.
- Security, Legal, Compliance, Support, Vendor Management, and Business Continuity join according to predefined triggers.
- Create a company-wide incident-command rotation with at least eight certified primary commanders and eight qualified backups. Use weekly rotations with explicit handoffs.
- Create similarly sustainable communications and scribe pools using Engineering Operations, Support, Customer Operations, and Communications personnel.
- Group service responders into approximately 8–12 coherent product or platform domains rather than creating 28 fragile rotations. Each domain rotation should normally contain at least six trained responders.
- Do not place an engineer into another team's responder pool without training, access, runbooks, shadow shifts, and explicit acceptance by both teams.
- Maintain dedicated database-platform and ledger-application escalation coverage for the shared PostgreSQL environment.
- Require distinct people for commander, communications, and primary technical lead during SEV0 and SEV1 incidents.
- Require a verbal and written handoff when any incident role changes. Record the exact time and the new role owner.
6. Implement compensation and fatigue safeguards (depends on: 5)
End unpaid on-call before expanding coverage. Treat availability, interrupted personal time, and overnight recovery as compensable work.
- Pay a fixed stipend for each primary on-call week and a secondary stipend equal to a defined percentage of the primary amount.
- Pay a higher holiday stipend. Apply overtime and call-out rules to non-exempt employees as required by law.
- Give exempt employees a minimum call-out credit or equivalent paid recovery time for material after-hours work.
- Provide a paid recovery day after prolonged overnight work, a SEV0, or a qualifying SEV1. Managers must arrange daytime coverage rather than expecting normal output.
- Have HR, Finance, and employment counsel publish dollar amounts, tax treatment, eligibility, and payroll procedures within 14 days. Apply the policy consistently across teams and locations.
- Target rotations no more frequent than one week in six. Exceptions require a time-limited staffing plan and executive risk acceptance.
- Avoid consecutive primary and secondary weeks. A person must not be primary for two simultaneous domain rotations.
- Track after-hours pages, sleep interruptions, swaps, missed acknowledgements, and reported burnout by rotation.
- Trigger a staffing or alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
- Permit responders to declare themselves temporarily unfit after overnight work without performance penalty.
7. Consolidate incident and paging tooling (depends on: 4)
Create one operational system of record while migrating safely from the six current alerting tools. Consolidation must reduce ambiguity without creating a monitoring gap.
- Select one enterprise paging and escalation platform and one integrated incident record.
- Initially ingest events from all six tools. Deduplicate, correlate, and route them through the new platform before retiring sources.
- Integrate paging with chat, conference bridges, ticket tracking, the service catalog, observability tools, and the customer status page.
- Automatically capture declaration time, acknowledgements, role assignments, severity changes, messages, decisions, mitigated time, and resolved time.
- Use role-based access, multifactor authentication, break-glass controls, immutable audit logs, and periodic access reviews.
- Provide mobile and telephone fallback paths if chat, identity, or the primary paging tool is unavailable.
- Test paging, escalation, status publication, and conference access every week.
- Retire a legacy alert path only after its signals have named owners, successful end-to-end tests, and at least two weeks of verified operation in the new platform.
8. Improve detection and enforce alert quality (depends on: 3, 7)
Shift detection toward customer journeys, payment outcomes, and ledger integrity. Infrastructure metrics alone will not solve the current customer-first detection problem.
- Instrument payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, and reporting with SLOs and business-level success metrics.
- Run external synthetic transactions and API checks from outside the production boundary and from both AWS regions.
- Monitor transaction failure rates, processing latency, queue age, unprocessed volume, reconciliation breaks, unexpected ledger balances, duplicate identifiers, regional asymmetry, and third-party response quality.
- Correlate application telemetry with Kubernetes, AWS, PostgreSQL, network, deployment, and feature-flag events.
- Route high-priority support cases and credible partner notifications into the same incident declaration path within five minutes.
- Define noise as a page that is duplicate, informational, unactionable, non-production, or requires no timely human action.
- Require every paging alert to name an owner, affected service, urgency, customer or SLO risk, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Send non-urgent conditions to a ticket queue rather than a pager.
- Run new alerts in shadow mode for at least seven days unless an emergency risk exception is approved. Test both firing and recovery behavior.
- Review any alert with less than 50% actionability or more than three firings in seven days within two business days.
- Never silently disable a noisy alert. Verify compensating detection, record the decision, and assign a correction owner first.
- Review alert actionability, false positives, missed detection, and page load with every responder group each month.
9. Codify acknowledgement and escalation paths (depends on: 3, 4, 5, 7, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts when a qualifying signal or customer report enters the system.
- For SEV0 and SEV1, page the owning primary immediately; page the secondary after five unacknowledged minutes; page the domain manager and company incident commander at 10 minutes; and engage the executive duty officer by 15 minutes.
- For SEV2, require primary acknowledgement within 10 minutes and incident-command assignment within 15 minutes. Escalate to the secondary and manager when either target is missed.
- For SEV3, require acknowledgement within 30 minutes when immediate production action is needed. Otherwise create a prioritized work item.
- Automatically page the company incident commander for any credible integrity or security concern, cross-team event, customer-visible Tier 0 failure, regional event, or unresolved ownership question.
- If impact remains unknown after 15 minutes, raise severity rather than waiting for certainty.
- Let the incident commander summon dependency owners, cloud support, database support, payment processors, banking partners, and vendors through maintained escalation contacts.
- Test vendor contacts and premium-support entitlements quarterly.
- Route alerts with no valid owner to the central command rotation, then treat the missing ownership record as a control defect.
- Require human acknowledgement. Delivery to a device or chat channel does not count.
- Record every missed acknowledgement, failed escalation, and manual contact workaround for review.
10. Standardize live incident execution (depends on: 4, 5, 7, 9)
Give responders one concise operating procedure for the first minutes through resolution. Prioritize limiting customer and financial harm before proving a root cause.
- Open a dedicated channel, bridge, incident record, and timeline immediately for SEV0 through SEV2.
- Have the incident commander state severity, known impact, current hypothesis, immediate objective, assigned roles, and next update time.
- Freeze unrelated production changes during SEV0 and SEV1 incidents. Record exceptions approved by the incident commander.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Use pre-approved runbooks for region failover, Kubernetes recovery, PostgreSQL failover, credential rotation, queue recovery, and payment suspension.
- Guard against split-brain, replay, duplication, and out-of-order processing during regional or database recovery.
- Require reconciliation and controlled backlog processing before declaring payment or ledger incidents resolved.
- Keep diagnosis and mitigation workstreams separate when enough responders are available.
- State decisions and owners aloud and in the incident record. Avoid unrecorded direct-message command paths.
- Require stability for a severity-specific observation period before closure. Reopen the incident if impact recurs during that period.
- Conduct an explicit operational handback to the owning team, Support, and Customer Success.
11. Standardize internal, customer, and regulatory communications (depends on: 4, 5, 7, 10)
Communicate known impact early without waiting for a root cause. Use approved facts, acknowledge uncertainty, and give the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish an initial customer status within 15 minutes when disclosure is operationally and legally appropriate, then update every 15 minutes.
- For SEV1, notify internal stakeholders within 15 minutes, publish an initial customer status within 15 minutes, and update at least every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send an initial notice within 30 minutes. Update at least every 60 minutes.
- Do not normally publish SEV3 events. Notify specifically affected customers if contracts or material impact require it.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. State affected capabilities, customer symptoms, workarounds, regions, and next update time.
- Do not speculate about root cause, blame, security scope, recovery time, or data integrity.
- Give account managers a single approved briefing and an affected-customer list. Prohibit contradictory or improvised incident explanations.
- Maintain templates for outages, delays, data-integrity investigation, third-party failure, regional failure, security events, and resolution.
- Issue a resolution notice only after operational recovery and required reconciliation. Provide a customer-facing incident summary within five business days for qualifying events.
- Have Legal and Compliance maintain a jurisdiction, regulator, sponsor-bank, network, cyber-insurer, partner, and contract notification matrix.
- Where applicable, explicitly track the current New York cybersecurity-event notification clock, including the 72-hour requirement, without assuming every incident is reportable.
- Have Legal record the reportability decision, decision time, evidence, approver, deadline, and submission confirmation.
- Allow Security or Legal to limit public detail during an active threat, but require the reason and an alternative stakeholder plan to be recorded.
- Coordinate service-credit calculations and contractual notices with Finance and Customer Success from the same incident record.
12. Make postmortems mandatory and actionable (depends on: 4, 7, 10)
Use postmortems to improve systems and controls, not to assign personal blame. Keep performance or misconduct processes separate from the learning review.
- Require a postmortem for every SEV0 and SEV1.
- Require one for a SEV2 that affected customers, incurred credits, breached an SLO or contract, involved financial or data integrity, repeated a prior failure, exposed a control gap, or lasted more than two hours.
- Permit incident command, Security, Compliance, or the service owner to require a review for a near miss.
- Produce a factual draft within three business days and hold the cross-functional review within five business days.
- Use one template covering executive summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, what worked, what failed, and lessons.
- Include why monitoring did or did not detect the event before customers.
- Avoid a single-root-cause assumption. Examine technical, organizational, process, dependency, testing, and incentive factors.
- Give every action one owner, due date, priority, expected risk reduction, verification method, and linked engineering item.
- Classify actions as containment due within 7 days, corrective work due within 30 days, or strategic work normally due within 90 days.
- Require director approval and documented residual-risk acceptance for overdue high-risk actions.
- Verify effectiveness after implementation. Closing a ticket without evidence does not close the action.
- Publish broadly useful reviews internally. Maintain access-restricted versions for security, privacy, personnel, or legally privileged details.
13. Measure performance and review it routinely (depends on: 7, 8, 11, 12)
Use outcome, process, quality, and human-sustainability measures together. Do not reward teams for suppressing declarations or hiding incidents.
- Measure detection time from first impact to first internal signal, declaration time, acknowledgement time, role-staffing time, mitigation time, resolution time, and recurrence.
- Report both median and 90th percentile. Break results down by severity, service tier, customer journey, region, detection source, and owning domain.
- Track customer-first detection, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, and role conflicts.
- Track availability, error-budget consumption, failed-payment volume, delayed value, reconciliation breaks, impacted customers, contractual breaches, and service credits.
- Track alert volume, actionability, duplicates, after-hours pages, missed pages, pages per responder, and tool-source distribution.
- Track required postmortems completed on time, actions completed by due date, action age, verified effectiveness, and repeat contributing factors.
- Track rotation size, on-call frequency, swaps, recovery days, attrition signals, and quarterly responder sentiment.
- Hold a weekly operational review for recent incidents, overdue actions, alert problems, and upcoming risk.
- Hold a monthly executive reliability review covering trends, investment decisions, accepted risks, and SLA exposure.
- Hold a quarterly resilience and control review with Security, Compliance, Risk, Internal Audit, and Product leadership.
- Use team scorecards to direct investment and assistance, not individual performance penalties.
- Reconcile dashboard data against a monthly sample of incident records and customer cases to detect metric gaming or missing incidents.
14. Train and certify participants (depends on: 4, 5, 9, 10, 11, 12)
Train people before assigning full independent duty. Use paid working time for training, shadowing, exercises, and certification.
- Train all employees to recognize impact, declare an incident, and find the incident channel and status page.
- Train engineers and Support on severity, escalation, customer-report handling, evidence preservation, and financial-integrity precautions.
- Certify incident commanders through instruction, tabletop exercises, shadow incidents, and observed command performance.
- Train communications leads in status writing, contractual communications, regulator escalation, and avoiding unsupported claims.
- Train scribes in timestamping, decision capture, evidence hygiene, and separating fact from hypothesis.
- Require subject-matter responders to demonstrate dashboard, runbook, rollback, failover, and access competence for their assigned domain.
- Add training to new-hire onboarding and repeat role-specific certification annually.
- Appoint an incident-management champion in each of the 28 teams to collect feedback and support local adoption.
- Conduct listening sessions focused on pager fairness, cross-team boundaries, tooling friction, and psychological safety.
- Publish that command duty is process coordination, not responsibility for understanding or repairing another team's code.
15. Pilot and expand the on-call model (depends on: 3, 5, 6, 8, 9, 14)
Pilot the model on the highest-risk customer journeys before expanding it. Correct staffing, alert, access, and compensation defects at each stage gate.
- Start with the ledger, payment orchestration, Kubernetes platform, PostgreSQL platform, authentication, settlement, and Support intake.
- Run the command, communications, and domain rotations in parallel with existing paths for two weeks.
- Verify primary and secondary coverage, handoffs, access, runbooks, paging, conference access, status publication, and compensation processing.
- Require at least two shadow shifts before independent primary duty.
- Review every pilot page within one business day for routing accuracy, actionability, responder load, and missing context.
- Expand by customer journey and dependency domain, not by arbitrary team order.
- Provide central first-line triage if useful, but keep technical remediation with the accepted service owner.
- Do not use contractors or a managed service as the sole incident commander or sole owner of payment and ledger remediation.
- Permit temporary shared domain rotations only after service owners document training, access, runbooks, and escalation boundaries.
- Set an executive-reviewed deadline and remediation plan for any production service that cannot provide sustainable 24x7 ownership.
16. Exercise regional, ledger, and communication failures (depends on: 10, 11, 14, 15)
Validate the process under realistic conditions before relying on it. Begin in tabletop and staging environments, then use controlled production tests where risk permits.
- Run a company-wide incident-command tabletop within 30 days of policy approval.
- Exercise loss of one AWS region, Kubernetes control-plane degradation, shared PostgreSQL failure, payment-processor failure, queue backlog, credential compromise, and suspected duplicate payments.
- Exercise a simultaneous operational and security event to test command boundaries and disclosure control.
- Exercise status-page failure and loss of the primary chat or paging provider.
- Exercise overnight staffing, role handoff, executive escalation, account-manager messaging, and a potential regulator-notification decision.
- Validate backups, restore procedures, RPO, RTO, failover prerequisites, and post-recovery reconciliation.
- Do not inject uncontrolled changes into the production ledger. Use replicas, staging, simulations, or tightly governed production tests.
- Record exercise observations as tracked actions under the same ownership and due-date rules as incident actions.
- Run at least one domain exercise per quarter and two cross-company exercises before the SOC 2 audit.
17. Execute a time-boxed enterprise rollout (depends on: 2, 6, 7, 11, 12, 13, 15)
Use fixed implementation waves so the audit deadline does not become the start date. Report progress weekly and escalate missed stage gates as business risks.
- Days 0–7: Establish governance, interim command coverage, one declaration path, provisional severity, and daily operational reviews.
- By day 14: Approve the core policy, role definitions, communications timings, compensation design, and historical-action triage.
- By day 30: Complete Tier 0 ownership, certify the first command roster, begin the on-call pilot, enable standard incident records, and run the first tabletop.
- By day 60: Provide 24x7 coverage for all Tier 0 and Tier 1 customer journeys, integrate the six alert sources, and enforce postmortem tracking.
- By day 90: Migrate critical paging, implement customer-journey detection, complete status and regulatory playbooks, and materially reduce alert noise.
- By day 120: Assign sustainable ownership and escalation for every production service and complete the first controlled regional or continuity exercise.
- By day 180: Complete tool retirement decisions, verify action closure, rerun weak scenarios, and demonstrate improving detection and mitigation trends.
- In month 7: Conduct a mock audit and executive readiness review, leaving at least one month to correct evidence or operating defects.
- Use exception records with owners, expiry dates, compensating controls, and executive approval. Do not allow indefinite verbal exceptions.
18. Build SOC 2 evidence as the process operates (depends on: 1)
Design evidence collection at the start rather than reconstructing it before the audit. Demonstrate both control design and sustained operation.
- Map the incident process to applicable SOC 2 criteria with Compliance and the auditor, including detection, response, communication, change management, access, availability, and corrective action.
- Maintain approved, version-controlled policies, procedures, severity definitions, role descriptions, and exception records.
- Preserve rotation schedules, compensation activation, training attendance, certification, paging tests, access reviews, and exercise results.
- Preserve incident declarations, timestamps, role assignments, communications, decisions, status updates, postmortems, and corrective-action evidence.
- Record regulatory and contractual notification assessments, including decisions that no notification was required.
- Define retention, confidentiality, legal-hold, and access requirements for operational and security records.
- Sample evidence monthly and trace incidents from initial signal through action verification.
- Have Internal Audit or an independent control owner test the process in months 4 and 6.
- Correct control failures through tracked actions rather than editing historical records.
- Conduct the formal mock audit in month 7 using the same evidence populations expected for the external audit.
19. Sustain accountability and continuous improvement (depends on: 13, 17, 18)
Make incident management an operating discipline rather than an audit project. Keep policy, staffing, tools, and investment aligned with changing customer and system risk.
- Assign permanent owners for the incident policy, paging platform, status page, service catalog, training program, and metrics.
- Review severity thresholds, communication timings, compensation, and staffing at least annually and after material incidents.
- Use incident trends to prioritize architectural work on the shared ledger, regional independence, deployment safety, dependency isolation, and graceful degradation.
- Review repeat incidents and repeat contributing factors quarterly. Require executive action when remediation repeatedly loses priority.
- Survey responders quarterly and publish actions addressing fatigue, fairness, psychological safety, and tool friction.
- Recognize effective incident leadership, early declaration, useful postmortems, and preventive work.
- Prohibit retaliation for good-faith incident declaration or escalation.
- Provide the board or risk committee a quarterly summary of severe incidents, SLA exposure, regulatory events, overdue high-risk actions, and resilience investment.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
Previous Proposal 3 (ID: c051663c-fdf6-4324-bfbe-a5f42ed62eed, Agent: qwen3.8-max_initial_3, LLM: alibaba/qwen3.8-max):
Estimated Complexity: high
Success Metrics: - **MTTD reduced from 22 minutes to <5 minutes** within 6 months of full rollout.
- **Customer-first detection rate reduced from 40% to <10%** within 6 months.
- **MTTM for SEV1 incidents reduced from 3 h 10 min to <60 minutes** within 9 months.
- **Monthly SLA credit payouts reduced from ~$325K to <$100K per quarter** within 12 months.
- **Alert volume reduced from 3,400/month to <600 actionable alerts/month** within 6 months; signal-to-noise ratio >80%.
- **Postmortem completion rate: 100% of SEV1/SEV2 incidents** have a blameless postmortem within 5 business days.
- **Postmortem action-item completion rate >90% within 30 days** of the postmortem (up from ~17%).
- **Zero incidents with >15 minutes of unowned command** (down from 2 incidents with >1 hour).
- **100% on-call coverage**: all 28 teams staffed with primary + secondary on-call 24×7 within 14 weeks.
- **On-call compensation adopted**: 100% of on-call engineers receiving stipends and page pay; on-call satisfaction score ≥4/5 in quarterly survey.
- **Status-page first-update within 15 minutes for SEV1 and 30 minutes for SEV2**, 100% compliance.
- **SOC 2 Type II audit passed** at month 8 with zero incident-response findings.
- **All 260 engineers trained**; 56+ certified ICs and 28+ certified CLs active within 14 weeks.
- **Six alerting tools consolidated to one** within 6 months; legacy tools decommissioned.
- **Regulator notification process tested**: at least one tabletop exercise includes a NY DFS / FinCEN notification drill, and the legal/compliance playbook is documented and approved.
- **Quarterly IMC reviews held consistently** with published KPI dashboards and action-item tracking.
- **On-call participation resistance resolved**: <10% of engineers report 'unwilling to participate' in the 6-month pulse survey (baseline to be measured in S1).
Steps (13):
1. Assess Current State and Baseline Metrics
Build an evidence-based picture of the current incident management reality before designing anything new.
Collect and catalog the last 12 months of incident data: all 31 customer-impacting incidents, 3,400 monthly alerts, on-call coverage gaps across the 28 teams, and the 11 of 64 closed postmortem action items. Interview one lead from each of the 28 teams to surface pain points, political concerns (the 'carrying a pager for other teams' pushback), and tool sprawl.
Deliverables to produce:
- **Alert inventory**: which of the six alert tools feed which teams, alert volume per team, noise rate per tool, and overlap between tools.
- **Incident timeline analysis**: median detection-to-notification-to-mitigation-to-resolution times, who detected first (internal vs. customer), who mitigated, and where handoff gaps occurred.
- **On-call coverage map**: which 12 teams have on-call, which 16 do not, rotation length, compensation status, and escalation paths (or lack thereof).
- **Postmortem audit**: format variance, action-item tracking gaps, and the two incidents with no clear owner for over an hour.
- **Tooling and integration audit**: Kubernetes observability stack, six alerting tools, status-page provider, communication channels (Slack, email, phone), and any existing runbooks.
- **Compliance gap analysis**: SOC 2 Type II CC7.3/CC7.4 requirements vs. current practice, with a risk register for the eight-month window.
- **Peer benchmarking**: incident management practices at 3–4 comparable B2B fintech platforms (e.g., Plaid, Stripe, Adyen) for severity scales, on-call comp, and MTTR targets.
2. Secure Executive Sponsorship and Form the IM Governance Body (depends on: 1)
Anchor the program with visible, top-down authority so that 28 teams adopt changes they did not individually request.
The CEO email about 'outages we hear about from clients' is a ready-made mandate. Convert it into a formal sponsorship structure.
- Appoint an **executive sponsor** (CTO or VP Engineering) who owns the program end-to-end and reports to the CEO monthly.
- Create an **Incident Management Office (IMO)**: one dedicated senior incident management lead, one tooling/platform engineer, and one part-time data analyst.
- Establish an **Incident Management Council (IMC)**: one engineering manager from each of the 28 teams, plus the VP of Customer Success, a compliance lead, and a security lead. The IMC meets bi-weekly during rollout, monthly thereafter.
- Draft and circulate an **executive mandate memo** that states: incident response is a shared operational obligation, not a per-team favor; participation in on-call rotations is a condition of employment for production-facing roles; and the program is not optional pending the SOC 2 audit.
- Allocate a dedicated budget line for on-call compensation, tooling consolidation, status-page licensing, training, and external facilitation.
3. Define Severity Levels and Automatic Triggers (depends on: 1)
Replace the current ad-hoc triage with a five-level severity taxonomy that every engineer, support agent, and account manager can apply in under 60 seconds.
- **SEV1 – Critical**: Ledger data corruption or loss, complete payment processing halt, confirmed data breach affecting customer PII or funds, or regulatory reporting breach. Triggers: automatic all-hands page to on-call, CTO + CEO paged within 10 minutes, dedicated bridge call within 5 minutes, status-page update within 15 minutes, regulator notification assessment within 1 hour, customer comms within 30 minutes.
- **SEV2 – High**: Payment processing degraded >30% throughput or >5% error rate, single-region failover failure, ledger read-only mode, or any condition likely to breach the 99.95% SLA within the current window. Triggers: primary + secondary on-call paged, incident commander assigned within 10 minutes, bridge call within 15 minutes, status-page update within 30 minutes, VP Engineering notified within 20 minutes.
- **SEV3 – Medium**: Non-critical service degradation affecting <30% of customers, single non-ledger microservice outage with working fallback, or elevated latency above SLA threshold but below halt. Triggers: primary on-call paged, IM notified within 30 minutes, status-page update within 1 hour if customer-visible, daily standup update.
- **SEV4 – Low**: Degraded internal tooling, minor UX bug with workaround, non-customer-facing alert. Triggers: next-business-day response, ticket created, no page unless on-call agrees.
- **SEV5 – Informational / Noise**: Cosmetic issues, planned-maintenance notifications, alert misfires logged for tuning. Triggers: no page, logged for weekly alert-quality review.
Define **escalation rules**: any SEV3 unresolved after 4 hours auto-escalates to SEV2; any SEV2 unresolved after 2 hours auto-escalates to SEV1. Severity can be **downgraded** only by the incident commander with IMC notification.
Publish the taxonomy as a one-page decision tree, a Slack slash-command (`/sev`), and an integration into the alerting tool so that every alert carries a suggested severity.
4. Define Incident Roles and Staffing Model (depends on: 3)
Codify four mandatory roles for every SEV1/SEV2 incident and optional roles for SEV3, then solve the 24×7 staffing problem across 28 teams.
**Roles**
- **Incident Commander (IC)**: owns the incident end-to-end, declares severity, assigns tasks, authorizes mitigations, decides when to escalate or stand down. Never writes code during the incident.
- **Communications Lead (CL)**: owns status-page updates, internal Slack channels, account-manager briefings, and regulator notifications. Separate from the IC so the IC can focus on mitigation.
- **Scribe / Timeline Keeper**: logs every decision, action, and timestamp in the incident channel and the incident-management tool. Produces the raw timeline for the postmortem.
- **Subject-Matter Responders (SMRs)**: 1–3 engineers from the owning team(s) who diagnose and fix. For the shared PostgreSQL ledger, a dedicated DBA responder is always required.
**24×7 Staffing via a Three-Tier Follow-the-Sun Model**
- **Tier 1 – Front-line on-call**: Primary + secondary responder per team, paged first. Covers the team's own services.
- **Tier 2 – Platform / SRE on-call**: A dedicated 6-person SRE rotation covering cross-cutting infrastructure: Kubernetes, the shared PostgreSQL ledger, networking, and the two AWS regions. This tier directly addresses the 'carrying a pager for other teams' concern by absorbing infrastructure incidents.
- **Tier 3 – IMC escalation**: Engineering managers and the IMO on-call for multi-team or SEV1 incidents. Provides the IC and CL when no team-level IC is available.
**Follow-the-Sun**: If any engineering hub exists in a second timezone, use it for overnight Tier 1 coverage. If not, partner with a managed on-call service for overnight first-response triage (severity declaration + paging the correct team), reducing 3 a.m. pages for NY-based engineers.
**IC and CL pools**: Nominate at least 2 ICs and 1 CL per team (56 ICs, 28 CLs minimum). ICs are trained and certified before they rotate. For SEV1 incidents, the IC must be a certified IC from the IMC pool, not just 'whoever is around'.
**Ledger-specific rule**: Because the PostgreSQL ledger is shared, a **Ledger Duty Officer** from the SRE Tier 2 is always on the bridge for any incident touching ledger services, regardless of which team owns the failing microservice.
5. Design On-Call Rotations, Compensation, and Alert-Quality Rules (depends on: 4)
Make on-call sustainable, fairly compensated, and free of alert noise so engineers stop resisting participation.
**Rotation Design**
- 7-day rotations, one primary + one secondary per team per week. No engineer is on-call more than one week in four.
- Minimum 48-hour rest between rotations. No on-call during approved PTO.
- All 28 teams participate. Teams without current on-call get a 90-day ramp with a shadow rotation before going live.
- Tier 2 SRE rotation: 6 engineers, one week on / five weeks off, with a dedicated backup.
**Compensation Package**
- **Base on-call stipend**: $500 per week of primary on-call, $250 for secondary, paid regardless of whether pages fire.
- **Page pay**: $75 per acknowledged page outside business hours; $150 if the page leads to active incident work.
- **Time-off-in-lieu (TOIL)**: Any engineer who works >4 hours overnight (00:00–06:00 local) gets a full TOIL day. >8 hours in a single incident gets 1.5 TOIL days.
- **SEV1 bonus**: $300 flat bonus for every engineer who actively works a SEV1 incident, paid within the next pay cycle.
- **Annual on-call cap**: No engineer exceeds 13 weeks of on-call per year. Exceeding the cap triggers a mandatory team-staffing review.
- Budget estimate: ~$420K/year for stipends and page pay across 28 teams; present this to the CFO as a fraction of the $1.3M annual SLA credit cost.
**Alert-Quality Rules (the '85% noise' problem)**
- Every alert must carry: owning team, suggested severity, runbook link, and a 30-day noise score.
- **Alert budget**: each team gets a maximum of 100 actionable alerts per month. Exceeding the budget triggers a mandatory alert-tuning session with the IMO.
- **Noise threshold**: any alert that fires >10 times in 7 days with no human action is auto-flagged for suppression or tuning within 14 days.
- **Alert review cadence**: weekly 30-minute alert-quality review per team; monthly cross-team alert review in the IMC.
- **Sunset rule**: alerts with no runbook are demoted to SEV5 after 30 days and suppressed after 60 days unless a runbook is written.
- Target: reduce monthly alert volume from 3,400 to <600 actionable alerts within 6 months.
6. Build Detection, Escalation, and Communication Paths (depends on: 3, 4)
Eliminate the 22-minute median detection gap and the 40% customer-first-detection rate with layered monitoring and a single escalation spine.
**Detection Layers**
- **Synthetic transactions**: run a payment end-to-end through the full stack (API → service → ledger → confirmation) every 60 seconds from both AWS regions. Alert if latency >2× baseline or any step fails. This catches what per-service metrics miss.
- **Customer-traffic anomaly detection**: monitor API error rates, payment success rates, and latency percentiles per customer cohort. Alert on >2σ deviation.
- **SLO-based alerting**: define SLIs for the 99.95% SLA (availability, latency p99, ledger consistency). Alert when error budget burn rate exceeds threshold, before the SLA actually breaches.
- **Infrastructure health**: Kubernetes node/pod health, PostgreSQL replication lag, disk I/O, and cross-region latency.
- **Support-ticket spike detection**: if >5 customers open tickets about the same symptom within 10 minutes, auto-create a SEV3 candidate.
**Escalation Path**
- Alert fires → PagerDuty routes to Tier 1 primary → 5-min no-ack → Tier 1 secondary → 10-min no-ack → Tier 2 SRE → 15-min no-ack → IMC on-call manager → 20-min no-ack → VP Engineering auto-page.
- Any SEV1 declaration auto-pages the CTO, opens a dedicated Slack channel + Zoom bridge, and notifies the CL.
- **No incident goes unowned for >15 minutes.** If no IC is assigned by minute 15, the IMC on-call manager assumes IC role by default.
**Internal Communications**
- Dedicated Slack channels: `#inc-sev1`, `#inc-sev2`, `#inc-sev3` (auto-created per incident), plus `#inc-updates` for broadcast.
- IC posts a structured update every 15 minutes (SEV1), 30 minutes (SEV2), 1 hour (SEV3) into the incident channel.
- CL posts a summary to `#inc-updates` and notifies relevant engineering managers.
**Customer Communications**
- **Status page**: auto-updated via API. SEV1: first update within 15 minutes, then every 30 minutes until resolved. SEV2: first update within 30 minutes, then every hour. SEV3: within 1 hour if customer-visible.
- **Account managers**: CL briefs AMs via a dedicated Slack channel within 30 minutes (SEV1) or 1 hour (SEV2). AMs contact their top-20 revenue accounts directly.
- **Customer email/SMS**: for SEV1 and SEV2, automated notification to all 2,100 customers via the status-page subscription system within 30 minutes.
- **Regulator notification**: Legal/Compliance assesses within 1 hour whether NY DFS, FinCEN, or card-network notification is required. If yes, file within the regulatory deadline (typically 72 hours for NY DFS cybersecurity events). Log the decision and filing in the incident record.
**Post-resolution**: CL publishes a 'resolved' update within 30 minutes of mitigation. For SEV1/SEV2, a preliminary customer-facing RCA summary is published within 5 business days.
7. Standardize Postmortems with Tracking and Accountability (depends on: 3, 6)
Fix the '11 of 64 action items closed' problem with a mandatory, uniform, blameless postmortem process backed by engineering-manager accountability.
**When Mandatory**
- All SEV1 and SEV2 incidents: postmortem required within 5 business days.
- SEV3 incidents: postmortem required if >50 customers affected, if the incident lasted >4 hours, or if it was customer-detected.
- SEV4/SEV5: optional, but any recurring SEV4 (≥3 times in 30 days) triggers a mandatory review.
**Format (single template, enforced by tooling)**
- Incident summary (severity, duration, customers affected, revenue impact, SLA credit exposure).
- Timeline (auto-generated from scribe notes + alert timestamps).
- Detection analysis: how was it detected, why did it take X minutes, could it have been faster.
- Root cause analysis using 5-Whys or fault-tree, not blame.
- Contributing factors (process, tooling, staffing, knowledge gaps).
- Impact quantification: customers affected, transactions failed, SLA credits triggered.
- Action items: each with a **named owner**, **due date**, **priority**, and **ticket in Jira**.
- Lessons learned and what went well.
**Blameless Review Meeting**
- Held within 5 business days, facilitated by the IMO or a trained facilitator (never the IC of that incident).
- All responders, the CL, relevant engineering managers, and an IMC representative attend.
- Ground rules: focus on system and process failures, not individual mistakes. The facilitator enforces this.
- Meeting recorded; notes published to the engineering-wide wiki within 48 hours.
**Action-Item Tracking and Accountability**
- Every action item is created as a Jira ticket with a due date and a named owner.
- **Engineering managers are accountable**: action-item completion is a standing agenda item in the bi-weekly IMC meeting. Any action item >7 days overdue is escalated to the VP Engineering.
- **Completion gate**: no team may close its postmortem until 100% of its action items have Jira tickets. Postmortem is 'closed' only when all tickets are resolved.
- **Quarterly audit**: the IMO audits action-item completion rates and reports to the IMC and the executive sponsor. Target: >90% completion within 30 days of the postmortem.
- Link postmortem quality and action-item completion to team health scores and engineering-manager performance reviews.
8. Define KPIs, Dashboards, and Governance Reviews (depends on: 7)
Create a measurable feedback loop so leadership can see whether the program is working and where to intervene.
**Primary KPIs (tracked weekly, reported monthly)**
- **MTTD** (Median Time to Detect): target <5 minutes (from 22).
- **MTTA** (Median Time to Acknowledge): target <5 minutes.
- **MTTM** (Median Time to Mitigate): target <60 minutes for SEV1 (from 3 h 10 min), <4 hours for SEV2.
- **Customer-first detection rate**: target <10% (from 40%).
- **SLA compliance**: maintain 99.95%; track monthly SLA credit payouts, target <$100K/quarter (from ~$325K/quarter).
- **Alert signal-to-noise ratio**: target >80% actionable (from ~15%).
- **Alert volume**: target <600/month (from 3,400).
- **Postmortem completion rate**: 100% for SEV1/SEV2 within 5 business days.
- **Action-item completion rate**: >90% within 30 days (from ~17%).
- **On-call health**: pages per engineer per week (target <5), TOIL usage, on-call satisfaction survey score.
- **Unowned incident duration**: target 0 incidents with >15 minutes without an IC.
**Dashboards**
- Real-time operational dashboard (Grafana): current incidents, active alerts, on-call roster, SLA error-budget burn.
- Weekly leadership dashboard (auto-generated): KPI trends, open action items, alert-noise report, on-call load distribution.
- Quarterly IMC scorecard per team.
**Review Cadence**
- **Weekly**: IMO publishes KPI snapshot to `#inc-updates`.
- **Bi-weekly IMC**: review open incidents, overdue action items, alert-quality exceptions, and on-call load.
- **Monthly executive review**: CTO presents KPI trends, SLA credit cost, and risk register to the CEO.
- **Quarterly incident-management review**: deep-dive into trends, training gaps, tooling needs, and process improvements. Output fed into the next quarter's roadmap.
9. Consolidate Tooling and Build the Incident Management Platform (depends on: 2, 3)
Replace six alerting tools and ad-hoc status-page updates with a single, integrated incident management stack.
**Target Tool Architecture**
- **Single alerting and on-call platform** (e.g., PagerDuty or Opsgenie): ingest all alerts, apply severity routing, manage on-call schedules, handle escalations, and send pages. Retire the other five tools within 6 months.
- **Observability consolidation**: standardize on one APM/metrics stack (e.g., Datadog or Grafana Cloud) for all 180 Kubernetes services across both AWS regions. Ensure the shared PostgreSQL ledger has dedicated dashboards.
- **Status page**: a dedicated, branded status page (e.g., Statuspage.io or Instatus) with API integration for auto-updates. Subscribe all 2,100 customers.
- **Incident coordination tool**: integrate incident-management workflows into Slack (auto-create channels, invite responders, post templates) and a dedicated incident record system (e.g., Jira Service Management, incident.io, or Rootly) for timelines, postmortems, and action-item tracking.
- **Runbook repository**: a central wiki (Confluence or Notion) with mandatory runbooks for every alert. No alert goes live without a linked runbook.
**Implementation Tasks**
- Migrate all 28 teams' alert rules into the single platform in three waves (highest-volume teams first).
- Build the severity-based routing rules and escalation policies per S3 and S6.
- Automate status-page updates triggered by severity declaration.
- Build the synthetic-transaction monitor and SLO-based alerting per S6.
- Integrate Jira for automatic action-item ticket creation from postmortems.
- Decommission legacy tools only after all teams have completed training on the new stack.
- Budget: allocate $150K–$250K/year for licensing, plus engineering time for migration.
10. Prepare for the SOC 2 Type II Audit (depends on: 7, 8, 9)
Ensure the incident management process produces the evidence the auditor will need, well before the audit window opens in eight months.
**SOC 2 Requirements to Address (CC7.3, CC7.4, CC7.5)**
- Documented incident response procedures (the severity taxonomy, role definitions, communication templates).
- Evidence of incident detection, response, and recovery for every SEV1/SEV2 incident during the audit period.
- Postmortem records with action-item tracking.
- On-call schedules, training records, and escalation evidence.
- Status-page update logs and customer notification records.
- Regulator notification logs (if any).
**Preparation Tasks**
- The IMO maintains a **SOC 2 evidence folder**: every incident record, postmortem, action-item ticket, status-page update, and training completion certificate is stored and indexed.
- Conduct a **mock SOC 2 audit** at month 5: an internal or external auditor reviews the incident management process end-to-end and identifies gaps.
- Remediate mock-audit findings before month 7.
- Ensure the incident management tool retains all records for at least 12 months (the SOC 2 Type II observation window).
- Document the **chain of custody** for incident records: who accessed, modified, or closed each record.
- Prepare a **narrative document** describing the incident management process, roles, and controls for the auditor.
- Coordinate with the compliance lead to align incident management evidence with the broader SOC 2 scope (access controls, change management, etc.).
11. Design and Deliver Training, Runbooks, and Change Management (depends on: 4, 5, 9)
Equip all 260 engineers, 28 team leads, account managers, and support staff with the knowledge and muscle memory to execute the new process.
**Training Tracks**
- **All 260 engineers** (2-hour session): severity taxonomy, how to acknowledge a page, how to join an incident bridge, how to hand off to an IC, and how to write a postmortem contribution. Delivered in team-level sessions over 4 weeks.
- **IC pool (56+ engineers)** (8-hour certification): incident command techniques, severity declaration, escalation decision-making, bridge facilitation, and blameless postmortem facilitation. Includes two tabletop exercises. Certification valid for 12 months, renewed annually.
- **CL pool (28+ staff)** (4-hour session): status-page writing, customer communication templates, regulator notification triggers, and AM briefing protocol.
- **Account managers and support staff** (1-hour session): how to read the status page, how to escalate a customer report into an incident, and what information to collect.
- **SRE Tier 2** (16-hour onboarding): Kubernetes and PostgreSQL ledger deep-dive, cross-region failover runbooks, and escalation authority.
**Runbooks**
- Every alert must have a runbook before it is routed to on-call. The IMO provides a runbook template and audits compliance weekly.
- Priority runbooks to write first: shared PostgreSQL ledger failover, Kubernetes cluster degradation, payment-processing pipeline failure, cross-region failover, and ledger data-integrity check.
- Runbooks are peer-reviewed and version-controlled.
**Change Management for Adoption**
- Address the 'carrying a pager for other teams' concern directly: publish an FAQ explaining the three-tier model, the SRE Tier 2 absorbing cross-team infrastructure, the compensation package, and the TOIL policy.
- Run **office hours** weekly for the first 8 weeks where any engineer can ask questions or raise concerns.
- Identify **team champions**: one engineer per team who volunteers as an early adopter and peer mentor.
- Publish a **weekly 'incident management newsletter'** during rollout: what changed, what improved, KPI trends, and success stories.
- Make on-call participation a documented expectation in job descriptions and performance reviews for production-facing roles.
12. Execute Phased Rollout, Tabletop Exercises, and Continuous Improvement (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11)
Introduce the process in three waves so teams are not overwhelmed, then validate with exercises and iterate continuously.
**Phase 1 – Weeks 1–6: Foundation**
- Publish the severity taxonomy, role definitions, and communication protocols (S3, S4, S6).
- Launch the single alerting platform for the 12 teams already on-call; begin migration for the other 16.
- Activate the SRE Tier 2 rotation for the shared PostgreSQL ledger and cross-cutting infrastructure.
- Deploy the status page and test the API integration.
- Begin IC and CL training (first cohort of 20 ICs, 10 CLs).
- Publish the on-call compensation package; HR integrates stipends into payroll.
- Write the top 10 priority runbooks.
**Phase 2 – Weeks 7–14: Expansion**
- All 28 teams live on the single alerting platform; legacy tools in read-only mode.
- All 28 teams on the on-call rotation schedule (the 16 new teams in shadow mode for the first 4 weeks).
- Second and third IC/CL training cohorts completed.
- First **tabletop exercise**: simulate a SEV1 ledger corruption scenario with all roles, test the escalation path, status-page updates, and AM briefings. Debrief and fix gaps.
- Postmortem template and Jira integration live; all new incidents use the standard process.
- Alert-tuning sprint: each team reduces its alert volume by 50%.
**Phase 3 – Weeks 15–24: Optimization**
- All teams fully live; legacy alerting tools decommissioned.
- Second **tabletop exercise**: simulate a SEV1 cross-region failure with regulator notification.
- First quarterly IMC review with full KPI dashboard.
- Mock SOC 2 audit (month 5) and remediation.
- Retrospective on the rollout: survey all 260 engineers for feedback, adjust compensation or rotation rules if needed.
- Establish the **continuous improvement cadence**: quarterly process review, annual severity-taxonomy review, and annual on-call compensation benchmarking.
**Ongoing Governance**
- The IMC owns the process document and approves changes.
- The IMO tracks all KPIs and reports to the CTO monthly.
- Any process change requires IMC approval and a 2-week notice period before enforcement.
- Annual external benchmarking against peer B2B payments platforms.
13. Establish Ongoing Governance, Annual Review, and Audit Readiness Cycle (depends on: 12)
Embed incident management as a permanent organizational capability, not a one-time project.
- **Annual process review**: the IMC reviews the severity taxonomy, role definitions, on-call structure, and compensation against industry benchmarks and internal KPIs. Update as needed.
- **Bi-annual tabletop exercises**: one SEV1 infrastructure scenario, one SEV1 data-breach/regulator scenario. Rotate the IC and CL assignments so everyone gets practice.
- **Quarterly alert-quality audit**: the IMO reviews alert volumes, noise rates, and runbook coverage across all 28 teams.
- **On-call health survey**: quarterly anonymous survey measuring burnout, fairness, and compensation satisfaction. Results reviewed by the IMC.
- **SOC 2 readiness cycle**: begin evidence collection immediately after each audit ends. The IMO maintains a rolling evidence folder. Mock audit at month 5 of every 12-month audit cycle.
- **Postmortem maturity tracking**: track the action-item completion rate monthly. If it drops below 80%, the VP Engineering intervenes.
- **Incident management maturity model**: adopt a 5-level maturity model (ad-hoc → defined → managed → optimized → predictive). Assess annually. Target: Level 3 within 12 months, Level 4 within 24 months.
- **Budget review**: annually review on-call compensation, tooling costs, and training budget against the reduction in SLA credits and incident frequency.
Previous Proposal 4 (ID: 31fbbc6c-8c88-4abb-9c6d-b67706d9da75, Agent: grok4.6_initial_4, LLM: xai/grok-4.6):
Estimated Complexity: high
Success Metrics: - Median time to detect customer-impacting incidents ≤ 5 minutes within 6 months of go-live.
- Share of SEV-1/SEV-2 incidents first detected by customers ≤ 5% (from 40%).
- Median time to mitigate SEV-1/SEV-2 ≤ 45 minutes (from 3 h 10 min).
- Named Incident Commander assigned within 5 minutes for ≥ 95% of SEV-1/SEV-2.
- First status-page update within policy time for ≥ 95% of SEV-1/SEV-2.
- SLA credits down ≥ 80% versus the trailing $1.3M within 12 months.
- Paging volume ≤ 500 per month and noise ≤ 15% (from 3,400 and 85%).
- 100% of production services have a named owning team and a paging policy.
- Postmortems filed within 5 business days for 100% of SEV-1/SEV-2; action-item close rate ≥ 80% within 30 days.
- 24×7 IC and critical-path coverage with zero unfilled shifts per quarter.
- Paid on-call live for every rotation before that rotation pages humans.
- SOC 2 Type II incident-response controls evidenced for ≥ 5 months before the auditor's report.
- On-call pulse: ≥ 70% of engineers agree rotations are fair and limited to their services.
Steps (30):
1. Secure executive mandate and budget
Get a written CEO/CTO mandate that incident command is a company process, not a team hobby.
The mandate must state that **paid on-call** is required for production ownership. "No pager for other teams' code" is solved by named ownership, not by refusing coverage.
- Approve budget for tooling, stipends, training, and a dedicated program lead for six months.
- Name an executive sponsor (CTO or VP Engineering) who will chair the weekly incident review.
- Tie the clock to SOC 2 Type II: the process must be live in about 10 weeks so ~6 months of evidence remain.
- Commit that the CEO will hear about outages from this process, not from customers.
2. Form the working group and decision rights (depends on: 1)
Stand up a small group that can decide. Do not form a 28-team committee.
**Core seats:** SRE/platform lead, payments/ledger engineering manager, support lead, legal/compliance, HR, one rotating team EM, and a program manager.
- Meet twice a week for 10 weeks, then weekly.
- RACI: the group proposes; the sponsor decides in 48 hours; teams implement.
- Publish one Slack channel and one source-of-truth doc on day one.
- Time-box design to four weeks. Ship v1 rather than wait for consensus.
3. Inventory services, owners, and on-call gaps (depends on: 1)
Build a living catalog of all ~180 services: owning team, criticality, current on-call, alert sources, and runbook link.
Walk the last 31 customer-impacting incidents and the two events where **nobody was in charge**. Record who detected, who led, time to mitigate, and which alerts fired.
- Tag each service as critical-path, customer-visible, or internal.
- List the 16 teams with no on-call and every orphan service with no owner.
- Map all six alerting tools and the 3,400 monthly alerts onto services.
- Flag the shared PostgreSQL ledger and two-region failover as named special cases.
4. Map regulatory and contractual notification duties (depends on: 1)
Legal and compliance list every duty an incident can trigger. Do not invent clocks that violate a contract.
Cover SOC 2 CC7, customer MSA/SLA credit terms, money-transmitter rules, NYDFS 23 NYCRR 500 if applicable, PCI if in scope, and breach clocks.
- Extract **notification timings** from the largest customer contracts (status page, named AM, written notice).
- Define when legal, regulators, insurers, or the board must be told.
- Feed these clocks into severity triggers and the communications playbook.
5. Approve paid on-call and incident pay (depends on: 1, 4)
Unpaid on-call is why 16 teams refuse the pager and why nights are uncovered. Fix the money before asking for coverage.
HR, legal, and finance design a New York–compliant package: weekly stipend for primary and secondary, extra stipend for company IC and comms, and **after-hours incident pay** or comp time.
- Treat exempt vs non-exempt staff explicitly under NY wage-hour rules.
- Put stipend in the next pay cycle after policy publish, not "later."
- Cap consecutive night weeks. Fund hiring if a team cannot rotate fairly (minimum six people for 24x7 primary plus secondary).
- Publish the package before any new rotation starts. This is the main answer to pager pushback.
6. Ratify four severity levels and their triggers (depends on: 3, 4)
Adopt a business-impact scale. Engineers do not invent severity in the moment.
**SEV-1:** material payments failure; ledger down or inconsistent; security or customer-data incident; both regions impaired; or many customers already in SLA-credit territory.
**SEV-2:** degraded payments or a contracted feature down for multiple customers; SLA at risk.
**SEV-3:** narrow or single-customer impact with a workaround; no fleet-wide SLA risk.
**SEV-4:** no customer impact; ticket only.
- SEV-1 pages IC, comms, scribe, owning SMEs, and an exec; war room in 5 minutes; status page in 10; AM outreach in 20.
- SEV-2 pages IC and owning SMEs; comms may be the IC; status page in 15 minutes; updates every 30 minutes.
- SEV-3 pages the owning team only; customer notice only if that customer is affected.
- Anyone may declare. Only the IC may downgrade. When unsure, start high.
7. Define incident roles and their authority (depends on: 6)
Four roles. Separate coordination from debugging so "who is in charge" cannot stall for an hour again.
- **Incident Commander:** owns severity, the room, the clock, and the next action. Does not write code. May page anyone, freeze deploys, and invoke failover. Staffed from a company-wide trained pool, not from the failing team.
- **Communications Lead:** status page, customers, AMs, execs, regulators. Speaks only from IC-approved facts.
- **Scribe:** timeline in the incident tool. Required for SEV-1 and SEV-2.
- **SME responders:** the owning team's on-call. They mitigate. They do not run the room.
Publish a one-page authority card. The IC stays in charge even if a VP joins.
8. Design 24x7 coverage without 28 night rotations (depends on: 3, 5, 7)
Do not put 28 teams on 24x7. That is what engineers are rejecting.
Use **three layers** so people page for their own code, plus a trained commander.
- Layer A — company IC and SEV-1 comms: 24x7; about 24 trained people; week-long primary and secondary.
- Layer B — critical-path team on-call (ledger, payments processing, auth, API edge, platform/Kubernetes, data stores): 24x7 primary plus secondary.
- Layer C — all other teams: business-hours on-call; after hours the IC pages the team EM, who has a written escalation list.
Platform on-call is the safety net for unknown-owner pages, never the permanent owner. Every service must have a named team within 60 days or be scheduled to shut off.
9. Set rotation, handoff, and load rules (depends on: 8)
Write mechanical rules so rotations are fair and load is visible.
Primary week, then secondary week, then at least two weeks off. No one holds primary on two rotations at once.
- Handoff is a 30-minute overlap covering open incidents, silenced alerts, and upcoming changes.
- Page-load SLO: p50 ≤ 4 pages per 12-hour night shift; p95 ≤ 10. A breach opens an alert-quality action.
- Require a shadow week before a first IC shift or a first critical-path rotation.
- Swaps live in the paging tool. Managers own coverage gaps, not the last person on the roster.
10. Write detection and escalation paths (depends on: 6, 9)
Customers currently detect 40% of incidents and median time to detect is 22 minutes. That is the first failure mode.
Detection path: synthetic full-payment probes in both regions, SLO burn-rate alerts, support-to-incident intake, and one customer callback path that can create a SEV.
- A page must be acked in **5 minutes** or it auto-escalates to secondary, then IC, then the EM, then the VP.
- Support may declare SEV-2 or higher without engineering permission.
- If ownership is unclear for 10 minutes, the IC keeps the incident and assigns a temporary owner. Never wait.
- An exec bridge auto-opens for every SEV-1 at T+15 minutes.
11. Write internal, customer, and regulator communications (depends on: 4, 6, 7)
Stop "whoever is around" from writing the status page. Comms follow the clock, not convenience.
Timings from declaration:
- Internal war room: immediate. Exec summary for SEV-1/2 at 15 minutes, then every 30 minutes.
- **Public status page:** SEV-1 in 10 minutes, SEV-2 in 15. Updates at least every 30 minutes until resolve. Templates only. No speculation.
- Account managers get an affected-customer list and a script at T+20 minutes for SEV-1/2.
- Resolve notice and credit assessment within one business day.
- Comms pages legal on SEV-1 security, ledger integrity, or any outage that will breach contractual notice. Legal owns outbound regulatory letters; the IC owns facts.
12. Standardize blameless postmortems and action tracking (depends on: 6)
A written postmortem is mandatory for every SEV-1 and SEV-2 within 5 business days. SEV-3 if the IC or EM requests it.
Use one template: timeline, customer impact (volume, duration, credits), detection gap, what went well, what did not, process-focused five whys, and numbered actions with owner and due date.
- Review is **blameless** and scheduled. The IC attends. The exec sponsor reads every SEV-1.
- Actions live in one tracker, not in the doc. No action without an owner and a date. Default due date 14 days; 30 days max unless architecture work with a milestone.
- Close rate is a published metric. The old 11-of-64 pattern is a process failure.
13. Set alert quality rules that make paging acceptable (depends on: 3, 6)
3,400 alerts a month and 85% noise is why on-call feels like punishment. Pages are a product with a quality bar.
A page (not a ticket) must map to a customer-facing SLO or a hard dependency of one. It must have an owner team, a runbook link, and a default severity. It must be actionable at 3 a.m. by the person who is paged.
- Ban parallel paging from six tools. One paging policy: symptom-based; burn-rate preferred over raw thresholds.
- Every team gets a monthly noise budget. Exceeding it is a sprint task, not heroics.
- A human may silence a flapping alert only with a linked ticket.
14. Publish Incident Management Policy v1 (depends on: 5, 6, 7, 8, 9, 10, 11, 12, 13)
Collapse the design into a short policy people will open during an outage.
Ten pages or fewer, plus one-page cards for severity, roles, and comms timings. Host it where the incident tool can link it.
- Include the compensation summary and the rule: you are **not on-call for other teams' services**.
- Version it. v1 is mandatory from the pilot start date.
- Legal, HR, and the exec sponsor sign. Announce in all-hands, not only in Slack.
15. Implement a single incident command tool (depends on: 7, 10, 11)
Put one tool in the path that creates the room, pages roles from severity, records the timeline, and prompts status-page updates.
Requirements: Slack (or equivalent) incident bot, severity in one click, role assignment, stakeholder groups, and timeline export for postmortems and auditors.
- Integrate with the pager so IC, comms, and SME pages are automatic.
- Retain artifacts at least one year for SOC 2.
- Ad-hoc Zoom/Slack threads are no longer the system of record.
16. Consolidate six alerting tools onto one pager (depends on: 9, 13)
Pick one paging product. Connect existing monitors to it. Migrate **pages** first, tickets second.
- Inventory every page-producing rule. Delete or downgrade the noisy majority in S23.
- Route by service label → owning team schedule → the escalation policy from S10.
- IC and comms schedules live in the same product.
- Set a hard date after which pages outside the chosen tool are not valid on-call obligations.
17. Operationalize the status page and AM path (depends on: 11, 15)
Put the status page behind the Comms Lead role. Use templates for investigating, identified, mitigating, and resolved.
Subscribe AMs and customers whose contracts require it. Generate the affected-customer list from the incident (tenant, region, payment method).
- Dry-run a SEV-2 update before the pilot goes live.
- Record every public update in the incident timeline for the audit.
- Define partial vs full-outage wording so impact cannot be understated.
18. Stand up one postmortem repo and action board (depends on: 12)
Create the template, the filing location, and one Jira/Linear board with states: open, in progress, blocked, done, and won't-do (with exec reason).
Wire the incident tool so a SEV-1/2 automatically opens a draft postmortem and action tickets.
- Each week the program manager reports actions past due to the exec sponsor.
- If it is not in the board, it does not exist.
19. Assign owners and write critical-path runbooks (depends on: 3, 9)
Close the ownership gaps that cause "pager for other people's code."
Every production service gets a team in the catalog. Unowned services get an owner in 30 days or a decommission date.
- Write runbooks for ledger Postgres, regional failover, payments API, auth, and the Kubernetes control plane: symptoms, dashboards, mitigate vs escalate, customer impact.
- Link runbooks from alerts. If there is no runbook, the alert cannot page at night unless the EM accepts the gap in writing.
20. Detect payments failures before customers do (depends on: 3, 13)
Build or finish end-to-end synthetics: create payment, ledger write, webhook, both regions, both critical payment methods.
Alert on SLO burn, not on a single 500. Page SEV-2 or SEV-1 from these probes. This is the fastest lever on the 22-minute MTTD and the 40% customer-detected rate.
- Add ledger lag, replication, disk, and failover-readiness as first-class pages to the ledger team.
- Every postmortem asks first: **why did a customer see this first?**
21. Train the first cadre of ICs, comms leads, and scribes (depends on: 7, 14, 15)
Train about 24 ICs and 12 comms leads before the pilot. Classroom plus a recorded shadow of a simulated SEV-1.
Curriculum: severity, authority card, tool, comms timings, when to call legal, how to run a room of 20, how to hand off at 2 a.m.
- Certification: pass a tabletop. No certificate, no rotation.
- Recertify yearly and after any SEV-1 where process failed.
- Managers of ICs protect calendar time. This is part of the job.
22. Pilot on the payments critical path for six weeks (depends on: 5, 14, 15, 16, 17, 18, 19, 21)
Go live with policy, tool, paid rotations, and IC coverage for ledger, payments, API, platform, and support intake.
Keep old paths as backup for one week, then cut over. Real incidents use the new process only.
- Staff the program lead in every SEV-2+ as coach, not as secret IC.
- Collect friction daily. Fix tooling and wording in 48 hours.
- Expansion gate: named IC in under 5 minutes, first status update on time, no unpaid pages, postmortem filed.
23. Cut alert noise with a forced burn-down (depends on: 13, 16)
Give every team a numbered list of their noisiest alerts. Move noise from 85% to **under 15%**, and monthly pages from 3,400 toward 500.
Each sprint, critical-path teams must delete, debounce, or convert to ticket a fixed quota. Platform provides burn-rate and grouping libraries.
- Publish a weekly noise leaderboard. Shame systems, not people.
- After eight weeks, any page without a runbook or with >30% false pages in 14 days is auto-downgraded until fixed.
24. Roll remaining teams onto the model by risk (depends on: 22)
After the pilot gate, add teams in waves of four to five every two weeks. Highest customer-impact first.
Layer C teams get business-hours schedules and the EM night list. Do not surprise anyone with a pager.
- Each wave: ownership confirmed, alerts routed, runbooks for paging alerts, paid rotation in HR, one tabletop.
- Finish all 28 teams at least five months before the SOC 2 report date so the observation window covers the company.
- Orphan services still unowned at wave end are escalated to the sponsor for shutdown or reassignment.
25. Run tabletops and multi-region game days (depends on: 21, 22)
Schedule a monthly tabletop: SEV-1 ledger, SEV-1 region loss, SEV-2 degraded payments, customer-detected incident, and a "who is in charge" chaos drill.
Quarterly game day: fail a region or a ledger replica in staging or a controlled production drill.
- Include support, AMs, legal, and an exec. Process fails if only engineers show up.
- Capture actions on the same board as real postmortems.
- Use results as SOC 2 evidence that IR is tested.
26. Resolve ownership fights and pager culture (depends on: 5, 8, 22)
Treat pushback as design input, not defiance. Repeat the contract in office hours: you carry a pager for **your** services; IC is coordination; nights are paid; noise is a defect.
- EMs who cannot staff a fair rotation get headcount or have services reassigned. Do not run two-person 24x7.
- Publicly close the two historical "nobody in charge" incidents with what would be different now.
- Pulse-survey on-call at 60 and 120 days. If load or fairness is red, stop expansion until fixed.
27. Launch metrics, weekly review, and error budgets (depends on: 15, 22)
The CEO email exists because there was no operating rhythm. Stand up a dashboard and use it.
Track incident count by SEV, MTTD, MTTA, MTTM, customer-detected %, time-to-IC, time-to-status, pages per shift, noise %, SLA credits, postmortem SLA, and action close rate.
- **Weekly 30-minute review** chaired by the exec sponsor. Last week's incidents, open actions, alert SLO breaches. Mandatory for EMs of teams that had SEV-2+.
- Error budgets on critical SLOs: burn too fast and the team stops features and pays down reliability.
- Monthly report to the CEO so customer-first news does not recur.
28. Collect SOC 2 Type II evidence as a side effect (depends on: 4, 14, 22, 27)
Do not create a parallel "audit process." The real process is the evidence.
From pilot start, retain policy versions, on-call schedules, incident records, status-page history, postmortems, action tickets, training attendance, tabletop notes, and access reviews of the incident and pager tools.
- Map each artifact to SOC 2 CC7.2–CC7.5. Compliance owns the matrix; engineering owns completeness.
- Internal audit or a consultant runs readiness reviews at week 12 and week 20, not the week before the auditor.
29. Inspect and adapt the process after 90 days (depends on: 22, 27)
After 90 days live, revise v1 using data, not opinions.
Look at MTTD/MTTM, uncovered shifts, IC burnout, missed status updates, action close rate, and survey results.
- Change severity definitions if teams inflate or deflate them.
- Adjust Layer B vs C membership from real pages.
- Cut steps nobody follows. Add only what the last 90 days proved missing.
- Freeze v2 as the SOC 2 described process unless a defect is found.
30. Rehearse the audit and freeze the observation window (depends on: 24, 28, 29)
Six to eight weeks before the auditor, run a mock interview with an IC, a random engineer, support, and compliance.
Show one SEV-1 and one SEV-2 end-to-end: detect, page, command, comms, postmortem, closed actions.
- Fix evidence gaps immediately.
- Freeze process wording for the rest of the observation window; log exceptions.
- Brief the CEO and customer success with the metrics so the "outages we hear about from clients" story is retired.
Previous Proposal 5 (ID: 53260d57-1d36-433f-83b4-7205e06869b2, Agent: deepseek-v4-pro_initial_5, LLM: deepseek/deepseek-v4-pro):
Estimated Complexity: high
Success Metrics: - Customer-detected incidents decrease from 40% to less than 15% within six months.
- Median time to detect (MTTD) is under 5 minutes for SEV1 and SEV2 incidents.
- Median time to mitigate (MTTM) is under 60 minutes for SEV1 and under 2 hours for SEV2.
- Alert noise decreases from 85% to below 10% within three months.
- 100% of SEV1 and SEV2 incidents have a completed blameless postmortem within 5 business days.
- 100% of postmortem action items are tracked with an owner and due date; 90% are completed on time.
- 24x7 on-call coverage achieved across all 28 teams with no unpaid on-call.
- 99% of pages are acknowledged within 5 minutes.
- SLA credits paid reduce by at least 50% over the next 12 months.
- SOC 2 readiness: all incident response controls are documented, tested, and evidence is produced by month 7.
Steps (14):
1. Baseline current incident response and align stakeholders
Collect data from the last 12 months of incidents, all six alert tools, on-call practices, and team interviews. Identify gaps against the target incident management process and secure executive sponsorship.
- Gather incident timeline, detection source, mitigation time, customer impact, SLA credit, and postmortem status for all 31 incidents.
- Survey 28 teams on on-call burden, alert quality, and operational pain.
- Map current tools, escalation paths, and communication workflows.
- Create baseline metrics and a stakeholder map with executive sponsor and audit owner.
2. Define severity levels and response triggers (depends on: 1)
Define a four-level severity scale with objective business-impact criteria so any engineer can classify an incident consistently.
- SEV1: widespread transaction processing outage, data breach, security incident, or severe SLA breach; triggers full incident command, executive notification, and a 5-minute status page update.
- SEV2: major feature outage, significant degradation without workaround, or customer financial risk; triggers incident commander, full communications role, and status page updates.
- SEV3: partial impairment with workaround or limited customer impact; triggers on-call response, internal communication, and optional status page update.
- SEV4: minor or internal issue, no customer impact; handled during business hours through ticketing.
- Include an escalation matrix showing who can declare, downgrade, and invoke regulatory or legal involvement.
3. Define incident roles and decision authority (depends on: 1, 2)
Define incident roles, responsibilities, and decision authority using RACI to remove ambiguity about who is in charge.
- Incident Commander: owns the incident, declares severity, and coordinates resolution.
- Communications Lead: owns internal and external messaging, status page updates, and account manager notifications.
- Scribe: maintains timeline, incident log, and postmortem notes.
- Subject-matter responders: diagnose and fix the incident; may come from multiple teams.
- Executive sponsor: optional for SEV1; customer liaison: handles account managers.
- Define decision rights for severity declaration, escalation, rollback, customer communications, and incident closure.
4. Design 24x7 staffing model across 28 teams (depends on: 3)
Design 24x7 coverage across 28 teams without overloading engineers. Use service-based on-call plus a central incident command pool.
- Each service or domain team assigns primary and secondary on-call for its own services.
- Create central incident commander, communications, and scribe rotations staffed from a trained incident response guild across all teams; use follow-the-sun between the two AWS regions and time zones.
- Define escalation layers: service on-call to team lead or manager to service owner to executive.
- Define handoff times, shadow shifts, and load balancing; target at most one week of on-call per engineer per month.
- Bridge the current 12-team paid on-call to 28-team paid coverage; no team remains uncovered.
5. Define on-call rotations, compensation, and alert quality rules (depends on: 4)
Define sustainable rotations, pay, and rules that eliminate noisy pages.
- Rotations: weekly or biweekly, at least one primary and one secondary, with 12-hour shifts where possible or 24-hour for low-volume services.
- Compensation: monthly on-call stipend for all on-call engineers, additional incident response bonus for after-hours work, and time off in lieu; align with market rates.
- Alert quality rules: every page must be actionable, have a runbook link, specify a service owner, include severity, and be based on SLO burn or known failure signals; no dashboard-only alerts.
- Noise budget: reject or downgrade non-actionable alerts; all pages must go to on-call only after suppression and deduplication.
- Weekly alert review removes the top noisy alerts.
6. Design detection, escalation, and alert routing (depends on: 2, 4, 5)
Define how incidents are detected, routed, and escalated so nothing waits on a human to notice.
- Consolidate the six alert tools into one alerting and paging platform with routing by service, severity, and tags.
- Detection sources: infrastructure metrics, application synthetic transactions, log-based anomalies, business transaction SLI monitoring, and customer-reported issues through support or account managers.
- Routing: alert is paged to service on-call within 30 seconds; primary must acknowledge within 5 minutes; if no ack, page secondary then on-call manager.
- Escalation timeouts: unresolved SEV1 escalates to service owner at 15 minutes and to leadership at 30 minutes; any engineer can escalate to the incident commander.
- Define customer-reported incident intake and classification in the same tool.
7. Define internal and external communication protocols (depends on: 2, 3)
Define communication channels, templates, and timing for internal, customer, and regulator audiences.
- Internal: dedicated incident Slack channel, internal status page mirror, and war room bridge for SEV1; incident commander and communications lead own these channels.
- Status page: SEV1 post within 5 minutes, updates every 30 minutes or on material change, resolution within 60 minutes of mitigation; SEV2 post within 15 minutes, updates hourly; SEV3 optional.
- Account managers: SEV1 and SEV2 notify account managers within 15 minutes with an approved customer-facing description and expected impact.
- Regulators: legal or compliance determines notification for data breaches, security incidents, funds availability issues, or regulatory reportable events; criteria and timing follow legal and regulatory requirements; communications lead coordinates.
- Use pre-approved message templates and an approval chain; no ad-hoc wording.
8. Define postmortem policy and action tracking (depends on: 3)
Define mandatory blameless postmortems and action tracking.
- Mandatory for all SEV1 and SEV2 incidents, and any SEV3 that breaches SLA or is customer-detected.
- Format: impact, timeline, root causes, contributing factors, detection and response gaps, what worked well, and action items.
- Blameless: focus on system and process causes, not individual blame; use trained facilitators.
- Ownership: each action has an owner, due date, and tracking ID in a single backlog.
- Review postmortems at the weekly incident review; track action closure; expect 100% completion.
- Complete postmortems within 5 business days for SEV1 and SEV2 incidents.
9. Define metrics, dashboards, and review cadence (depends on: 2, 3, 8)
Define metrics and review cadence to measure process health.
- Metrics: MTTD, MTTM, customer detected percentage, alert noise percentage, on-call response time, on-call load, SLA credits paid, and postmortem action completion.
- Dashboards: real-time operational dashboard for on-call engineers and management.
- Weekly incident review: review all SEV1 and SEV2 incidents, action items, and noisy alerts.
- Monthly trends with leadership; quarterly review against SLOs and audit controls.
- Success thresholds: MTTD under 5 minutes, MTTM under 60 minutes for SEV1, customer detected under 15%, and alert noise under 10%.
10. Configure incident tooling and integrations (depends on: 5, 6, 7, 8, 9)
Implement and integrate the tools that automate the defined process.
- Aggregate alerts from the existing six tools into PagerDuty, Opsgenie, or a similar platform.
- Configure on-call schedules, escalation policies, and paging targeted at service owners.
- Integrate status page API for automated or one-click updates.
- Add Slack commands to declare incidents, start war rooms, assign roles, and post status updates.
- Integrate runbook and service catalog access; create postmortem templates in Jira or Notion with action item tracking.
- Ensure audit trails and role assignments are logged for SOC 2.
11. Pilot with 2-3 volunteer teams and iterate (depends on: 10)
Run a controlled pilot before full rollout to validate and refine the process.
- Select 2-3 volunteer teams with representative services and on-call patterns.
- Run the new severity, roles, on-call, alerting, and communication process for 2 weeks.
- Track metrics and gather feedback from on-call engineers, incident commanders, and communications leads.
- Iterate severity thresholds, alert rules, templates, and runbooks based on findings.
- Exit criteria: no SEV1 without a declared incident commander, alert noise below target, and positive on-call survey results.
12. Train and certify all 28 teams (depends on: 11)
Prepare all 28 teams so the process is not a paper exercise.
- Create role-specific training: all engineers, incident commanders, communications leads, scribes, subject-matter responders, and managers.
- Deliver interactive workshops and online modules.
- Run at least two game-day drills per quarter simulating SEV1 and SEV2 incidents, testing escalation, communication, and postmortem.
- Certify incident commanders and communications leads; only certified staff can hold those roles in rotations.
- Include scenario-specific training for account managers and legal or regulatory handling.
13. Full rollout and enforce process (depends on: 12)
Move every team onto the new process and retire old habits.
- Migrate all 28 teams into consolidated on-call schedules and alert routing.
- Retire old alerting tools or methods that bypass the process.
- Activate status page templates and postmortem action tracking globally.
- Monitor the first 30 days with a daily operational review and quick fixes.
- Communicate rollout status to leadership and all engineers; provide office hours and support.
14. SOC 2 readiness and continuous improvement (depends on: 13)
Map the process to SOC 2 and keep improving after go-live.
- Map each process component, including severity, roles, detection, communications, postmortems, and metrics, to SOC 2 control requirements.
- Produce runbooks, evidence of training, incident logs, postmortems, and action item tracking.
- Conduct an internal readiness assessment at month 6 and remediate gaps before the audit at month 8.
- Maintain weekly and monthly metrics reviews to drive continuous improvement.
- Refresh on-call compensation and alert quality rules based on operational data.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Executive mandate, program funding, and governance", "description": "Convert the CEO email into a **written mandate within 48 hours**. Name one accountable program owner and a small decision group. Time-box design to five weeks so rollout starts well before the audit.\n\n- Appoint the CTO as executive sponsor and a Head of Reliability or Incident Management as program owner with full-time authority.\n- Create an 8–10 person design working group: SRE/platform lead, payments and ledger engineering managers, support lead, compliance, legal, HR, finance, and two rotating engineering managers.\n- Approve budget lines: tooling consolidation, on-call compensation, training, coaching, and 2–3 dedicated program staff.\n- State the non-negotiables: paid on-call, named service ownership, one severity scale, one paging platform, mandatory postmortems, and protected engineering capacity for reliability actions.\n- Fix the master timeline: interim controls week 1, design weeks 1–5, pilot weeks 6–11, full rollout weeks 12–22, internal audit rehearsal month 6, audit-ready month 7.\n- Publish a one-page charter to the company that says incident response is a company operating process, not an optional team practice.", "dependencies": []}, {"step_id": "S2", "title": "Evidence baseline from incidents, alerts, and coverage gaps", "description": "Before changing anything, create an auditable **before picture** from the last 12 months. This baseline drives the severity design, staffing model, and executive reporting.\n\n- Reconstruct all 31 customer-impacting incidents: detection source, first owner, severity, mitigation time, credits paid, and whether command was clear.\n- Write a specific case review of the two incidents with unclear ownership for more than one hour.\n- Inventory the six alert tools, alert volume by team, noise rate, rules with no owner, and rules with no runbook.\n- Map the 16 teams without on-call and the 12 teams with unpaid on-call.\n- Review the 53 open postmortem actions and triage the highest-risk items first.\n- Freeze baseline metrics: 22-minute detection, 40% customer-first detection, 3h10 mitigation, $1.3M credits, 3,400 alerts, 85% noise, and 17% action closure.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Stakeholder listening, resistance mapping, and co-design group", "description": "Treat engineer pushback as **design input, not an attitude problem**. The objection to carrying a pager for other teams' code must shape ownership, routing, compensation, and command staffing.\n\n- Interview leads from all 28 teams plus support, customer success, sales, compliance, and HR within two weeks.\n- Separate the real causes of resistance: unpaid night work, unfamiliar systems, poor runbooks, unfair load, fear of blame, or unclear authority.\n- Recruit 10–15 respected engineers and managers as co-designers so the model is built with the teams.\n- Survey baseline on-call sentiment, trust in alerts, and psychological safety; repeat at 90, 180, and 365 days.\n- Document the central promise: responders are paged for services they own, trained incident commanders coordinate, and all on-call work is compensated.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Service ownership catalog, criticality tiers, and dependency map", "description": "No alert can page the right team until every service has a **named owner**. Build the catalog as the routing foundation for on-call, severity impact mapping, status-page components, and audit evidence.\n\n- Assign one accountable team to each of the 180 services, with an engineering manager, Slack channel, escalation policy, dashboard, and runbook link.\n- Tier services: Tier 0 for money movement, ledger integrity, authentication, settlement, and shared PostgreSQL; Tier 1 for customer-facing degradable services; Tier 2 for internal or batch services; Tier 3 for non-critical services.\n- Map customer journeys to services, databases, regions, third parties, and contractual SLA components.\n- Mark orphan and shared services; require ownership, reassignment, or decommissioning within 30 days.\n- Document cross-region dependencies, failover constraints, and services that defeat nominal regional redundancy.\n- Treat missing ownership for Tier 0 or Tier 1 services as an executive escalation and a release-blocking risk.", "dependencies": ["S2"]}, {"step_id": "S5", "title": "Severity scale, declaration rights, and automatic triggers", "description": "Adopt one severity scale so declaration is a lookup, not a debate. **Anyone may declare; only the incident commander may downgrade.** When uncertain, start higher.\n\n- SEV1: money movement stopped or incorrect, ledger integrity in doubt, confirmed security or data event, both regions impaired, or broad SLA-credit exposure.\n- SEV2: major degradation, settlement window at risk, one or more critical customers fully down, or likely contractual breach.\n- SEV3: limited impact with a workaround, no financial-integrity or security risk.\n- SEV4: no customer impact; handled by ticket during business hours.\n- Define fixed triggers for each level: pages, command staffing, bridge call, status-page timing, account-manager outreach, executive notification, and postmortem obligation.\n- Add automatic escalation: SEV3 open more than 2 hours becomes SEV2; any incident touching the shared ledger or unclear ownership becomes SEV2 unless the commander documents otherwise.\n- Publish a decision tree and 10–12 worked examples from the actual 31 incidents.", "dependencies": ["S2", "S4"]}, {"step_id": "S6", "title": "Incident roles, authority, and handover rules", "description": "Solve the **nobody-was-in-charge failure** by making command explicit, trained, and transferable. Separate coordination from technical remediation so responders are not asked to debug unfamiliar code.\n\n- Incident Commander: owns severity, priorities, escalation, mitigation strategy, role assignment, handoffs, and closure. Does not write code during the incident. May freeze deploys, invoke pre-approved failover, and pull any on-call responder.\n- Communications Lead: owns status page, internal updates, account-manager briefs, and coordination with legal or compliance.\n- Scribe: maintains the timestamped record of decisions, actions, role changes, and customer communications.\n- Subject-matter responders: engineers from owning teams who diagnose and remediate within their accepted domain.\n- Executive liaison: required for SEV1; shields the commander from executive questions and owns board or regulator escalation.\n- Require the commander to claim the role within 5 minutes of SEV1 or SEV2 declaration, state it in the incident channel, and record every handover.\n- Allow role combination only for SEV3 or SEV4; require separate people for command, communications, and primary technical work at SEV1.", "dependencies": ["S5"]}, {"step_id": "S7", "title": "Paid on-call, fatigue safeguards, and HR compliance", "description": "Unpaid on-call is a retention, fairness, and legal risk in New York. **Pay must be approved before any new mandatory rotation starts.** This is the fastest way to reduce pager resistance.\n\n- Create a weekly stipend for primary and secondary on-call, differentiated by tier and night responsibility.\n- Add per-page or per-incident pay for out-of-hours activation, plus guaranteed recovery time after night work or major incidents.\n- Create a separate stipend for duty incident commanders and communications leads because command is a heavier burden.\n- Verify FLSA, New York wage-hour, overtime, holiday, and payroll treatment with HR, legal, and finance.\n- Cap rotation frequency: no more than one primary week in four or one week in six depending on staffing; no simultaneous primary assignments.\n- Require at least six trained responders for any 24x7 rotation; fund hiring or service reassignment where teams are too small.\n- Publish the compensation package and payroll start date before asking engineers to join rotations.", "dependencies": ["S3"]}, {"step_id": "S8", "title": "On-call staffing model and rotation rules", "description": "Do not force 28 identical night rotations. Use a **layered model**: central command coverage, critical-path team coverage, and business-hours coverage for lower-tier services.\n\n- Company incident-command and communications rotation: 24x7 pool of 25–35 trained volunteers and designated senior staff, giving primary plus secondary coverage at all times.\n- Critical-path teams: 24x7 primary and secondary on-call for Tier 0 and Tier 1 domains, including payments, ledger, authentication, API edge, Kubernetes platform, and shared PostgreSQL.\n- Other teams: business-hours on-call with a documented night escalation list owned by the engineering manager.\n- Group services into 8–12 coherent domains so rotations are sustainable; no rotation with fewer than six people is approved without an executive exception.\n- Require handoff overlap, shadow shifts before first primary duty, and no on-call during approved PTO.\n- Define cross-team pull rules: the commander may page another team's on-call with a 10-minute acknowledgement obligation; this is coordinated by command, not pushed onto responders.\n- Publish schedules, swap rules, and load limits in the paging tool.", "dependencies": ["S4", "S6", "S7"]}, {"step_id": "S9", "title": "Alert quality standard, page budget, and noise controls", "description": "With 3,400 alerts and 85% noise, detection fails because people stop trusting pages. Make alert quality a **condition of paging anyone**.\n\n- Every page must have a named owning team, customer or SLO impact, severity, dashboard, runbook, expected action, and escalation policy.\n- Page on customer-visible symptoms: payment success rate, latency, settlement deadlines, error-budget burn, ledger integrity, and replication health.\n- Demote cause-based CPU, memory, or infrastructure-only alerts to dashboards or tickets unless they map to a customer journey.\n- Set a page budget: maximum 2 out-of-hours pages per person per week; breach triggers a mandatory alert-tuning sprint.\n- Auto-flag alerts that fire repeatedly without action or have high no-action acknowledgement rates.\n- Require shadow mode for new alerts before they page humans, except for documented emergencies.\n- Review alert quality monthly by team and publish a noise leaderboard.\n- Target fewer than 500 actionable pages per month and above 80% actionability within six months.", "dependencies": ["S2", "S4"]}, {"step_id": "S10", "title": "Detection uplift across payments, ledger, and customer signals", "description": "Customers detected 40% of incidents first. Detection must shift to **payment outcomes, ledger integrity, and inbound customer signals**, not host metrics alone.\n\n- Define SLIs and SLOs for payment initiation, authorization, settlement, reconciliation, refunds, API availability, and reporting freshness.\n- Run synthetic end-to-end payment tests from outside the platform in both AWS regions every 60 seconds.\n- Add continuous ledger assurance: double-entry reconciliation, replication lag, failover readiness, disk pressure, and settlement-window countdown alerts.\n- Monitor per-customer anomalies for top accounts so a single-tenant outage is detected before the account manager calls.\n- Convert high-priority support tickets and account-manager reports into incident candidates within 5 minutes.\n- Track detection source for every incident; make customer-first detection a reviewed defect with a corrective action.\n- Add partner, banking, and card-network notification intake into the same declaration path.", "dependencies": ["S4", "S9"]}, {"step_id": "S11", "title": "Single incident platform and alert-tool consolidation", "description": "Collapse six tools into **one paging and incident-management system** with one queue, one timeline, and one audit trail. Do not create a parallel audit process.\n\n- Select an integrated stack: paging and on-call scheduling, incident workflow, Slack or chat integration, bridge calling, status-page API, and ticketing integration.\n- Implement one-command declaration that creates the incident record, channel, bridge, severity label, role prompts, and clock.\n- Migrate alert sources by wave; retire a legacy paging path only after routing, ownership, and acknowledgement tests pass.\n- Automate evidence capture: timestamps, acknowledgements, role assignments, severity changes, communications, and postmortem links.\n- Integrate with the service catalog, on-call schedules, Jira or equivalent action tracker, customer-success tooling, and the status page.\n- Verify the platform works during a single-region failure, including SMS, phone, and offline fallback paths.\n- Set a hard date after which pages outside the chosen platform are not valid on-call obligations.", "dependencies": ["S5", "S8", "S9"]}, {"step_id": "S12", "title": "Escalation paths and five-minute command rule", "description": "Create one path from **signal to named commander in under five minutes**, any hour of the day. Make escalation automatic and time-bound.\n\n- Accept declarations from automated alerts, engineers, support, account managers, partners, and customers through the same command.\n- For SEV1 and SEV2, page the duty incident commander and owning-team primary immediately.\n- Acknowledgement ladder: primary 5 minutes, secondary 10 minutes, manager 15 minutes, director or executive 20 minutes.\n- If no commander claims the incident within 5 minutes, the platform assigns and announces one; the assignee may hand over but cannot leave the incident unowned.\n- For SEV3, require team acknowledgement within 30 minutes; otherwise create a tracked work item.\n- Give the commander pre-approved authority to invoke regional failover, ledger read-only mode, feature kill switches, and partner notifications without waiting for executive sign-off.\n- Record every missed acknowledgement and escalation failure for weekly review.", "dependencies": ["S6", "S8", "S11"]}, {"step_id": "S13", "title": "Internal communications protocol", "description": "Separate the **working incident channel from the audience channel** so responders can work and executives, support, and sales stay informed without interrupting the commander.\n\n- Create one incident channel and one bridge per SEV1–SEV3 incident; use a read-only broadcast channel for executives and support.\n- Set update cadence: every 15 minutes for SEV1, 30 minutes for SEV2, and at state changes for SEV3.\n- Use a fixed update template: impact, customer-visible symptoms, current action, ETA or next update, commander, and communications lead.\n- Brief support and customer success with a live affected-customer list and approved holding statements within 15 minutes of SEV1 or SEV2.\n- Require executives to route questions through the executive liaison; the commander is not interrupted.\n- Define fatigue rules and formal handover for incidents lasting more than 4 hours.", "dependencies": ["S6", "S11"]}, {"step_id": "S14", "title": "Customer status page, account-manager outreach, and SLA credit workflow", "description": "Replace ad-hoc status updates with a **timed, owned, template-driven customer communication process**. Link incidents to SLA credits so finance and customer success are not surprised.\n\n- Publish status-page updates within 15 minutes of SEV1 declaration and 30 minutes for SEV2; update every 30 or 60 minutes until resolution.\n- Assign the communications lead as the single author; use pre-approved templates reviewed by legal and communications.\n- Map status-page components to customer capabilities: payments, settlement, reporting, API, onboarding, and regional availability.\n- For top accounts, require direct account-manager outreach within 30 minutes of SEV1 with an approved briefing pack.\n- Send proactive email or webhook notifications to subscribed customers for SEV1 and SEV2.\n- Publish a resolution notice within 30 minutes of mitigation and a customer-facing incident summary within 5 business days for SEV1.\n- Automate SLA credit calculation from incident duration and affected capability; review with finance and legal within 5 business days.\n- Track credits by incident and root-cause family to guide reliability investment.", "dependencies": ["S5", "S13"]}, {"step_id": "S15", "title": "Regulatory, legal, and partner notification playbook", "description": "In payments, some incidents are reportable and the clock starts at detection. Make regulatory assessment a **mandatory step in the incident process**, not an afterthought.\n\n- Map obligations: NYDFS cybersecurity-event rules, state breach laws, GLBA safeguards, PCI DSS if in scope, money-transmitter duties, card-network and sponsor-bank contracts, and cyber-insurer notice.\n- Require a legal or compliance reportability assessment within 2 hours for every SEV1 and every security-related SEV2, even when the answer is not reportable.\n- Maintain a 24x7 contact matrix for regulators, sponsor banks, card networks, outside counsel, insurers, and law enforcement.\n- Pre-draft notification templates and preserve legal review.\n- Encode customer-contract notification deadlines into account tiers and the communications workflow.\n- Record every reportability decision, approver, deadline, and submission confirmation in the incident record.", "dependencies": ["S5", "S14"]}, {"step_id": "S16", "title": "Postmortem standard and blameless review", "description": "Replace inconsistent postmortems with a **mandatory, blameless, single-format process**. The discipline comes from deadlines, facilitation, and action tracking.\n\n- Require postmortems for all SEV1 and SEV2 incidents, customer-detected incidents, incidents over 2 hours, repeat failures, and ledger near-misses.\n- Draft within 3 business days, peer review within 5, publish within 10 for SEV1 and SEV2.\n- Use one template: summary, impact, timeline, detection analysis, response analysis, contributing factors, what worked, what failed, and action items.\n- Make blamelessness explicit: focus on systems and decisions, not individual fault; never use postmortems in performance discipline.\n- Hold a weekly incident review board to review postmortems, ratify severity, and challenge weak actions.\n- Maintain a searchable postmortem library and quarterly recurring-cause analysis.\n- Require a trained facilitator for major reviews; the incident commander attends but does not facilitate.", "dependencies": ["S5", "S6"]}, {"step_id": "S17", "title": "Action-item tracking, ownership, and delivery gates", "description": "Only 11 of 64 actions were closed. Give postmortem actions the same status as **customer commitments**, with named owners and visible escalation.\n\n- Create every action as a ticket with one named individual owner, priority, due date, and verification method.\n- Use delivery classes: containment within 7 days, corrective work within 30 days, strategic work within 90 days.\n- Reserve 15–20% of team sprint capacity for reliability and incident actions.\n- Escalate overdue items: manager at 7 days, director at 14 days, CTO dashboard at 30 days.\n- Block related feature releases when overdue P0 actions prevent recurrence of a severe incident.\n- Require director approval and documented residual risk acceptance for overdue high-risk items.\n- Verify effectiveness after completion; closing a ticket without evidence does not close the action.\n- Target 90% of high-priority actions completed on time within two quarters.", "dependencies": ["S16", "S11"]}, {"step_id": "S18", "title": "Runbooks, critical-incident playbooks, and readiness bar", "description": "Poor runbooks are a real cause of pager resistance and slow mitigation. Define a **minimum readiness bar** before a service is allowed to page anyone at night.\n\n- Require for every Tier 0 and Tier 1 service: architecture summary, dependencies, dashboards, alert-to-runbook map, rollback procedure, feature flags, escalation contacts, and customer-impact statement.\n- Write major playbooks for shared PostgreSQL ledger failure, regional failover, Kubernetes control-plane loss, payment-processor outage, settlement-window breach, duplicate-payment suspicion, and security compromise.\n- Define ledger recovery rules: failover procedure, read-only degraded mode, reconciliation, RPO/RTO, and data-loss tolerance approved by executives.\n- Test runbooks in drills at least twice a year; mark untested runbooks stale.\n- Prevent paging alerts for services without readiness sign-off unless the engineering manager accepts the gap in writing.\n- Keep runbooks linked from every alert and incident template.", "dependencies": ["S4", "S8"]}, {"step_id": "S19", "title": "Training, certification, and role readiness", "description": "Command and communications are skills. Build a **tiered certification path** so rotations are staffed by people who have practiced, not by whoever is around.\n\n- All employees: 1-hour module on declaring incidents, finding the incident channel, and reading the status page.\n- All responders: half-day training on severity, acknowledgement, escalation, runbooks, and evidence hygiene.\n- Incident commanders: 2-day course plus two shadowed incidents and one simulation before certification.\n- Communications leads: training on status-page writing, customer language, account-manager briefs, and regulatory triggers.\n- Scribes: training on timeline discipline and audit evidence.\n- Certify for 12 months; renew through a simulation.\n- Require shadow shifts before independent primary duty; no new hire holds primary within 90 days.\n- Publish the certification register as an audit artifact.", "dependencies": ["S6", "S12", "S13", "S16"]}, {"step_id": "S20", "title": "Simulation program and game days", "description": "Rehearse the process before it meets a real SEV1. Simulations build commander confidence, expose runbook gaps, and produce audit evidence.\n\n- Run monthly 60-minute tabletops using real incidents from the 31-incident baseline.\n- Run quarterly game days covering regional failover, ledger replica promotion, dependency failure, partner outage, and security event.\n- Run twice-yearly unannounced paging drills to measure night acknowledgement times.\n- Include support, account managers, legal, compliance, and executives in at least one exercise per quarter.\n- Produce tracked action items from every exercise using the same board as real incidents.\n- Measure time to commander, time to first status update, and time to mitigation decision.", "dependencies": ["S19", "S11", "S18"]}, {"step_id": "S21", "title": "Critical-path pilot and gate review", "description": "Prove the model on the highest-risk services with willing teams before full rollout. Run a **six-week pilot with daily feedback and public exit criteria**.\n\n- Pilot with payments, ledger/database, platform/Kubernetes, API edge, authentication, and support intake.\n- Activate severity scale, duty commanders, paid rotations, single paging platform, alert budget, status-page policy, and postmortem process.\n- Hold a weekly pilot retrospective and fix process defects quickly.\n- Validate night acknowledgement, cross-team pull response, severity clarity, and compensation payroll.\n- Exit gate: commander assigned within 5 minutes in 95% of incidents, status page on time, page noise down at least 50%, postmortems on time, and positive on-call sentiment.\n- Publish pilot results to the whole company as the main adoption argument.", "dependencies": ["S7", "S10", "S11", "S14", "S18", "S19"]}, {"step_id": "S22", "title": "Wave rollout across all 28 teams", "description": "Roll out by criticality and dependency, not by calendar alone. Use **readiness gates** so teams are not forced live without coverage.\n\n- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1 teams. Wave 3: Tier 2 teams. Wave 4: Tier 3 and internal platforms.\n- Gate per team: catalog entry complete, alerts migrated, runbooks ready, six-person rotation staffed, one commander candidate nominated, compensation in payroll, and one tabletop passed.\n- Assign a program coach to each wave for three weeks.\n- Freeze legacy paging paths for each team after successful onboarding.\n- Publish a live adoption scoreboard by team.\n- Complete all teams by week 22, leaving months of operating evidence before the audit.", "dependencies": ["S21"]}, {"step_id": "S23", "title": "Metrics, dashboards, and review cadence", "description": "Instrument the process itself. Leadership must see whether the program is working, and auditors must see **operating evidence, not retrospective paperwork**.\n\n- Track response metrics: MTTD, time to declare, time to commander, acknowledgement time, mitigation time, resolution time, and customer-first detection rate.\n- Track quality metrics: page volume, alert actionability, missed pages, postmortem timeliness, action closure rate, and status-page compliance.\n- Track business metrics: availability against 99.95%, SLA credits, repeat incidents, and error-budget consumption.\n- Track people metrics: on-call load, night pages per person, recovery time usage, sentiment, and attrition signals.\n- Hold weekly incident review, monthly reliability review, quarterly executive review, and annual policy review.\n- Publish dashboards internally and give every metric a target and owner.", "dependencies": ["S11", "S16", "S21"]}, {"step_id": "S24", "title": "Change management, incentives, and culture", "description": "The process will be judged on fairness. Communicate the deal repeatedly and make participation **recognized, compensated, and safe**.\n\n- Core message: you are paid, paged only for what you own, supported by a trained commander, and given real capacity for actions.\n- Run CTO all-hands, team roadshows, office hours, FAQ, and an internal incident-management hub.\n- Add incident response and reliability work to promotion criteria and manager objectives.\n- Recognize good postmortems, alert-noise reduction, and calm incident leadership.\n- Provide a written path for engineers who cannot do nights; cover those shifts with paid volunteers or adjusted staffing.\n- Prohibit retaliation for good-faith declaration or escalation.\n- Publish sentiment survey results, including bad news, to maintain credibility.", "dependencies": ["S3", "S7", "S21"]}, {"step_id": "S25", "title": "SOC 2 evidence design and internal dry-run audit", "description": "Design the process so audit evidence is a **by-product of normal operations**. Test it internally before the external auditor does.\n\n- Map the process to SOC 2 criteria: incident identification, response, recovery, monitoring, communication, control activities, availability, and corrective action.\n- Approve versioned policies: incident response policy, severity standard, on-call policy, communications policy, postmortem standard, and exception process.\n- Retain incident records, paging logs, status-page history, postmortems, action tickets, training records, drill records, and access reviews for the audit period.\n- Record regulatory reportability decisions even when no notification is required.\n- Run an internal dry-run audit at month 6: sample at least 15 incidents and trace evidence end to end.\n- Fix gaps with at least four weeks before the audit.\n- Maintain an exception and remediation log instead of claiming perfection.", "dependencies": ["S16", "S17", "S22", "S23"]}, {"step_id": "S26", "title": "Program risk register and contingency planning", "description": "Name the likely failure modes now and pre-commit responses. Review the register monthly with the executive sponsor.\n\n- Commander volunteer shortfall: roster command duty for senior engineers and managers until the pool reaches 25 trained people.\n- Compensation delay: use immediate time-off-in-lieu plus phased stipend, but do not start mandatory night rotations without pay.\n- Tool migration slip: preserve single paging and routing scope; defer advanced automation if needed.\n- Alert pruning causing missed incidents: move noisy alerts to ticket first, observe 30 days, then delete.\n- Burnout in experienced teams: monitor load weekly and cap pages per person.\n- Major SEV1 during rollout: program owner shifts to incident support, wave schedule slips one wave, sponsor is informed same day.\n- Shared ledger concentration: track blast-radius reduction and failover improvements as top strategic actions.", "dependencies": ["S1"]}, {"step_id": "S27", "title": "Continuous improvement, maturity roadmap, and post-audit sustainability", "description": "Prevent the classic failure where the process decays after the audit. Build the second-year operating plan before the first year ends.\n\n- Hold quarterly process retrospectives with commanders, responders, support, and communications leads.\n- Re-baseline targets every six months; tighten goals once current targets are met.\n- Move from lagging metrics to leading indicators: error-budget burn, near-miss rate, drill performance, and action ageing.\n- Plan year-two improvements: follow-the-sun coverage, automated mitigation, ledger blast-radius reduction, error-budget release gates, and per-customer real-time impact reporting.\n- Keep annual policy review, certification renewal, drills, and board reporting on a permanent calendar independent of the audit cycle.\n- Report quarterly to the board or risk committee on availability, credits, severe incidents, overdue actions, and resilience investment.", "dependencies": ["S23", "S25"]}], "estimated_complexity": "high", "success_metrics": "- Median time to detect customer-impacting incidents falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.\n- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.\n- Median mitigation time falls from 3h10 to under 90 minutes by day 120 and under 60 minutes by month 9.\n- A named incident commander is assigned within 5 minutes in at least 95% of SEV1 and SEV2 incidents; zero incidents remain unowned for more than 10 minutes.\n- First status-page update occurs within 15 minutes of SEV1 declaration and 30 minutes of SEV2 declaration in at least 95% of qualifying incidents.\n- Monthly paging volume falls from 3,400 to under 500 actionable pages within six months, with alert actionability above 80%.\n- Out-of-hours pages average no more than 2 per responder per week; sustained breaches trigger mandatory alert remediation.\n- Six legacy alerting tools are consolidated into one paging and incident platform, with legacy paging paths disabled by week 16.\n- 100% of Tier 0 and Tier 1 services have a named owning team, escalation path, dashboard, and runbook by day 60.\n- All 28 teams are onboarded with readiness gates by week 22; every 24x7 critical-path rotation has at least six trained responders.\n- 30 or more certified incident commanders and 20 or more certified communications leads provide continuous primary and secondary coverage.\n- Paid on-call is approved by HR, legal, and finance and active in payroll before any new mandatory night rotation begins.\n- 100% of required SEV1 and SEV2 postmortems are drafted within 3 business days and published within 10 business days in the standard template.\n- Postmortem action closure rises from 17% to at least 90% of high-priority actions completed by their due date within two quarters.\n- All 53 open historical actions are triaged within 30 days; high-risk unaccepted items are completed or formally risk-accepted within 90 days.\n- Repeat incidents from a known unaddressed contributing factor decline by at least 50% within six months.\n- Annualized SLA credits fall from $1.3M to under $400k within 12 months.\n- Monthly availability meets or exceeds 99.95% by month 6, with exceptions reviewed at the executive reliability meeting.\n- At least two cross-company exercises, including regional and ledger scenarios, are completed before the SOC 2 audit, with critical findings tracked.\n- The month-6 internal dry-run audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete evidence.\n- SOC 2 Type II incident-response controls pass with zero exceptions at the month-8 audit.\n- On-call sentiment improves quarter over quarter; fewer than 10% of engineers report unwillingness to participate in their owned-service rotation by month 6.\n- No increase in attrition among engineers on active rotations compared with baseline."}Consolidated 30 loose steps into 24 denser ones (postmortems and actions merged into S17, runbooks and readiness folded into S14, all comms into S15) and imported the financial and training machinery it was missing. Remains the most readable plan, but it dropped two of its own best R0 steps.
- S2 adds a 7-day operating floor with an interim IC, retroactive pay and triage of all 53 open actions — R0 had no interim state.
- S16 adds the SLA credit workflow R0 lacked, tying credits to root-cause families and turning the $1.3M into a standing business case.
- S18 adds a tiered training academy with certification valid 12 months and a register as audit artefact.
- S21 now reports median and p90 by tier/journey/region, keeps error budgets gating features, and reconciles dashboards monthly against customer cases.
- S7 and S13 preserve ledger dual control and reconciliation even under the IC's pre-authorised failover powers.
- Fixed R0's arithmetic: S22 now claims 'roughly three months of operating evidence' after a week-24 finish rather than R0's impossible five months.
- Tightest metric set in the round (16 items), each tied to a date.
- Dropped R0 S14 'Publish Incident Management Policy v1' — the ten-page, signed, all-hands artefact is now only a bullet in S24, losing the mid-outage usable document.
- Dropped R0 S29 'inspect and adapt after 90 days'; there is no data-driven revision point before the annual review in S21, and severity calibration has no correction loop.
- R0 S30's mock auditor interviews and observation-window freeze are compressed to a single line in S24.
- No risk register or contingency step anywhere, unlike P1 S31 and P3 S26.
- S20 pilot does not depend on S15 (communications) although it activates the status page.
- Proposal 1 : Forensic re-coding of the 31 incidents and profiling of the 3,400 alerts before design is locked.
- Proposal 1 : SLA credit workflow, tiered commander academy, and the metrics/review cadence.
- Proposal 2 : Seven-day interim controls with a named IC in 10 minutes and a daily 15-minute ops review.
- Proposal 2 : Financial dual control during incidents, mitigation-versus-resolution semantics, and 8–12 coherent responder domains.
- Proposal 2 : Ingest all six tools first and retire a legacy path only after two weeks of verified operation; shadow-mode new alerts.
- Proposal 2 : SEV0 crisis tier.
- Proposal 1 : Separate steps for internal comms, customer comms and the regulatory playbook.
- Proposal 3 : Managed on-call service for overnight first-response triage.
+ Immediate 7-day operating floor+ Roles, authority, and ledger dual-control+ Three-layer 24x7 coverage model+ Customer-journey detection and ledger assurance+ Single incident and paging platform+ Detection-to-command escalation path+ Live execution and major-incident playbooks+ Internal, customer, and regulatory communications+ Blameless postmortems and action tracking+ Simulations and game days+ Critical-path pilot+ Metrics, reviews, and error budgets+ SOC 2 evidence, mock audit, and sustainabilityIncident roles, decision authority and handover rules24x7 incident command coverage modelTeam on-call structure, rotations and routing rulesDetection uplift: SLOs, synthetic journeys and ledger assuranceTool consolidation and incident platform implementationDetection-to-escalation path and the five-minute command ruleInternal communications protocolCustomer communications and status page policyRegulatory, partner and legal notification playbookPostmortem standard and blameless review forumAction item ownership and tracking systemRunbooks, major-incident playbooks and the on-call readiness barSimulation programme: game days, drills and wheel of misfortunePilot with wave 0 teamsMetrics, dashboards and the review cadenceSOC 2 control mapping, evidence automation and internal dry-run auditProgram risk register and contingency planningContinuous improvement, maturity roadmap and post-audit sustainability
The plan produced
1. Program charter, mandate, and funding from P1 step 1
Convert the CEO email into a named program with one owner, a budget, and a deadline earlier than the audit.
The charter must state that incident management is a company-level operating process, not a per-team choice. Without that, 28 teams will opt out.
- Appoint a Program Lead (Head of Reliability) with direct CTO and CEO sponsorship.
- Form a small steering group: CTO, VP Eng, Head of CS/Support, CISO, Legal, Finance, HR. Not a 28-team committee.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Timeline: operating floor in 7 days; design weeks 1–6; pilot weeks 7–12; rollout weeks 13–24; mock audit week 28; SOC 2 at month 8.
- Fund tooling ($150–250k/yr), on-call pay (~$600k–1M/yr), 2–3 program FTEs, and reserved engineering capacity. Anchor the ask against $1.3M in credits plus unmeasured incident cost.
- Freeze baselines now: 31 incidents, MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, $1.3M credits, 3,400 alerts/month at 85% noise, 11 of 64 actions closed.
2. Immediate 7-day operating floor (after 1) new
Do not wait for tooling, compensation, or the audit. Put a minimum process in place this week so the next outage has a named commander.
Start retaining artifacts on day one. This week becomes the first audit evidence.
- Publish a one-page interim severity guide and a single declaration path through Slack, phone, and pager.
- Staff interim primary and backup Incident Commander 24x7 from existing on-call veterans and engineering managers. Compensate this duty retroactively.
- Require a named IC within 10 minutes for every suspected major incident. If nobody claims it, the duty manager is IC.
- Use one channel, one bridge, one timeline doc, and one naming convention for every major incident.
- Direct Support to escalate credible customer reports immediately. They do not wait for engineering confirmation.
- Triage all 53 open historical actions. Close, re-plan, or formally risk-accept. Ledger integrity, payment duplication, regional resilience, security, and detection come first.
- Hold a daily 15-minute ops review until the permanent process is live.
3. Forensic baseline of incidents and alerts (after 1) from P1 step 2
Rebuild the facts before locking design. This is the before picture for the CEO and the auditor.
Every later design choice should trace to this evidence.
- Re-code each of the 31 incidents: trigger, service, detection source, timestamps, who led, credits paid, root-cause family.
- Quantify the 40% customer-first detections and name the missing signal in each case.
- Reconstruct the two nobody-in-charge incidents minute by minute. Use them as the burning-platform story.
- Audit the six alerting tools: volume per tool and team, top 50 noisy rules, rules with no owner or runbook.
- Freeze the baseline numbers. Do not let them drift during design.
4. Listening tour and the fairness contract (after 1) from P1 step 3
Engineer pushback against carrying a pager for other teams' code is the main delivery risk. Treat it as a design constraint, not an attitude problem.
The answer that will actually stick is the deal: you are paid; you are paged only for services you own; a trained commander runs the room; your postmortem actions get real sprint capacity.
- Interview all 28 teams plus Support, CS, and Sales in two weeks.
- Separate the objections: unpaid work, nights, unfamiliar code, bad runbooks, fear of blame. Each needs a different fix.
- Collect informal practices from the 12 teams already on-call. They are the pilot candidates and the veteran pool.
- Recruit 10–15 credible engineers as a design working group so the process is co-authored.
- Baseline sentiment on alert trust, on-call willingness, and burnout. Re-measure at 6 and 12 months.
5. Service ownership catalog and criticality tiers (after 3) from P1 step 4
You cannot page the right person across 180 services until each one has a named owner. This is the foundation of fairness, routing, and audit evidence.
Build a machine-readable catalog as the single source of truth.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, the shared PostgreSQL cluster, queues, partner banks, and customer-facing endpoints.
- Assign one owning team, named engineering manager, Slack channel, escalation policy, and dependency list per service.
- Tier 0: money movement, ledger, auth, shared Postgres, regional control plane. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map Tier 0/1 services to customer capabilities: initiation, authorization, settlement, reporting, onboarding.
- Assign coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
6. Severity scale and declaration rules (after 3, 5) from P1 step 5
Replace in-the-moment debate with a lookup. Four payments-specific levels, with worked examples drawn from the 31 real incidents.
Anyone may declare. Only the Incident Commander may downgrade, with recorded rationale. When unsure, start high.
- SEV1: money movement stopped or incorrect; ledger integrity in doubt; material security or data exposure; both regions impaired; or more than 10% of customers impacted. Pages IC, comms, scribe, SMEs, and exec liaison. Bridge in 5 minutes. Status page in 15 minutes. Mandatory postmortem and regulator assessment.
- SEV2: severe degradation; settlement window at risk; a strategic customer fully down; SLA breach likely. IC and SMEs paged. Status page in 30 minutes. Mandatory postmortem.
- SEV3: partial impact with a workaround; no credit exposure. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, or a repeat.
- SEV4: internal or minor. Ticket only. No page.
- Auto-escalate: any SEV3 open more than 2 hours, or any incident touching the shared ledger, becomes SEV2. Unknown impact after 15 minutes is raised, not sat on.
- Nobody is punished for over-declaring. Publish that rule in writing and repeat it.
7. Roles, authority, and ledger dual-control (after 6) new
Solve nobody-in-charge-for-an-hour by making command explicit, transferable, and logged.
The Incident Commander owns the incident, not the fix, and never types in a production terminal.
- Incident Commander: declares severity, pulls anyone, freezes deploys, invokes failover, authorises spend. One IC at a time. Assumed within 5 minutes and announced in channel.
- Communications Lead: single voice to customers, status page, account managers, and the exec summary.
- Scribe: timestamped timeline, decisions, and open questions. Feeds the postmortem and the audit trail. Required for SEV1 and SEV2.
- Subject-matter responders: diagnose and mitigate only services they own, with access and runbooks.
- Executive Liaison (SEV1): shields the IC from exec questions; owns regulator and board escalation.
- The IC may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access rules.
- Handover is verbal and written, with time and new owner recorded. Roles may combine below SEV2; never at SEV1. The IC stays in charge if a VP joins.
8. Three-layer 24x7 coverage model (after 5, 7) new
Do not put 28 teams on night rotation. That is what engineers are rejecting.
Staff command centrally. Page engineers only for code their team owns.
- Layer A, company command: Duty IC plus secondary, Duty Comms, and a scribe pool. Target 25–35 certified ICs and 15–20 comms leads. Roughly one week per person every 6–8 months. Five-minute ack SLA, then secondary, then on-call director.
- Layer B, Tier 0/1 domains: group 28 teams into 8–12 product and platform domains (ledger app, Postgres platform, payments orchestration, auth, API edge, Kubernetes/platform, settlement). Primary plus secondary 24x7. Minimum six trained people. Target one week in six, never worse than one in four.
- Layer C, Tier 2/3: business-hours on-call. After hours the IC pages the EM, who holds a written escalation list.
- No engineer joins another team's responder pool without training, access, runbooks, shadow shifts, and both teams' acceptance.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner.
- Evaluate follow-the-sun coverage as a 12-month option, not a year-1 dependency.
9. Paid on-call, NY labor compliance, and fatigue rules (after 4, 8) from P1 step 9
Unpaid on-call is a retention risk and a New York legal exposure. Pay must be live in payroll before any mandatory night rotation starts.
Treat interrupted nights and recovery time as compensable work.
- Weekly stipend by layer, higher for 24x7 Tier 0/1 and Duty IC, lower for business-hours, with holiday and weekend premiums.
- After-hours call-out pay or TOIL. A paid recovery day after prolonged overnight work, a SEV1, or a qualifying SEV2. Managers cover the next day.
- HR, Legal, and Finance publish dollar amounts, FLSA exempt/non-exempt treatment, NY wage-hour rules, tax treatment, and payroll timing within 14 days of charter.
- Load rules: no primary on two rotations; no consecutive primary weeks; no on-call the week after a SEV1 you commanded.
- A person may declare temporarily unfit after overnight work with no performance penalty.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- If a rotation cannot staff six people, merge domains or hire. Do not run two-person 24x7.
- Model annual cost against $1.3M in credits and get it as a CFO/board line item.
10. Alert quality contract and page budget (after 3, 5) from P1 step 10
3,400 alerts a month at 85% noise is why detection takes 22 minutes. Make quality a condition of paging a human.
A page is a product with a quality bar, not a dump of host metrics.
- Every paging alert must have a named owner, customer or SLO impact, runbook, tested threshold, severity mapping, dashboard, expected action, and dedup key. Fail any of these and it becomes a ticket or is deleted.
- Page on customer symptoms: SLO burn, error budget, settlement-queue age. Cause-based CPU and memory alerts become dashboards or tickets.
- Set a page budget of at most two out-of-hours pages per person per week. A breach triggers a mandatory tuning sprint and blocks new paging alerts for that team.
- Auto-quarantine alerts that fire more than five times a month with no action, or that have more than 70% no-action acks. Never silently disable without compensating detection and a recorded decision.
- Run new alerts in shadow for seven days unless an emergency exception is approved.
- Monthly per-team kill, tune, or keep review. Target under 500 pages a month and actionability above 75% within six months.
11. Customer-journey detection and ledger assurance (after 5, 10) new
Stop customers being the monitoring system. Detect on money-path outcomes, not host metrics.
Every postmortem will ask why a customer saw it first.
- Define SLOs per Tier 0/1 capability: initiation success, auth latency, settlement timeliness, API availability, reporting freshness. Set an internal target stricter than 99.95%.
- Run synthetic full-lifecycle payments from outside the platform, in both regions, every 60 seconds. Include a small real-value canary where legally feasible.
- Add ledger assurance: continuous double-entry reconciliation, replication lag, failover-readiness, unexpected balances, duplicate identifiers, settlement-window countdown.
- Add per-customer anomaly detection for the top 100 accounts.
- Auto-create a triage incident within five minutes when support tickets or AM reports match impact keywords.
12. Single incident and paging platform (after 6, 8, 10) from P2 step 7
Collapse six alerting tools into one queue, one timeline, and one audit record.
The platform must still work if one AWS region is down. Verify SMS and phone fallback, plus an offline runbook.
- Select one paging product, one incident layer, and one hosted status page.
- One Slack command declares the incident, creates the channel and bridge, pages the Duty IC, sets severity, and starts the clock.
- Ingest the six existing tools first. Deduplicate and route. Retire a legacy path only after named owners and two weeks of verified operation.
- Auto-capture timestamps, acknowledgements, role assignments, severity changes, comms, mitigated time, and resolved time. Export for SOC 2.
- Integrate the service catalog, Jira, Salesforce or CS tooling for affected-customer lists, and Zoom or Slack Huddle.
- Test paging, escalation, status publication, and conference access every week.
- After a team's wave, old paging paths are disabled, not left as a fallback.
13. Detection-to-command escalation path (after 7, 8, 12)
Write the single path from something looks wrong to someone is in charge. Target: a commander in under five minutes, any hour.
If nobody claims IC in five minutes, the platform assigns the Duty IC. The assignee can hand over, not decline.
- Converge every entry point on the same declare command: alert, engineer, support, account manager, partner bank, SEV hotline.
- SEV1: page the owning primary immediately; secondary at 5 minutes unacked; domain manager and Duty IC at 10; exec liaison at 15.
- SEV2: primary ack in 10 minutes; IC assigned in 15.
- Human acknowledgement is required. Delivery to a device does not count.
- The IC can page any team's on-call, with a 10-minute ack obligation. This reciprocity makes single-team ownership viable.
- Unowned alerts go to Layer A command, then the missing owner record is a control defect.
- Pre-authorise regional failover, ledger read-only mode, and partner-bank notice so the IC does not wait for an executive. Dual-control still applies to ledger writes.
14. Live execution and major-incident playbooks (after 7, 13) new
Limit customer and financial harm before proving root cause. One procedure from the first minute to handback.
A service cannot page at night until it meets the readiness bar.
- Open channel, bridge, record, and timeline immediately for SEV1 and SEV2. The IC states severity, known impact, hypothesis, objective, roles, and next update time.
- Freeze unrelated production changes during SEV1. Record exceptions the IC approves.
- Prefer reversible mitigation: rollback, feature flag, traffic isolation, rate limit, partner reroute.
- Write playbooks first for Postgres ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank outage, settlement-window breach, suspected security compromise, and suspected duplicate payments.
- Guard against split-brain, replay, and duplication during regional or database recovery. Reconcile and process backlog before calling the incident resolved.
- Mitigation means customer impact ended. Resolution means stable, backlog processed, and ledger reconciled.
- Formal IC handover after 4 hours. Reopen if impact recurs during the stability window.
- Readiness bar per Tier 0/1 service: diagram, dependencies, dashboard, runbook, rollback, kill switch, escalation contacts, RPO/RTO. Untested runbooks are marked stale.
15. Internal, customer, and regulatory communications (after 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved process. The Communications Lead is the single author.
State impact and the next update time. Never speculate on cause.
- Internal: one working channel, one bridge, one read-only broadcast for execs, Support, and Sales. SEV1 update every 30 minutes even if unchanged; SEV2 every 60 minutes.
- Executives ask questions only of the Executive Liaison. Publish this as a signed exec behaviour rule.
- Status page: SEV1 within 15 minutes, SEV2 within 30 minutes; then 30/60-minute updates; resolve notice within 30 minutes of mitigation. Templates pre-approved with Legal.
- Top 100 accounts: named AM contact within 30 minutes of SEV1, with a briefing pack from Comms. Long tail gets the status page plus email or webhook.
- Customer-facing summary within 5 business days for SEV1.
- Mandatory regulatory checkpoint on every SEV1 and every security SEV2, within 2 hours, recorded even when not reportable. Cover NYDFS Part 500 72-hour clock, state breach laws, GLBA/FTC, PCI if in scope, sponsor-bank and card-network windows, FinCEN/OFAC if relevant.
- Legal owns outbound regulatory letters. The IC owns facts. Security or Legal may limit public detail during an active threat, with the reason recorded.
- Encode bespoke customer-contract notice SLAs into account tiering.
16. SLA credit and financial-impact workflow (after 6, 15) from P1 step 17
Link incidents to money so severity, credits, and investment stay consistent. Finance should not learn about outages from invoices.
Make credit calculation an output of the incident record, not a negotiation.
- Agree availability measurement per contract and component with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and capability telemetry. Produce a proposed credit schedule within 5 business days of resolution.
- Posture: proactive credits for top-tier accounts; claims-based for the rest. Document the approval chain.
- Track credits per incident and root-cause family. Quarterly report which reliability investments would have prevented which credits.
- Target: cut credits from $1.3M to under $400k in 12 months. Use that delta as the ongoing business case.
17. Blameless postmortems and action tracking (after 6, 12)
Eleven of 64 actions closed is the clearest process failure. Make learning mandatory and actions as binding as customer commitments.
Closing a ticket without evidence of effectiveness does not close the action.
- Mandatory for every SEV1 and SEV2, every customer-first detection, every incident over 2 hours, every repeat of a known cause, and every ledger near-miss.
- Draft within 3 business days, review within 5, published internally within 10. The IC owns the draft. The owning EM is accountable.
- One template: timeline, customer and financial impact, why detection was late, why mitigation took that long, contributing factors, what went well, actions.
- Blameless in writing. No individual named as a cause. HR and management commit that postmortems are never used in performance reviews.
- Every action gets a named person, priority, due date, Jira ticket, and verification method. P0 (prevents SEV1 recurrence) due in 30 days and committed into the next sprint before roadmap work. P1 in 60 days. P2 in 90 days.
- Teams reserve 15–20% of sprint capacity for reliability and incident actions.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Weekly Incident Review Board with engineering directors. Target 90% of P0/P1 actions closed on time within two quarters.
18. Training, certification, and commander academy (after 7, 13, 15, 17) from P1 step 21
Command is a skill, not a title. Nobody holds independent duty untrained. Training happens in paid working time.
The certification register is an audit artefact.
- All employees, 1 hour: recognise impact, declare, find the channel and status page.
- Responder, half day: severity, escalation, runbooks, comms hygiene. Mandatory before joining a rotation.
- Scribe, 2 hours: timeline discipline. The entry point.
- Incident Commander, two days plus shadowing: command presence, decisions under uncertainty, severity, handover, exec management. Two shadowed incidents and one simulated SEV1 before certification.
- Communications Lead, one day: status writing, customer tiering, legal boundaries, regulator triggers.
- Certification valid 12 months, renewed via simulation.
- New joiners shadow two shifts and do not hold primary in their first 90 days.
- Appoint a champion in each of the 28 teams.
19. Simulations and game days (after 14, 15, 18) new
Rehearse before the next real SEV1. Use a standing calendar, not a one-off exercise.
Do not inject uncontrolled changes into the production ledger.
- Monthly 60-minute tabletop per engineering group, using a real incident from the 31.
- Quarterly full-scale game day: regional failover, ledger replica promotion, dependency failure. Whole role structure, timed.
- Twice-yearly unannounced paging drill, including nights, to measure real acknowledgement times.
- One security incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Also exercise status-page failure and loss of the primary chat or pager.
- Use replicas, staging, or tightly governed tests for ledger scenarios.
- Every exercise produces actions in the same tracker as real incidents.
- Complete two cross-company exercises before the SOC 2 audit.
20. Critical-path pilot (after 9, 11, 12, 18, 19) new
Prove the process on the highest-risk surface with willing teams before asking 28 teams to adopt it.
Publish a one-page result company-wide. That result is the adoption argument.
- Six-week Wave 0: ledger app, Postgres platform, payments orchestration, Kubernetes/platform, API gateway, Support intake, plus two of the 12 teams already on-call.
- Activate the full stack: severity scale, Duty IC, one pager, page budget, status page, mandatory postmortems, paid on-call.
- Parallel-run old paths for one week, then cut over. The program lead coaches every SEV2+ and does not secretly take command.
- Weekly retro. Expect 20–40 process defects and fix them in the standard before rollout.
- Exit gate: MTTD under 10 minutes on pilot services; IC assigned within 5 minutes in 95% of incidents; page volume down 50%; all postmortems on time; no unpaid pages; sentiment not worse.
21. Metrics, reviews, and error budgets (after 12, 17, 20)
Instrument the process itself. Do not reward hiding incidents or suppressing pages.
Every metric has a target and a named owner. Dashboards are public inside the company.
- Response: MTTD, time to declare, time to IC, MTTA, MTTM, MTTR, incidents by severity, percent customer-first. Report median and p90, by tier, journey, and region.
- Quality: pages per person per week, actionability, budget breaches, postmortem on-time rate, action closure and age, missed acknowledgements.
- Business: credits, 99.95% per capability, error-budget burn, failed-payment volume, reconciliation breaks, repeat-incident rate.
- People: rotation size, frequency, after-hours pages, recovery days, sentiment, attrition among on-call staff.
- Cadence: weekly Incident Review Board; monthly Reliability Review; quarterly exec and board review; annual policy review.
- Error budgets on Tier 0/1 SLOs. Burn too fast and the team pauses features to pay down reliability.
- Reconcile dashboards monthly against a sample of incident records and customer cases so missing incidents cannot hide.
22. Wave rollout with readiness gates (after 20, 21) from P1 step 25
Roll out in four waves by criticality, every three weeks. Gates keep the standard credible. A missed gate is rescheduled, not waived.
Finish all 28 teams by week 24 so roughly three months of operating evidence remain before audit fieldwork.
- Wave 1: remaining Tier 0. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team kit: catalog complete, alerts migrated and inside budget, runbooks at the readiness bar, rotation of 6+ certified responders, one IC nominee, one drill passed, paid rotation in payroll.
- Named coach for three weeks. The director signs the gate.
- Freeze legacy paging per wave.
- Publish a live adoption scoreboard.
- Any production service that cannot provide sustainable ownership gets an executive-reviewed deadline, compensating control, and expiry date. No indefinite verbal exceptions.
23. Change management, incentives, and the deal (after 1, 4, 9) from P1 step 27
Run this from day one in parallel with design. Engineers will judge fairness. Executives will judge visible results.
Repeat the deal until it is muscle memory: paid on-call, paged only for what you own, trained commander, real sprint capacity for actions.
- Launch via CTO all-hands, per-team roadshows, a one-page laptop card, an internal wiki, and a Slack help channel with a 4-hour answer SLA.
- Put incident-response contribution into promotion criteria. Award best postmortem and biggest noise cut quarterly. Thank people publicly after every SEV1.
- Put adoption, alert hygiene, action closure, and on-call load fairness into every engineering manager's quarterly objectives.
- Write an exception path for engineers who cannot do nights because of caring responsibilities or health, covered by stipended volunteers.
- Prohibit retaliation for good-faith declaration or escalation.
- Pulse-survey at 60 and 120 days. If fairness or load is red, pause expansion until fixed.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
24. SOC 2 evidence, mock audit, and sustainability (after 17, 21, 22) new
Evidence is a by-product of doing the work, not a reconstruction the week before the auditor. Protect the process after SOC 2 is signed.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to CC7.2–CC7.4, CC2.2/CC2.3, CC5, and availability A1.2. Confirm interpretation early.
- Maintain versioned, signed policies for incident response, severity, on-call, communications, and postmortems, reviewed annually.
- Automate evidence: incident records, paging and ack logs, status history, postmortem library, action closure, training register, drill records, and reportability decisions including not reportable.
- Internal dry-run at month 6 on a 15-incident evidence walkthrough. Mock audit at month 7. Fix with weeks to spare.
- Assign permanent owners for policy, pager, status page, catalog, training, and metrics.
- Year-two roadmap: follow-the-sun cell, self-healing for the top recurring causes, error-budget release gates, and blast-radius reduction on the shared ledger.
- Quarterly board summary of severe incidents, credits, overdue P0s, and resilience investment so attention does not die after the audit.
- Median time to detect falls from 22 minutes to under 10 minutes by month 3 and under 5 minutes by month 9.
- Customer-first detection falls from 40% to under 20% by month 3 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by month 4 and under 60 minutes by month 12.
- A named Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes after week 4.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases by month 3, with required update cadence met in 95% of cases.
- Monthly pages fall from 3,400 to under 500 within 6 months, with actionability above 75%; out-of-hours pages at or under 2 per person per week.
- All six legacy alerting tools route through one paging platform by week 16; legacy paging paths disabled per wave.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by due date within two quarters; all 53 currently open historical actions triaged within 30 days of S2.
- SLA credits fall from $1.3M to under $400k in the first 12 months; customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known unaddressed cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar; every Layer B rotation has 6+ certified responders or a time-limited executive exception.
- Paid on-call policy approved by HR, Legal, and Finance and in payroll before any mandatory night rotation starts.
- All 28 teams onboarded by week 24 with 24x7 command cover from 25+ certified ICs and 15+ certified Comms Leads.
- Internal dry-run at month 6 passes a 15-incident evidence walkthrough; month-7 mock audit finds no unowned high-risk control gap; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment improves quarter over quarter; no increase in attrition among engineers on rotation; pulse at 60 and 120 days shows at least 70% agree rotations are fair and limited to services they own.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked actions; two cross-company exercises completed before the audit.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 223f0a8a-068a-4034-ad20-8e2ff8553040, Agent: opus5_initial_1, LLM: anthropic/claude-opus-5):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 9 months.
- Customer-first detection falls from 40% of incidents to under 10% within 6 months and under 5% within 12.
- Median time to mitigate falls from 3 h 10 min to under 60 minutes within 12 months.
- An Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes.
- Status page updated within 15 minutes of SEV1 declaration and 30 minutes of SEV2 in 95% of cases.
- Monthly alert volume falls from 3,400 to under 500 pages, with actionability above 75%; out-of-hours pages under 2 per person per week.
- All six legacy alerting tools consolidated into one paging platform, legacy paging paths disabled, by week 16.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by their due date.
- SLA credits fall from $1.3M to under $400k in the first 12 months.
- Customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar.
- All 28 teams onboarded by week 24, with 24x7 rotations of 6+ certified responders for every Tier 0/1 team.
- 30+ certified Incident Commanders and 20+ certified Communications Leads, giving 24x7 primary and secondary command cover.
- Paid on-call policy approved by HR, Legal and Finance and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among engineers on rotation.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked action items.
Steps (29):
1. Program charter, executive mandate and funding
Convert the CEO's frustration into a named program with one accountable owner, a budget and a deadline that is earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, the 28 teams will opt out.
- Appoint a single Incident Management Program Lead (Head of Reliability/SRE) with direct exec sponsorship from CTO and CEO.
- Form a steering group: CTO, VP Eng, Head of Support/CS, CISO/Compliance, Legal, Finance (SLA credits), HR (on-call pay).
- Set the non-negotiables: one severity scale, one paging tool, one postmortem format, mandatory action tracking, paid on-call.
- Fix the timeline: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready by week 28 (four weeks of buffer before the audit).
- Approve budget lines: tooling (~$150–250k/yr), on-call compensation (~$600k–1M/yr), 2–3 dedicated program FTEs. Anchor it against $1.3M of credits plus incident cost.
2. Forensic baseline of the 31 incidents and the alert estate (depends on: 1)
Before designing anything, rebuild the facts. Re-open all 31 incidents and profile the 3,400 monthly alerts so every later design decision is evidence-based.
This also creates the "before" picture the exec and the auditor will compare against.
- Re-code each incident: trigger, service, detection source (customer vs monitor), timestamps for detect/acknowledge/declare/mitigate/resolve, who led, credits paid, root cause family.
- Quantify the 40% customer-first detections: which signal was missing in each case.
- Classify the two "nobody in charge" incidents minute by minute; use them as the burning-platform story.
- Audit the six alerting tools: volume per tool, per team, per alert rule; identify the top 50 rules that produce most of the 85% noise; find rules with no owner and no runbook.
- Baseline the numbers formally: MTTD 22 min, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed. Freeze them as the reference line.
3. Stakeholder listening tour and resistance map (depends on: 1)
Engineer pushback against "carrying a pager for other teams' code" is the main delivery risk. Treat it as a design input, not an attitude problem.
Run structured interviews across all 28 teams, plus Support, CS and Sales, in two weeks.
- Test the real objection: is it unpaid work, night sleep, unfamiliar code, poor runbooks, or fear of blame? Each has a different fix.
- Collect current informal practices — the 12 teams already on-call are the pilot candidates and the source of veterans.
- Document the promise that answers the objection: **you are only paged for services your team owns**, plus a trained commander who runs the incident and pulls in others.
- Map influencers and blockers by team; recruit 10–15 credible engineers as a design working group so the process is co-authored, not imposed.
- Survey baseline sentiment (trust in alerts, willingness to be on-call, burnout) to re-measure at 6 and 12 months.
4. Service ownership catalog and criticality tiering (depends on: 2)
You cannot page the right person across 180 services until each service has a named owning team. This is the foundation of both on-call fairness and severity mapping.
Build a machine-readable catalog (Backstage or equivalent) that is the single source of truth for routing.
- One owning team per service, a named engineering manager, a Slack channel, a paging escalation policy, a dependency list.
- Tier services by business impact: Tier 0 (money movement, ledger, auth, shared PostgreSQL cluster), Tier 1 (customer-facing but degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map each Tier 0/1 service to the customer-visible capability it supports (payment initiation, settlement, reporting, onboarding).
- Flag orphan services and cross-team shared components; force an ownership decision for each within 30 days, or schedule decommissioning.
- Publish coverage gaps to the steering group: any Tier 0 service without an owner is an executive escalation.
5. Severity scale and declaration criteria (depends on: 2, 4)
Define a five-level scale with objective, payments-specific triggers so declaration is a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade.
Each level triggers a fixed bundle of response, comms and postmortem obligations.
- **SEV1**: money movement stopped or incorrect, ledger integrity in doubt, data breach, full region loss, >10% of customers impacted. Triggers: immediate 24x7 page of IC + comms + exec, bridge within 5 min, status page within 15 min, mandatory postmortem, regulator assessment.
- **SEV2**: severe degradation, settlement at risk of missing a window, single large/strategic customer fully down, SLA breach likely. Triggers: IC paged, status page within 30 min, mandatory postmortem.
- **SEV3**: partial or workaround-available degradation, no credit exposure. Team-led, business-hours comms, postmortem optional but encouraged.
- **SEV4/5**: minor or internal-only; ticket-tracked, no paging.
- Add auto-escalation rules: any SEV3 open >2 h, or any incident touching the shared ledger cluster, becomes SEV2 automatically. Include a severity decision tree and 12 worked examples drawn from the 31 real incidents.
6. Incident roles, decision authority and handover rules (depends on: 5)
Solve the "nobody in charge for an hour" failure by making command explicit, transferable and logged.
Define five roles with written responsibilities, entry criteria and explicit authority.
- **Incident Commander**: owns the incident, not the fix. Authority to declare severity, pull any engineer, approve customer-impacting mitigations, invoke failover and authorise spend. The IC never types in the terminal.
- **Communications Lead**: owns status page, internal updates, account-manager briefings and the exec summary. Single voice to customers.
- **Scribe**: maintains the timeline, decisions and open questions; feeds the postmortem and the audit evidence trail.
- **Subject-Matter Responders**: engineers from owning teams; they investigate and remediate, and report to the IC.
- **Executive Liaison** (SEV1 only): shields the IC from exec questions and owns regulator/board escalation.
- Rules: the IC role is assumed within 5 minutes of declaration, stated explicitly in the channel ("I am IC"), and any handover is announced and logged. Roles may be combined below SEV2; never at SEV1.
7. 24x7 incident command coverage model (depends on: 6, 4)
Command is staffed by a small, trained, cross-team pool — not by 28 teams individually. This is what makes 24x7 realistic in one New York time zone.
Design a central rotation that scales with a growing certified pool.
- Create a **Duty Incident Commander** rotation of 25–35 certified volunteers (target ~1 per team, plus managers and senior engineers), giving each person roughly one week per 6–8 months.
- Pair with a Duty Comms Lead rotation (Support/CS leads plus engineering managers, ~15–20 people) and a Scribe pool (rotating, lowest barrier, used as the training entry point).
- Coverage: primary + secondary IC at all times; hard 5-minute acknowledgement SLA with automatic failover to secondary, then to the on-call engineering director.
- Night coverage options to evaluate in writing: US-only rotation with paid night stipend now; a Lisbon/Dublin or APAC follow-the-sun cell as a 12-month option; a 24x7 NOC-style triage desk for first-line detection.
- Eligibility: certification required (S21); commanders are volunteers with manager approval and can step out with 30 days' notice.
8. Team on-call structure, rotations and routing rules (depends on: 4, 6)
Rebuild team on-call around the principle that answers the pushback: **you are only paged for code your team owns**.
Apply a tiered obligation so 28 teams are not treated identically.
- Tier 0/1 owning teams (expected 14–18 teams): 24x7 primary + secondary, minimum 6 people per rotation, one-week shifts, handover Wednesday mornings.
- Tier 2/3 teams: business-hours on-call with a best-effort out-of-hours escalation path, no night paging.
- Platform/Infrastructure and Database teams: 24x7, since they own the shared PostgreSQL ledger cluster and the Kubernetes/regional layer.
- Rotations under 6 people are merged across teams or backfilled by hiring; no rotation of fewer than 4 is approved.
- Routing: every page resolves through the service catalog to the owning team's escalation policy; cross-team pages are made by the IC, never by an alert.
- Guardrails: maximum one week in four, no on-call in the week after a SEV1 you led, protected recovery time after any night page, and a per-person page budget (see S10).
9. On-call compensation, labour compliance and fairness policy (depends on: 8, 3)
Unpaid on-call is both a retention risk and a legal exposure in New York. Paying for it is the fastest way to convert resistance into participation.
Design the scheme with HR, Legal, Finance and Payroll, and publish it before asking anyone to sign up.
- Base stipend per week on rotation, differentiated by tier: e.g. $800–1,200 for 24x7 Tier 0/1, $300–500 for business-hours rotations, with premiums for holidays and weekends.
- Per-incident payment for out-of-hours activation (e.g. $150 per night page plus hourly beyond one hour) and guaranteed time-off-in-lieu after night work.
- Separate Duty IC stipend, since command is a distinct and heavier burden.
- Verify FLSA exempt/non-exempt treatment, NY State wage rules and overtime exposure for non-exempt staff; document the legal review.
- Budget and model the annual cost; get board/CFO approval as a line item, benchmarked against $1.3M of credits.
- Add non-cash elements: on-call time counted as delivery load (teams reduce sprint commitment by ~15%), incident leadership recognised in promotion criteria, and a public quarterly report of on-call load per team.
10. Alert quality standard and page budget (depends on: 2, 4)
3,400 alerts a month at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a contractual condition of paging someone.
Publish a standard, then enforce it mechanically.
- Every paging alert must have: a named owning team, a documented customer impact, a runbook link, a tested threshold, and a severity mapping. Alerts failing this are demoted to ticket or deleted.
- Page only on symptoms that affect customers (SLO burn rate, error budget, queue depth against settlement deadlines); cause-based CPU/memory alerts become dashboards or tickets.
- Set a **page budget**: maximum 2 out-of-hours pages per person per week. Breach triggers a mandatory alert-tuning sprint for the owning team and blocks new alert creation.
- Auto-quarantine: any alert that fires more than 5 times a month without action, or has >70% no-action acknowledgements, is silenced automatically and returned to its owner.
- Monthly alert review per team: kill, tune, or keep, with the numbers on screen.
- Target: 3,400 → under 500 pages a month, with actionability above 75% within six months.
11. Detection uplift: SLOs, synthetic journeys and ledger assurance (depends on: 10, 4)
The goal is to stop customers telling you first. Detection must be driven by customer-visible outcomes, not host metrics.
Instrument the money path end to end and alert on it.
- Define SLOs for each Tier 0/1 customer capability: payment initiation success rate, authorisation latency, settlement file timeliness, API availability and reporting freshness. Tie them to the 99.95% contractual SLA with a stricter internal target.
- Deploy synthetic transactions from outside the platform, in both regions, every 60 seconds, covering the full payment lifecycle including a small real-value canary flow where feasible.
- Add ledger assurance checks: continuous double-entry balance reconciliation, replication lag and failover-readiness alarms on the shared PostgreSQL cluster, and settlement-window countdown alerts.
- Build per-customer anomaly detection for the top 100 accounts (volume drop, error spike) so a single-tenant outage is detected before the account manager calls.
- Create an "inbound signal" bridge: any support ticket or account-manager report matching impact keywords auto-creates a triage incident within 5 minutes.
- Track every incident's detection source; make "customer detected first" a reviewed defect with its own follow-up action.
12. Tool consolidation and incident platform implementation (depends on: 5, 7, 8, 10)
Collapse six alerting tools into one paging and incident platform so there is a single queue, a single timeline and a single audit record.
Run a short, time-boxed selection and migrate within the pilot window.
- Select an integrated stack: paging/on-call scheduling plus an incident management layer (e.g. PagerDuty + incident.io/FireHydrant, or a single vendor) and a hosted status page.
- Implement one-command declaration in Slack (`/incident declare`) that creates the channel and bridge, pages the Duty IC, sets severity, opens the timeline and starts the clock.
- Migrate all monitoring sources to route into the one platform; decommission direct paging from the legacy six and block new integrations that bypass it.
- Automate the evidence trail: timestamps, role assignments, severity changes, comms sent, and postmortem linkage exported for SOC 2.
- Integrate with the service catalog for routing, Jira for actions, Salesforce/CS tooling for affected-customer lists, and Zoom/Slack Huddle for the bridge.
- Hard requirement: the platform must work when AWS in one region is down — verify out-of-band paging (SMS/phone) and a printed/offline fallback runbook.
13. Detection-to-escalation path and the five-minute command rule (depends on: 6, 7, 12)
Write the single path from "something looks wrong" to "someone is in charge", and make it impossible to skip.
The design target is detection to commander in under five minutes, any hour.
- Entry points: automated alert, engineer observation, support ticket, account manager, partner bank, customer-facing SEV hotline. All converge on the same declaration command.
- Anyone in the company may declare up to SEV2; nobody is punished for over-declaring. Publish that rule in writing and repeat it.
- Auto-page ladder: Duty IC (5 min) → secondary IC (5 min) → on-call Director (10 min) → CTO. Same ladder for the owning team's responder.
- Cross-team pull: the IC can page any team's on-call directly, with a 10-minute acknowledgement obligation. This is the reciprocal commitment that makes single-team ownership viable.
- Explicit takeover protocol: if no one claims IC within 5 minutes, the platform assigns it and announces it; the assignee cannot decline, only hand over.
- Define standing severity triggers for immediate regional failover, ledger read-only mode and partner-bank notification, with pre-authorised decision rights so the IC does not wait for an executive.
14. Internal communications protocol (depends on: 6, 12)
Standardise the internal channel so responders, executives and support see the same picture without interrupting the IC.
Separate the working channel from the audience channel.
- One incident channel per incident (auto-created), one bridge, and a read-only broadcast channel for executives, Support and Sales.
- Update cadence by severity: SEV1 every 30 minutes even if nothing has changed; SEV2 every 60 minutes; SEV3 at state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, ETA or next update time, current IC and Comms Lead.
- Exec briefing rule: executives ask questions only to the Executive Liaison; the IC is not interrupted. Publish this as a behavioural expectation signed by the exec team.
- Support/CS enablement: a live affected-customer list and a holding statement within 15 minutes of SEV1/SEV2 so the front line is never guessing.
- Handover protocol for incidents beyond 4 hours: formal IC handover checklist, fatigue rule, and staffing of a second shift.
15. Customer communications and status page policy (depends on: 5, 14)
Customers currently learn of outages from their own monitoring and hear from whoever happens to be around. Replace that with a timed, owned, pre-approved process.
The Comms Lead is the single author; templates remove the need to write under pressure.
- Timing commitments: status page posted within 15 minutes of SEV1 declaration and 30 minutes for SEV2; updates every 30/60 minutes; resolution notice within 30 minutes of mitigation; customer-facing summary within 5 business days for SEV1.
- Pre-approve 12–15 templates with Legal and Comms (degradation, delay in settlement, API errors, security event, third-party failure) so nothing needs legal review mid-incident.
- Subscription-based status page with per-component granularity mapped to the customer capabilities from S11, plus an email/webhook/RSS feed.
- Tiered outreach: top 100 accounts get a direct named call or email from their account manager within 30 minutes of SEV1, with a briefing pack from the Comms Lead; long tail gets the status page and a proactive email.
- Rules of language: state impact and next update time, never speculate on cause, never assign blame to a vendor before facts are confirmed.
- Run a quarterly customer-perception check with the top accounts on whether comms were timely and useful.
16. Regulatory, partner and legal notification playbook (depends on: 5, 15)
In payments, some incidents are reportable and the clock starts at detection. Build the assessment into the process so it is never an afterthought.
Work with Legal, Compliance and the CISO to produce a decision tree and contact matrix.
- Map obligations: NYDFS Part 500 (72-hour cybersecurity event notification), state breach laws, GLBA/FTC Safeguards, PCI DSS if card data is in scope, sponsor-bank and card-network contractual notice windows, and any FinCEN/OFAC implications.
- Add a mandatory regulatory-assessment checkpoint to every SEV1 and every security-related SEV2, owned by the Executive Liaison, completed within 2 hours of declaration and recorded even when the answer is "not reportable".
- Build the contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel, with 24x7 numbers and named backups.
- Pre-draft notification letters and hold them under legal privilege review.
- Check customer contracts for bespoke notification SLAs (often 1–4 hours for enterprise accounts) and encode them in the customer tiering.
- Test the playbook once per quarter as part of the simulation programme.
17. SLA credit and financial impact workflow (depends on: 5, 15)
Link incidents to money so severity, credits and prioritisation stay consistent — and so Finance stops being surprised.
Make credit calculation an automated output of the incident record, not a negotiation.
- Define the availability measurement method per contract, per component, and agree it with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and the per-capability telemetry; generate a proposed credit schedule within 5 business days of resolution.
- Decide the posture: proactive credits for the top tier (reputation upside) versus claims-based for the rest; document the approval chain.
- Track credits per incident and per root-cause family; feed a quarterly report showing which reliability investments would have prevented which credits.
- Set a target: reduce credits from $1.3M to under $400k in year one, and use that delta as the ongoing business case.
18. Postmortem standard and blameless review forum (depends on: 5, 6)
Replace "some incidents, various formats" with a mandatory, single-format, blameless process with fixed deadlines.
The discipline is in the deadlines and the forum, not in the template.
- Mandatory for: every SEV1 and SEV2, every incident where a customer detected it first, every incident over 2 hours, every repeat of a known cause, and every near-miss involving the ledger. Optional but templated for SEV3.
- Fixed timeline: draft within 3 business days, peer review within 5, published company-wide within 10. The IC owns delivery; the owning team's manager is accountable.
- One template: timeline, customer and financial impact, detection analysis (why not sooner), response analysis (why mitigation took as long as it did), contributing factors, what went well, action items with owner and due date.
- Blameless rules in writing: describe systems and decisions in the context available at the time; no individual named as a cause; HR and management commit that postmortems are never used in performance reviews.
- Weekly 60-minute Incident Review Board: reviews all postmortems from the prior week, challenges quality, ratifies severity, and approves or rejects action items. Attendance by engineering directors is mandatory.
- Publish a searchable postmortem library and a quarterly "top five recurring causes" analysis.
19. Action item ownership and tracking system (depends on: 18, 12)
11 of 64 actions closed is the single clearest symptom of a process nobody enforces. Give actions the same status as customer commitments.
Track them where engineering work already lives, with visible escalation.
- Every action gets: a named individual owner (not a team), a priority class, a due date and a Jira ticket auto-created from the postmortem.
- Priority classes with hard SLAs: P0 prevents recurrence of a SEV1, due in 30 days; P1 in 60 days; P2 in 90 days. P0s are committed into the next sprint before any roadmap work.
- Capacity rule: teams reserve a standing 15–20% of sprint capacity for reliability and incident actions. Without reserved capacity, the actions will not land.
- Escalation ladder for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Monthly reporting of closure rate by team in the engineering leadership review; include it in manager performance objectives.
- Target: 90% of P0/P1 actions closed on time within two quarters.
20. Runbooks, major-incident playbooks and the on-call readiness bar (depends on: 4, 8)
Nobody can respond well to unfamiliar systems at 3 a.m. without runbooks — and poor runbooks are a real part of the pager resistance.
Define a minimum readiness bar that a service must meet before it is allowed to page anyone.
- Readiness checklist per Tier 0/1 service: current architecture diagram, dependency map, dashboard link, alert-to-runbook mapping, rollback procedure, feature-flag kill switches, escalation contacts, and a data-loss/latency impact statement.
- Write major-incident playbooks for the top failure modes derived from S2: shared PostgreSQL ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank/third-party outage, settlement-window breach, and suspected security compromise.
- Prioritise the shared ledger: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and a clearly stated data-loss tolerance (RPO/RTO) signed off by the exec.
- Runbooks must be tested at least twice a year in a drill; untested runbooks are marked stale in the catalog.
- Enforcement: a service without readiness sign-off cannot create paging alerts, and the gap is reported to its director.
21. Training, certification and the commander academy (depends on: 6, 13, 14, 18)
Command is a skill, not a title. Build a certification path so 24x7 coverage is staffed by people who have practised.
Use a tiered curriculum with real assessment.
- **Scribe** (2 hours): timeline discipline and tooling. The entry point for everyone.
- **Responder** (half a day): severity scale, declaration, escalation, runbook use, comms hygiene. Mandatory for every engineer joining an on-call rotation.
- **Incident Commander** (two days plus shadowing): command presence, delegation, decision-making under uncertainty, severity calls, handover, exec management. Requires two shadowed incidents and one simulated SEV1 before certification.
- **Communications Lead** (one day): status-page writing, customer tiering, legal boundaries, regulator triggers.
- Certification is valid 12 months and renewed via a simulation; the register of certified people is an audit artefact.
- Add on-call onboarding per team: a new joiner shadows two shifts before holding primary, and never holds primary in their first 90 days.
22. Simulation programme: game days, drills and wheel of misfortune (depends on: 21, 12, 20)
The process must be rehearsed before it meets a real SEV1. Simulations also build the commander pool and expose runbook gaps cheaply.
Run a standing calendar rather than one-off exercises.
- Monthly 60-minute tabletop ("wheel of misfortune") per engineering group, using a real past incident from the 31.
- Quarterly full-scale game day in production or a production-like environment: regional failover, ledger replica promotion, dependency failure, with the whole role structure activated and timed.
- Twice-yearly unannounced paging drill to measure real acknowledgement times at night.
- One security-incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Every exercise produces a lightweight postmortem and action items in the same system as real incidents.
- Measure and publish drill metrics: time to IC, time to first status update, time to correct mitigation decision.
23. Pilot with wave 0 teams (depends on: 22, 9, 11)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Run a six-week pilot with tight measurement and a public verdict.
- Select 5–6 teams: core payments, ledger/database, platform/Kubernetes, API gateway, plus two of the 12 teams already on-call.
- Activate the full stack for them: new severity scale, Duty IC rotation, single paging tool, alert budget, status-page policy, mandatory postmortems, paid on-call.
- Hold a weekly pilot retro; expect and document 20–40 process defects, and fix them in the standard before rollout.
- Validate the hard questions: does the 5-minute IC rule hold at 3 a.m.? Do cross-team pulls get answered? Is the severity tree unambiguous?
- Exit criteria: MTTD under 10 minutes for pilot services, IC assigned within 5 minutes in 95% of incidents, page volume down 50%, all postmortems on time, positive on-call sentiment.
- Publish a one-page pilot result to the whole company — this is the main adoption argument for the remaining teams.
24. Metrics, dashboards and the review cadence (depends on: 12, 18, 23)
Instrument the process itself, so improvement is visible and the audit has evidence of monitoring and review.
Define a small set of metrics with owners and a fixed meeting rhythm.
- Response metrics: MTTD, time to declare, time to IC assigned, MTTA, MTTM, MTTR, incidents per month by severity, % detected by customers first.
- Quality metrics: page volume per person per week, alert actionability rate, page budget breaches, postmortem on-time rate, action closure rate and ageing.
- Business metrics: SLA credits paid, availability against 99.95% per capability, error-budget consumption, repeat-incident rate.
- People metrics: on-call load distribution across teams, out-of-hours pages per person, on-call sentiment and attrition among on-call staff.
- Cadence: weekly Incident Review Board (postmortems and actions), monthly Reliability Review (metrics per team, alert hygiene, on-call load), quarterly Executive/Board review (credits, trends, investment asks), annual policy review.
- Every metric gets a target and a named owner; dashboards are self-serve and public inside the company.
25. Wave rollout across all 28 teams with readiness gates (depends on: 23, 24)
Roll out in four waves of six to eight teams, every three weeks, ordered by criticality. Each wave passes an explicit gate rather than a deadline.
Gates keep quality high and make the standard credible.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with 6+ certified responders, one IC candidate nominated, one drill passed.
- Assign each wave a named coach from the program team for three weeks of hands-on support.
- Gate criteria are checked and signed by the director; teams that fail are re-scheduled, not waived.
- Freeze legacy tooling per wave: after onboarding, the old alerting paths are disabled, not left as a fallback.
- Publish a live adoption scoreboard by team so progress is social, not administrative.
26. SOC 2 control mapping, evidence automation and internal dry-run audit (depends on: 18, 19, 25)
Design the process so audit evidence is a by-product of doing the work, then test it before the auditors do.
Engage the auditor early to confirm the interpretation of controls.
- Map the process to the Trust Services Criteria: CC7.3 and CC7.4 (incident identification, response, recovery), CC7.2 (monitoring), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), plus availability criteria A1.2.
- Produce and approve formal policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard — versioned, signed, annually reviewed.
- Automate evidence: incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, postmortem library, action-item closure reports, training and certification register, drill records.
- Confirm the observation window with the auditor and ensure the process is operating for a minimum of three months before fieldwork.
- Run an internal dry-run audit at month six: sample 15 incidents and walk the full evidence chain; fix gaps with 8 weeks to spare.
- Keep a remediation log for any incident where the process was not followed, with the corrective action — auditors respond better to documented exceptions than to a claim of perfection.
27. Change management, incentives and communications campaign (depends on: 3, 9, 23)
Run this in parallel from day one. The process will be judged by engineers on fairness, and by executives on visible results.
Communicate the deal explicitly and repeatedly.
- The deal in one sentence: **you are paid for on-call, you are paged only for what you own, a trained commander runs the incident, and your postmortem actions get real sprint capacity**.
- Launch communications: CTO all-hands, per-team roadshows, a one-page process card for laptops, an internal wiki hub and a Slack support channel with a 4-hour answer SLA.
- Recognition: incident-response contribution in promotion criteria and performance frameworks, quarterly awards for best postmortem and biggest alert-noise reduction, public thanks after every SEV1.
- Manager accountability: adoption, alert hygiene, action closure and on-call load in each engineering manager's quarterly objectives.
- Handle the exceptions: a written path for engineers who cannot do nights (caring responsibilities, health), covered by stipended volunteers elsewhere.
- Track sentiment quarterly and publish the results, including bad news, to keep credibility.
28. Program risk register and contingency planning (depends on: 1)
Name the ways this program fails and pre-commit the response. Review it monthly in the steering group.
The main risks are predictable.
- **Volunteer shortfall for the IC pool**: contingency is to make command a rostered duty for engineering managers and senior engineers until the pool reaches 25.
- **Compensation not approved in time**: fall back to time-off-in-lieu plus a phased stipend, but do not launch mandatory night on-call without some compensation.
- **Tool migration slipping**: keep the single-queue requirement and cut scope on the incident-management layer, not on paging consolidation.
- **Alert pruning causing a missed incident**: prune from paging to ticket first, observe for 30 days, then delete; keep a recovery path.
- **Burnout or attrition among the 12 experienced on-call teams**: monitor load weekly and cap individual page counts.
- **A major SEV1 mid-rollout**: pre-agree that the program lead becomes a full-time responder and the wave schedule slips by one wave, with the steering group informed the same day.
29. Continuous improvement, maturity roadmap and post-audit sustainability (depends on: 24, 25, 26)
Protect against the classic failure: the process decays once the audit is passed. Build the second-year plan before the first year ends.
Set a maturity model and a forward roadmap with owners.
- Quarterly process retrospective with the IC pool: what in the process itself slowed us down, what needs simplifying, is the severity scale calibrated?
- Re-baseline targets every six months; a process that hits all targets is under-ambitious.
- Year-two roadmap candidates: follow-the-sun coverage cell, automated mitigation and self-healing for the top three recurring causes, error-budget policy that gates releases, per-customer real-time impact reporting, and blast-radius reduction for the shared ledger cluster (the largest single structural risk).
- Move from lagging metrics (MTTR) to leading ones (error-budget burn, near-miss rate, drill performance).
- Make the annual policy review, certification renewal and drill calendar permanent calendar items owned by the Head of Reliability, independent of the audit cycle.
- Report to the board quarterly on availability, credits and incident trends so the process keeps executive attention after SOC 2 is signed.
Previous Proposal 2 (ID: 7ee53cee-6357-4ed7-bd99-535e92e8b1e6, Agent: gpt5.6-sol_initial_2, LLM: openai/gpt-5.6-sol):
Estimated Complexity: high
Success Metrics: - Within 7 days, every suspected SEV0–SEV2 has one incident record, one channel, and a named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named incident commander for more than 10 minutes.
- By day 30, 100% of Tier 0 and Tier 1 services have a named owner, primary escalation, secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and subject-matter coverage.
- By day 120, 100% of production services have sustainable ownership and tested escalation paths.
- At least 95% of SEV0 and SEV1 pages are acknowledged within 5 minutes by month 3.
- At least 95% of SEV2 pages are acknowledged within 10 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to 10 minutes by day 90 and 5 minutes by month 6.
- The proportion of incidents first detected by customers falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to below 90 minutes by day 120 and below 60 minutes by month 6.
- At least 95% of qualifying incidents meet their initial customer-communication deadline by month 3.
- At least 95% of published incidents meet their required update cadence by month 3.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without loss of Tier 0 or Tier 1 detection coverage.
- Monthly pages fall from 3,400 to no more than 1,500 by day 90, with alert actionability and missed-detection reviews used as countermeasures against unsafe suppression.
- 100% of new paging alerts satisfy the owner, runbook, dashboard, action, severity, and escalation quality rules by day 60.
- 100% of required SEV0 and SEV1 postmortems are drafted within 3 business days and reviewed within 5 business days by month 2.
- At least 90% of postmortem actions are completed by their approved due dates by month 6.
- All 53 currently open historical actions are triaged within 30 days; all unaccepted high-risk items are completed within 90 days.
- Repeat incidents with the same unaddressed contributing factor decline by at least 50% within 6 months.
- Every primary rotation has at least six trained responders or a documented, time-limited executive exception by day 120.
- No responder is routinely scheduled more frequently than one primary week in six by day 120.
- Two end-to-end cross-company exercises, including regional and ledger scenarios, are completed before the audit, with all critical findings assigned and tracked.
- Monthly availability meets or exceeds the 99.95% contractual target by month 6, with exceptions reviewed at the executive reliability meeting.
- SLA credits decline by at least 50% on an annualized trailing basis by month 8.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (19):
1. Establish ownership, authority, and funding
Launch the program within 48 hours under an executive sponsor. Give one program owner authority to standardize incident management across all 28 teams.
- Name the CTO or equivalent as executive sponsor and a Head of Incident Management or Reliability as directly accountable owner.
- Form a working group with Engineering, SRE or Platform, Product, Support, Customer Success, Communications, Security, Legal, Compliance, Risk, HR, Finance, and Internal Audit.
- Approve authority for an incident commander to stop deployments, roll back releases, disable features, shift traffic, invoke continuity plans, and pause payment processing when integrity is at risk.
- Preserve financial controls. The incident commander may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access requirements.
- Fund paging tools, compensation, training, observability work, exercises, and dedicated reliability capacity.
- Reserve engineering capacity for incident remediation. Start with 10% of capacity and adjust through quarterly risk reviews.
- Record the current baselines: 31 customer-impacting incidents, 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M in credits, 3,400 monthly alerts, 85% noise, and 11 of 64 actions closed.
- Maintain a risk register for staffing gaps, shared-ledger concentration, regional failover, alert coverage, third parties, and audit readiness.
2. Install immediate minimum controls (depends on: 1)
Put an interim process in place during the first seven days. Do not wait for tool consolidation, policy perfection, or the SOC 2 audit.
- Publish a one-page interim severity guide and incident declaration procedure.
- Establish one continuously monitored incident declaration path through chat, telephone, and the paging system.
- Create a standard incident channel, conference bridge, incident document, and event naming convention.
- Staff an interim primary and backup incident commander at all times. Compensate this duty retroactively under the final compensation policy.
- Give trained duty personnel access to the status page, paging system, dashboards, support queue, service catalog, and emergency contacts.
- Require an incident commander to be named within 10 minutes for every suspected major incident.
- Direct Support to escalate credible customer reports immediately rather than waiting for engineering confirmation.
- Triage all 53 open historical postmortem actions. Complete, re-plan, or formally risk-accept the items affecting ledger integrity, payment duplication, regional resilience, security, and detection first.
- Hold a daily 15-minute operational review until permanent controls are working.
3. Create the service and dependency catalog (depends on: 1)
Build a reliable ownership map for all production services and customer journeys. This is the basis for paging, escalation, impact assessment, and audit evidence.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, data stores, queues, external processors, banking partners, and customer-facing endpoints.
- Assign each component a single accountable team, primary responder group, secondary escalation group, engineering manager, and product owner.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 according to financial integrity, customer impact, dependency centrality, and contractual obligations.
- Treat the ledger, payment orchestration, authentication, settlement, reconciliation, and critical shared infrastructure as Tier 0 or Tier 1.
- Map every important customer journey to its service, database, cloud-region, and third-party dependencies.
- Record SLOs, RTOs, RPOs, data classification, dashboards, runbooks, deployment controls, feature flags, and failover methods.
- Assign separate but coordinated responders for the ledger application and the shared PostgreSQL platform.
- Document whether each service is active-active, active-passive, or region-bound. Identify dependencies that make nominal regional redundancy ineffective.
- Make missing ownership or missing runbooks a release-blocking risk for Tier 0 and Tier 1 services.
4. Adopt severity and incident lifecycle standards (depends on: 1)
Approve one impact-based severity model for operational, security, data, and third-party incidents. When evidence is incomplete, start at the higher credible severity and downgrade later.
- SEV0, crisis: Use for actual or credible unauthorized, lost, duplicated, or corrupted movement of money; ledger integrity loss; material security compromise; material data exposure; both-region failure; or an event likely to require crisis or regulatory management. Page all roles immediately, engage executives, Security, Legal, Compliance, and Risk, and consider pausing payment activity.
- SEV1, critical: Use for widespread inability to initiate, process, settle, or reconcile payments; a core journey failing without a viable workaround; material regional impact; fast SLA-budget exhaustion; or an imminent integrity risk. Staff all incident roles, notify the executive duty officer, and publish customer communications.
- SEV2, major: Use for a material customer subset, one or more critical customers, significant degradation with a workaround, partial transaction failure, or a likely contractual impact. Assign an incident commander and subject-matter responders; add communications and scribe roles whenever customers are affected.
- SEV3, minor: Use for localized, low-impact degradation with no financial-integrity, security, regulatory, or material contractual risk. The owning team leads the response and keeps an internal record; external communication is not normally required.
- Base severity on actual or credible impact, not the seniority of the reporter, number of alerts, or presumed complexity of the fix.
- Permit any employee to declare an incident. Only the incident commander may lower severity after recording the evidence and rationale.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of customer impact. Define resolution only after stability, backlog processing, transaction recovery, and required ledger reconciliation are complete.
- Start measurement from the earliest reliable indication of impact, including telemetry, customer reports, and partner notifications.
5. Define roles and sustainable 24x7 staffing (depends on: 3, 4)
Separate command from technical remediation. This allows trained commanders to coordinate any incident without asking engineers to debug code they do not own.
- Incident commander: Owns severity, priorities, role assignment, escalation, decision cadence, mitigation strategy, handoffs, and final closure. One person has command at a time.
- Communications lead: Owns internal notices, status-page updates, account-manager briefs, approved customer language, and coordination with Legal or regulators.
- Scribe: Maintains a timestamped timeline of observations, decisions, commands, owners, and status changes. Automation may assist but does not replace human validation for SEV0 and SEV1.
- Subject-matter responders: Diagnose and mitigate only services or domains for which they have accepted ownership, training, access, and runbooks.
- Executive duty officer: Removes organizational obstacles and approves exceptional business decisions. This role does not take command unless a formal transfer occurs.
- Security, Legal, Compliance, Support, Vendor Management, and Business Continuity join according to predefined triggers.
- Create a company-wide incident-command rotation with at least eight certified primary commanders and eight qualified backups. Use weekly rotations with explicit handoffs.
- Create similarly sustainable communications and scribe pools using Engineering Operations, Support, Customer Operations, and Communications personnel.
- Group service responders into approximately 8–12 coherent product or platform domains rather than creating 28 fragile rotations. Each domain rotation should normally contain at least six trained responders.
- Do not place an engineer into another team's responder pool without training, access, runbooks, shadow shifts, and explicit acceptance by both teams.
- Maintain dedicated database-platform and ledger-application escalation coverage for the shared PostgreSQL environment.
- Require distinct people for commander, communications, and primary technical lead during SEV0 and SEV1 incidents.
- Require a verbal and written handoff when any incident role changes. Record the exact time and the new role owner.
6. Implement compensation and fatigue safeguards (depends on: 5)
End unpaid on-call before expanding coverage. Treat availability, interrupted personal time, and overnight recovery as compensable work.
- Pay a fixed stipend for each primary on-call week and a secondary stipend equal to a defined percentage of the primary amount.
- Pay a higher holiday stipend. Apply overtime and call-out rules to non-exempt employees as required by law.
- Give exempt employees a minimum call-out credit or equivalent paid recovery time for material after-hours work.
- Provide a paid recovery day after prolonged overnight work, a SEV0, or a qualifying SEV1. Managers must arrange daytime coverage rather than expecting normal output.
- Have HR, Finance, and employment counsel publish dollar amounts, tax treatment, eligibility, and payroll procedures within 14 days. Apply the policy consistently across teams and locations.
- Target rotations no more frequent than one week in six. Exceptions require a time-limited staffing plan and executive risk acceptance.
- Avoid consecutive primary and secondary weeks. A person must not be primary for two simultaneous domain rotations.
- Track after-hours pages, sleep interruptions, swaps, missed acknowledgements, and reported burnout by rotation.
- Trigger a staffing or alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
- Permit responders to declare themselves temporarily unfit after overnight work without performance penalty.
7. Consolidate incident and paging tooling (depends on: 4)
Create one operational system of record while migrating safely from the six current alerting tools. Consolidation must reduce ambiguity without creating a monitoring gap.
- Select one enterprise paging and escalation platform and one integrated incident record.
- Initially ingest events from all six tools. Deduplicate, correlate, and route them through the new platform before retiring sources.
- Integrate paging with chat, conference bridges, ticket tracking, the service catalog, observability tools, and the customer status page.
- Automatically capture declaration time, acknowledgements, role assignments, severity changes, messages, decisions, mitigated time, and resolved time.
- Use role-based access, multifactor authentication, break-glass controls, immutable audit logs, and periodic access reviews.
- Provide mobile and telephone fallback paths if chat, identity, or the primary paging tool is unavailable.
- Test paging, escalation, status publication, and conference access every week.
- Retire a legacy alert path only after its signals have named owners, successful end-to-end tests, and at least two weeks of verified operation in the new platform.
8. Improve detection and enforce alert quality (depends on: 3, 7)
Shift detection toward customer journeys, payment outcomes, and ledger integrity. Infrastructure metrics alone will not solve the current customer-first detection problem.
- Instrument payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, and reporting with SLOs and business-level success metrics.
- Run external synthetic transactions and API checks from outside the production boundary and from both AWS regions.
- Monitor transaction failure rates, processing latency, queue age, unprocessed volume, reconciliation breaks, unexpected ledger balances, duplicate identifiers, regional asymmetry, and third-party response quality.
- Correlate application telemetry with Kubernetes, AWS, PostgreSQL, network, deployment, and feature-flag events.
- Route high-priority support cases and credible partner notifications into the same incident declaration path within five minutes.
- Define noise as a page that is duplicate, informational, unactionable, non-production, or requires no timely human action.
- Require every paging alert to name an owner, affected service, urgency, customer or SLO risk, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Send non-urgent conditions to a ticket queue rather than a pager.
- Run new alerts in shadow mode for at least seven days unless an emergency risk exception is approved. Test both firing and recovery behavior.
- Review any alert with less than 50% actionability or more than three firings in seven days within two business days.
- Never silently disable a noisy alert. Verify compensating detection, record the decision, and assign a correction owner first.
- Review alert actionability, false positives, missed detection, and page load with every responder group each month.
9. Codify acknowledgement and escalation paths (depends on: 3, 4, 5, 7, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts when a qualifying signal or customer report enters the system.
- For SEV0 and SEV1, page the owning primary immediately; page the secondary after five unacknowledged minutes; page the domain manager and company incident commander at 10 minutes; and engage the executive duty officer by 15 minutes.
- For SEV2, require primary acknowledgement within 10 minutes and incident-command assignment within 15 minutes. Escalate to the secondary and manager when either target is missed.
- For SEV3, require acknowledgement within 30 minutes when immediate production action is needed. Otherwise create a prioritized work item.
- Automatically page the company incident commander for any credible integrity or security concern, cross-team event, customer-visible Tier 0 failure, regional event, or unresolved ownership question.
- If impact remains unknown after 15 minutes, raise severity rather than waiting for certainty.
- Let the incident commander summon dependency owners, cloud support, database support, payment processors, banking partners, and vendors through maintained escalation contacts.
- Test vendor contacts and premium-support entitlements quarterly.
- Route alerts with no valid owner to the central command rotation, then treat the missing ownership record as a control defect.
- Require human acknowledgement. Delivery to a device or chat channel does not count.
- Record every missed acknowledgement, failed escalation, and manual contact workaround for review.
10. Standardize live incident execution (depends on: 4, 5, 7, 9)
Give responders one concise operating procedure for the first minutes through resolution. Prioritize limiting customer and financial harm before proving a root cause.
- Open a dedicated channel, bridge, incident record, and timeline immediately for SEV0 through SEV2.
- Have the incident commander state severity, known impact, current hypothesis, immediate objective, assigned roles, and next update time.
- Freeze unrelated production changes during SEV0 and SEV1 incidents. Record exceptions approved by the incident commander.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Use pre-approved runbooks for region failover, Kubernetes recovery, PostgreSQL failover, credential rotation, queue recovery, and payment suspension.
- Guard against split-brain, replay, duplication, and out-of-order processing during regional or database recovery.
- Require reconciliation and controlled backlog processing before declaring payment or ledger incidents resolved.
- Keep diagnosis and mitigation workstreams separate when enough responders are available.
- State decisions and owners aloud and in the incident record. Avoid unrecorded direct-message command paths.
- Require stability for a severity-specific observation period before closure. Reopen the incident if impact recurs during that period.
- Conduct an explicit operational handback to the owning team, Support, and Customer Success.
11. Standardize internal, customer, and regulatory communications (depends on: 4, 5, 7, 10)
Communicate known impact early without waiting for a root cause. Use approved facts, acknowledge uncertainty, and give the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish an initial customer status within 15 minutes when disclosure is operationally and legally appropriate, then update every 15 minutes.
- For SEV1, notify internal stakeholders within 15 minutes, publish an initial customer status within 15 minutes, and update at least every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send an initial notice within 30 minutes. Update at least every 60 minutes.
- Do not normally publish SEV3 events. Notify specifically affected customers if contracts or material impact require it.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. State affected capabilities, customer symptoms, workarounds, regions, and next update time.
- Do not speculate about root cause, blame, security scope, recovery time, or data integrity.
- Give account managers a single approved briefing and an affected-customer list. Prohibit contradictory or improvised incident explanations.
- Maintain templates for outages, delays, data-integrity investigation, third-party failure, regional failure, security events, and resolution.
- Issue a resolution notice only after operational recovery and required reconciliation. Provide a customer-facing incident summary within five business days for qualifying events.
- Have Legal and Compliance maintain a jurisdiction, regulator, sponsor-bank, network, cyber-insurer, partner, and contract notification matrix.
- Where applicable, explicitly track the current New York cybersecurity-event notification clock, including the 72-hour requirement, without assuming every incident is reportable.
- Have Legal record the reportability decision, decision time, evidence, approver, deadline, and submission confirmation.
- Allow Security or Legal to limit public detail during an active threat, but require the reason and an alternative stakeholder plan to be recorded.
- Coordinate service-credit calculations and contractual notices with Finance and Customer Success from the same incident record.
12. Make postmortems mandatory and actionable (depends on: 4, 7, 10)
Use postmortems to improve systems and controls, not to assign personal blame. Keep performance or misconduct processes separate from the learning review.
- Require a postmortem for every SEV0 and SEV1.
- Require one for a SEV2 that affected customers, incurred credits, breached an SLO or contract, involved financial or data integrity, repeated a prior failure, exposed a control gap, or lasted more than two hours.
- Permit incident command, Security, Compliance, or the service owner to require a review for a near miss.
- Produce a factual draft within three business days and hold the cross-functional review within five business days.
- Use one template covering executive summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, what worked, what failed, and lessons.
- Include why monitoring did or did not detect the event before customers.
- Avoid a single-root-cause assumption. Examine technical, organizational, process, dependency, testing, and incentive factors.
- Give every action one owner, due date, priority, expected risk reduction, verification method, and linked engineering item.
- Classify actions as containment due within 7 days, corrective work due within 30 days, or strategic work normally due within 90 days.
- Require director approval and documented residual-risk acceptance for overdue high-risk actions.
- Verify effectiveness after implementation. Closing a ticket without evidence does not close the action.
- Publish broadly useful reviews internally. Maintain access-restricted versions for security, privacy, personnel, or legally privileged details.
13. Measure performance and review it routinely (depends on: 7, 8, 11, 12)
Use outcome, process, quality, and human-sustainability measures together. Do not reward teams for suppressing declarations or hiding incidents.
- Measure detection time from first impact to first internal signal, declaration time, acknowledgement time, role-staffing time, mitigation time, resolution time, and recurrence.
- Report both median and 90th percentile. Break results down by severity, service tier, customer journey, region, detection source, and owning domain.
- Track customer-first detection, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, and role conflicts.
- Track availability, error-budget consumption, failed-payment volume, delayed value, reconciliation breaks, impacted customers, contractual breaches, and service credits.
- Track alert volume, actionability, duplicates, after-hours pages, missed pages, pages per responder, and tool-source distribution.
- Track required postmortems completed on time, actions completed by due date, action age, verified effectiveness, and repeat contributing factors.
- Track rotation size, on-call frequency, swaps, recovery days, attrition signals, and quarterly responder sentiment.
- Hold a weekly operational review for recent incidents, overdue actions, alert problems, and upcoming risk.
- Hold a monthly executive reliability review covering trends, investment decisions, accepted risks, and SLA exposure.
- Hold a quarterly resilience and control review with Security, Compliance, Risk, Internal Audit, and Product leadership.
- Use team scorecards to direct investment and assistance, not individual performance penalties.
- Reconcile dashboard data against a monthly sample of incident records and customer cases to detect metric gaming or missing incidents.
14. Train and certify participants (depends on: 4, 5, 9, 10, 11, 12)
Train people before assigning full independent duty. Use paid working time for training, shadowing, exercises, and certification.
- Train all employees to recognize impact, declare an incident, and find the incident channel and status page.
- Train engineers and Support on severity, escalation, customer-report handling, evidence preservation, and financial-integrity precautions.
- Certify incident commanders through instruction, tabletop exercises, shadow incidents, and observed command performance.
- Train communications leads in status writing, contractual communications, regulator escalation, and avoiding unsupported claims.
- Train scribes in timestamping, decision capture, evidence hygiene, and separating fact from hypothesis.
- Require subject-matter responders to demonstrate dashboard, runbook, rollback, failover, and access competence for their assigned domain.
- Add training to new-hire onboarding and repeat role-specific certification annually.
- Appoint an incident-management champion in each of the 28 teams to collect feedback and support local adoption.
- Conduct listening sessions focused on pager fairness, cross-team boundaries, tooling friction, and psychological safety.
- Publish that command duty is process coordination, not responsibility for understanding or repairing another team's code.
15. Pilot and expand the on-call model (depends on: 3, 5, 6, 8, 9, 14)
Pilot the model on the highest-risk customer journeys before expanding it. Correct staffing, alert, access, and compensation defects at each stage gate.
- Start with the ledger, payment orchestration, Kubernetes platform, PostgreSQL platform, authentication, settlement, and Support intake.
- Run the command, communications, and domain rotations in parallel with existing paths for two weeks.
- Verify primary and secondary coverage, handoffs, access, runbooks, paging, conference access, status publication, and compensation processing.
- Require at least two shadow shifts before independent primary duty.
- Review every pilot page within one business day for routing accuracy, actionability, responder load, and missing context.
- Expand by customer journey and dependency domain, not by arbitrary team order.
- Provide central first-line triage if useful, but keep technical remediation with the accepted service owner.
- Do not use contractors or a managed service as the sole incident commander or sole owner of payment and ledger remediation.
- Permit temporary shared domain rotations only after service owners document training, access, runbooks, and escalation boundaries.
- Set an executive-reviewed deadline and remediation plan for any production service that cannot provide sustainable 24x7 ownership.
16. Exercise regional, ledger, and communication failures (depends on: 10, 11, 14, 15)
Validate the process under realistic conditions before relying on it. Begin in tabletop and staging environments, then use controlled production tests where risk permits.
- Run a company-wide incident-command tabletop within 30 days of policy approval.
- Exercise loss of one AWS region, Kubernetes control-plane degradation, shared PostgreSQL failure, payment-processor failure, queue backlog, credential compromise, and suspected duplicate payments.
- Exercise a simultaneous operational and security event to test command boundaries and disclosure control.
- Exercise status-page failure and loss of the primary chat or paging provider.
- Exercise overnight staffing, role handoff, executive escalation, account-manager messaging, and a potential regulator-notification decision.
- Validate backups, restore procedures, RPO, RTO, failover prerequisites, and post-recovery reconciliation.
- Do not inject uncontrolled changes into the production ledger. Use replicas, staging, simulations, or tightly governed production tests.
- Record exercise observations as tracked actions under the same ownership and due-date rules as incident actions.
- Run at least one domain exercise per quarter and two cross-company exercises before the SOC 2 audit.
17. Execute a time-boxed enterprise rollout (depends on: 2, 6, 7, 11, 12, 13, 15)
Use fixed implementation waves so the audit deadline does not become the start date. Report progress weekly and escalate missed stage gates as business risks.
- Days 0–7: Establish governance, interim command coverage, one declaration path, provisional severity, and daily operational reviews.
- By day 14: Approve the core policy, role definitions, communications timings, compensation design, and historical-action triage.
- By day 30: Complete Tier 0 ownership, certify the first command roster, begin the on-call pilot, enable standard incident records, and run the first tabletop.
- By day 60: Provide 24x7 coverage for all Tier 0 and Tier 1 customer journeys, integrate the six alert sources, and enforce postmortem tracking.
- By day 90: Migrate critical paging, implement customer-journey detection, complete status and regulatory playbooks, and materially reduce alert noise.
- By day 120: Assign sustainable ownership and escalation for every production service and complete the first controlled regional or continuity exercise.
- By day 180: Complete tool retirement decisions, verify action closure, rerun weak scenarios, and demonstrate improving detection and mitigation trends.
- In month 7: Conduct a mock audit and executive readiness review, leaving at least one month to correct evidence or operating defects.
- Use exception records with owners, expiry dates, compensating controls, and executive approval. Do not allow indefinite verbal exceptions.
18. Build SOC 2 evidence as the process operates (depends on: 1)
Design evidence collection at the start rather than reconstructing it before the audit. Demonstrate both control design and sustained operation.
- Map the incident process to applicable SOC 2 criteria with Compliance and the auditor, including detection, response, communication, change management, access, availability, and corrective action.
- Maintain approved, version-controlled policies, procedures, severity definitions, role descriptions, and exception records.
- Preserve rotation schedules, compensation activation, training attendance, certification, paging tests, access reviews, and exercise results.
- Preserve incident declarations, timestamps, role assignments, communications, decisions, status updates, postmortems, and corrective-action evidence.
- Record regulatory and contractual notification assessments, including decisions that no notification was required.
- Define retention, confidentiality, legal-hold, and access requirements for operational and security records.
- Sample evidence monthly and trace incidents from initial signal through action verification.
- Have Internal Audit or an independent control owner test the process in months 4 and 6.
- Correct control failures through tracked actions rather than editing historical records.
- Conduct the formal mock audit in month 7 using the same evidence populations expected for the external audit.
19. Sustain accountability and continuous improvement (depends on: 13, 17, 18)
Make incident management an operating discipline rather than an audit project. Keep policy, staffing, tools, and investment aligned with changing customer and system risk.
- Assign permanent owners for the incident policy, paging platform, status page, service catalog, training program, and metrics.
- Review severity thresholds, communication timings, compensation, and staffing at least annually and after material incidents.
- Use incident trends to prioritize architectural work on the shared ledger, regional independence, deployment safety, dependency isolation, and graceful degradation.
- Review repeat incidents and repeat contributing factors quarterly. Require executive action when remediation repeatedly loses priority.
- Survey responders quarterly and publish actions addressing fatigue, fairness, psychological safety, and tool friction.
- Recognize effective incident leadership, early declaration, useful postmortems, and preventive work.
- Prohibit retaliation for good-faith incident declaration or escalation.
- Provide the board or risk committee a quarterly summary of severe incidents, SLA exposure, regulatory events, overdue high-risk actions, and resilience investment.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
Previous Proposal 3 (ID: c051663c-fdf6-4324-bfbe-a5f42ed62eed, Agent: qwen3.8-max_initial_3, LLM: alibaba/qwen3.8-max):
Estimated Complexity: high
Success Metrics: - **MTTD reduced from 22 minutes to <5 minutes** within 6 months of full rollout.
- **Customer-first detection rate reduced from 40% to <10%** within 6 months.
- **MTTM for SEV1 incidents reduced from 3 h 10 min to <60 minutes** within 9 months.
- **Monthly SLA credit payouts reduced from ~$325K to <$100K per quarter** within 12 months.
- **Alert volume reduced from 3,400/month to <600 actionable alerts/month** within 6 months; signal-to-noise ratio >80%.
- **Postmortem completion rate: 100% of SEV1/SEV2 incidents** have a blameless postmortem within 5 business days.
- **Postmortem action-item completion rate >90% within 30 days** of the postmortem (up from ~17%).
- **Zero incidents with >15 minutes of unowned command** (down from 2 incidents with >1 hour).
- **100% on-call coverage**: all 28 teams staffed with primary + secondary on-call 24×7 within 14 weeks.
- **On-call compensation adopted**: 100% of on-call engineers receiving stipends and page pay; on-call satisfaction score ≥4/5 in quarterly survey.
- **Status-page first-update within 15 minutes for SEV1 and 30 minutes for SEV2**, 100% compliance.
- **SOC 2 Type II audit passed** at month 8 with zero incident-response findings.
- **All 260 engineers trained**; 56+ certified ICs and 28+ certified CLs active within 14 weeks.
- **Six alerting tools consolidated to one** within 6 months; legacy tools decommissioned.
- **Regulator notification process tested**: at least one tabletop exercise includes a NY DFS / FinCEN notification drill, and the legal/compliance playbook is documented and approved.
- **Quarterly IMC reviews held consistently** with published KPI dashboards and action-item tracking.
- **On-call participation resistance resolved**: <10% of engineers report 'unwilling to participate' in the 6-month pulse survey (baseline to be measured in S1).
Steps (13):
1. Assess Current State and Baseline Metrics
Build an evidence-based picture of the current incident management reality before designing anything new.
Collect and catalog the last 12 months of incident data: all 31 customer-impacting incidents, 3,400 monthly alerts, on-call coverage gaps across the 28 teams, and the 11 of 64 closed postmortem action items. Interview one lead from each of the 28 teams to surface pain points, political concerns (the 'carrying a pager for other teams' pushback), and tool sprawl.
Deliverables to produce:
- **Alert inventory**: which of the six alert tools feed which teams, alert volume per team, noise rate per tool, and overlap between tools.
- **Incident timeline analysis**: median detection-to-notification-to-mitigation-to-resolution times, who detected first (internal vs. customer), who mitigated, and where handoff gaps occurred.
- **On-call coverage map**: which 12 teams have on-call, which 16 do not, rotation length, compensation status, and escalation paths (or lack thereof).
- **Postmortem audit**: format variance, action-item tracking gaps, and the two incidents with no clear owner for over an hour.
- **Tooling and integration audit**: Kubernetes observability stack, six alerting tools, status-page provider, communication channels (Slack, email, phone), and any existing runbooks.
- **Compliance gap analysis**: SOC 2 Type II CC7.3/CC7.4 requirements vs. current practice, with a risk register for the eight-month window.
- **Peer benchmarking**: incident management practices at 3–4 comparable B2B fintech platforms (e.g., Plaid, Stripe, Adyen) for severity scales, on-call comp, and MTTR targets.
2. Secure Executive Sponsorship and Form the IM Governance Body (depends on: 1)
Anchor the program with visible, top-down authority so that 28 teams adopt changes they did not individually request.
The CEO email about 'outages we hear about from clients' is a ready-made mandate. Convert it into a formal sponsorship structure.
- Appoint an **executive sponsor** (CTO or VP Engineering) who owns the program end-to-end and reports to the CEO monthly.
- Create an **Incident Management Office (IMO)**: one dedicated senior incident management lead, one tooling/platform engineer, and one part-time data analyst.
- Establish an **Incident Management Council (IMC)**: one engineering manager from each of the 28 teams, plus the VP of Customer Success, a compliance lead, and a security lead. The IMC meets bi-weekly during rollout, monthly thereafter.
- Draft and circulate an **executive mandate memo** that states: incident response is a shared operational obligation, not a per-team favor; participation in on-call rotations is a condition of employment for production-facing roles; and the program is not optional pending the SOC 2 audit.
- Allocate a dedicated budget line for on-call compensation, tooling consolidation, status-page licensing, training, and external facilitation.
3. Define Severity Levels and Automatic Triggers (depends on: 1)
Replace the current ad-hoc triage with a five-level severity taxonomy that every engineer, support agent, and account manager can apply in under 60 seconds.
- **SEV1 – Critical**: Ledger data corruption or loss, complete payment processing halt, confirmed data breach affecting customer PII or funds, or regulatory reporting breach. Triggers: automatic all-hands page to on-call, CTO + CEO paged within 10 minutes, dedicated bridge call within 5 minutes, status-page update within 15 minutes, regulator notification assessment within 1 hour, customer comms within 30 minutes.
- **SEV2 – High**: Payment processing degraded >30% throughput or >5% error rate, single-region failover failure, ledger read-only mode, or any condition likely to breach the 99.95% SLA within the current window. Triggers: primary + secondary on-call paged, incident commander assigned within 10 minutes, bridge call within 15 minutes, status-page update within 30 minutes, VP Engineering notified within 20 minutes.
- **SEV3 – Medium**: Non-critical service degradation affecting <30% of customers, single non-ledger microservice outage with working fallback, or elevated latency above SLA threshold but below halt. Triggers: primary on-call paged, IM notified within 30 minutes, status-page update within 1 hour if customer-visible, daily standup update.
- **SEV4 – Low**: Degraded internal tooling, minor UX bug with workaround, non-customer-facing alert. Triggers: next-business-day response, ticket created, no page unless on-call agrees.
- **SEV5 – Informational / Noise**: Cosmetic issues, planned-maintenance notifications, alert misfires logged for tuning. Triggers: no page, logged for weekly alert-quality review.
Define **escalation rules**: any SEV3 unresolved after 4 hours auto-escalates to SEV2; any SEV2 unresolved after 2 hours auto-escalates to SEV1. Severity can be **downgraded** only by the incident commander with IMC notification.
Publish the taxonomy as a one-page decision tree, a Slack slash-command (`/sev`), and an integration into the alerting tool so that every alert carries a suggested severity.
4. Define Incident Roles and Staffing Model (depends on: 3)
Codify four mandatory roles for every SEV1/SEV2 incident and optional roles for SEV3, then solve the 24×7 staffing problem across 28 teams.
**Roles**
- **Incident Commander (IC)**: owns the incident end-to-end, declares severity, assigns tasks, authorizes mitigations, decides when to escalate or stand down. Never writes code during the incident.
- **Communications Lead (CL)**: owns status-page updates, internal Slack channels, account-manager briefings, and regulator notifications. Separate from the IC so the IC can focus on mitigation.
- **Scribe / Timeline Keeper**: logs every decision, action, and timestamp in the incident channel and the incident-management tool. Produces the raw timeline for the postmortem.
- **Subject-Matter Responders (SMRs)**: 1–3 engineers from the owning team(s) who diagnose and fix. For the shared PostgreSQL ledger, a dedicated DBA responder is always required.
**24×7 Staffing via a Three-Tier Follow-the-Sun Model**
- **Tier 1 – Front-line on-call**: Primary + secondary responder per team, paged first. Covers the team's own services.
- **Tier 2 – Platform / SRE on-call**: A dedicated 6-person SRE rotation covering cross-cutting infrastructure: Kubernetes, the shared PostgreSQL ledger, networking, and the two AWS regions. This tier directly addresses the 'carrying a pager for other teams' concern by absorbing infrastructure incidents.
- **Tier 3 – IMC escalation**: Engineering managers and the IMO on-call for multi-team or SEV1 incidents. Provides the IC and CL when no team-level IC is available.
**Follow-the-Sun**: If any engineering hub exists in a second timezone, use it for overnight Tier 1 coverage. If not, partner with a managed on-call service for overnight first-response triage (severity declaration + paging the correct team), reducing 3 a.m. pages for NY-based engineers.
**IC and CL pools**: Nominate at least 2 ICs and 1 CL per team (56 ICs, 28 CLs minimum). ICs are trained and certified before they rotate. For SEV1 incidents, the IC must be a certified IC from the IMC pool, not just 'whoever is around'.
**Ledger-specific rule**: Because the PostgreSQL ledger is shared, a **Ledger Duty Officer** from the SRE Tier 2 is always on the bridge for any incident touching ledger services, regardless of which team owns the failing microservice.
5. Design On-Call Rotations, Compensation, and Alert-Quality Rules (depends on: 4)
Make on-call sustainable, fairly compensated, and free of alert noise so engineers stop resisting participation.
**Rotation Design**
- 7-day rotations, one primary + one secondary per team per week. No engineer is on-call more than one week in four.
- Minimum 48-hour rest between rotations. No on-call during approved PTO.
- All 28 teams participate. Teams without current on-call get a 90-day ramp with a shadow rotation before going live.
- Tier 2 SRE rotation: 6 engineers, one week on / five weeks off, with a dedicated backup.
**Compensation Package**
- **Base on-call stipend**: $500 per week of primary on-call, $250 for secondary, paid regardless of whether pages fire.
- **Page pay**: $75 per acknowledged page outside business hours; $150 if the page leads to active incident work.
- **Time-off-in-lieu (TOIL)**: Any engineer who works >4 hours overnight (00:00–06:00 local) gets a full TOIL day. >8 hours in a single incident gets 1.5 TOIL days.
- **SEV1 bonus**: $300 flat bonus for every engineer who actively works a SEV1 incident, paid within the next pay cycle.
- **Annual on-call cap**: No engineer exceeds 13 weeks of on-call per year. Exceeding the cap triggers a mandatory team-staffing review.
- Budget estimate: ~$420K/year for stipends and page pay across 28 teams; present this to the CFO as a fraction of the $1.3M annual SLA credit cost.
**Alert-Quality Rules (the '85% noise' problem)**
- Every alert must carry: owning team, suggested severity, runbook link, and a 30-day noise score.
- **Alert budget**: each team gets a maximum of 100 actionable alerts per month. Exceeding the budget triggers a mandatory alert-tuning session with the IMO.
- **Noise threshold**: any alert that fires >10 times in 7 days with no human action is auto-flagged for suppression or tuning within 14 days.
- **Alert review cadence**: weekly 30-minute alert-quality review per team; monthly cross-team alert review in the IMC.
- **Sunset rule**: alerts with no runbook are demoted to SEV5 after 30 days and suppressed after 60 days unless a runbook is written.
- Target: reduce monthly alert volume from 3,400 to <600 actionable alerts within 6 months.
6. Build Detection, Escalation, and Communication Paths (depends on: 3, 4)
Eliminate the 22-minute median detection gap and the 40% customer-first-detection rate with layered monitoring and a single escalation spine.
**Detection Layers**
- **Synthetic transactions**: run a payment end-to-end through the full stack (API → service → ledger → confirmation) every 60 seconds from both AWS regions. Alert if latency >2× baseline or any step fails. This catches what per-service metrics miss.
- **Customer-traffic anomaly detection**: monitor API error rates, payment success rates, and latency percentiles per customer cohort. Alert on >2σ deviation.
- **SLO-based alerting**: define SLIs for the 99.95% SLA (availability, latency p99, ledger consistency). Alert when error budget burn rate exceeds threshold, before the SLA actually breaches.
- **Infrastructure health**: Kubernetes node/pod health, PostgreSQL replication lag, disk I/O, and cross-region latency.
- **Support-ticket spike detection**: if >5 customers open tickets about the same symptom within 10 minutes, auto-create a SEV3 candidate.
**Escalation Path**
- Alert fires → PagerDuty routes to Tier 1 primary → 5-min no-ack → Tier 1 secondary → 10-min no-ack → Tier 2 SRE → 15-min no-ack → IMC on-call manager → 20-min no-ack → VP Engineering auto-page.
- Any SEV1 declaration auto-pages the CTO, opens a dedicated Slack channel + Zoom bridge, and notifies the CL.
- **No incident goes unowned for >15 minutes.** If no IC is assigned by minute 15, the IMC on-call manager assumes IC role by default.
**Internal Communications**
- Dedicated Slack channels: `#inc-sev1`, `#inc-sev2`, `#inc-sev3` (auto-created per incident), plus `#inc-updates` for broadcast.
- IC posts a structured update every 15 minutes (SEV1), 30 minutes (SEV2), 1 hour (SEV3) into the incident channel.
- CL posts a summary to `#inc-updates` and notifies relevant engineering managers.
**Customer Communications**
- **Status page**: auto-updated via API. SEV1: first update within 15 minutes, then every 30 minutes until resolved. SEV2: first update within 30 minutes, then every hour. SEV3: within 1 hour if customer-visible.
- **Account managers**: CL briefs AMs via a dedicated Slack channel within 30 minutes (SEV1) or 1 hour (SEV2). AMs contact their top-20 revenue accounts directly.
- **Customer email/SMS**: for SEV1 and SEV2, automated notification to all 2,100 customers via the status-page subscription system within 30 minutes.
- **Regulator notification**: Legal/Compliance assesses within 1 hour whether NY DFS, FinCEN, or card-network notification is required. If yes, file within the regulatory deadline (typically 72 hours for NY DFS cybersecurity events). Log the decision and filing in the incident record.
**Post-resolution**: CL publishes a 'resolved' update within 30 minutes of mitigation. For SEV1/SEV2, a preliminary customer-facing RCA summary is published within 5 business days.
7. Standardize Postmortems with Tracking and Accountability (depends on: 3, 6)
Fix the '11 of 64 action items closed' problem with a mandatory, uniform, blameless postmortem process backed by engineering-manager accountability.
**When Mandatory**
- All SEV1 and SEV2 incidents: postmortem required within 5 business days.
- SEV3 incidents: postmortem required if >50 customers affected, if the incident lasted >4 hours, or if it was customer-detected.
- SEV4/SEV5: optional, but any recurring SEV4 (≥3 times in 30 days) triggers a mandatory review.
**Format (single template, enforced by tooling)**
- Incident summary (severity, duration, customers affected, revenue impact, SLA credit exposure).
- Timeline (auto-generated from scribe notes + alert timestamps).
- Detection analysis: how was it detected, why did it take X minutes, could it have been faster.
- Root cause analysis using 5-Whys or fault-tree, not blame.
- Contributing factors (process, tooling, staffing, knowledge gaps).
- Impact quantification: customers affected, transactions failed, SLA credits triggered.
- Action items: each with a **named owner**, **due date**, **priority**, and **ticket in Jira**.
- Lessons learned and what went well.
**Blameless Review Meeting**
- Held within 5 business days, facilitated by the IMO or a trained facilitator (never the IC of that incident).
- All responders, the CL, relevant engineering managers, and an IMC representative attend.
- Ground rules: focus on system and process failures, not individual mistakes. The facilitator enforces this.
- Meeting recorded; notes published to the engineering-wide wiki within 48 hours.
**Action-Item Tracking and Accountability**
- Every action item is created as a Jira ticket with a due date and a named owner.
- **Engineering managers are accountable**: action-item completion is a standing agenda item in the bi-weekly IMC meeting. Any action item >7 days overdue is escalated to the VP Engineering.
- **Completion gate**: no team may close its postmortem until 100% of its action items have Jira tickets. Postmortem is 'closed' only when all tickets are resolved.
- **Quarterly audit**: the IMO audits action-item completion rates and reports to the IMC and the executive sponsor. Target: >90% completion within 30 days of the postmortem.
- Link postmortem quality and action-item completion to team health scores and engineering-manager performance reviews.
8. Define KPIs, Dashboards, and Governance Reviews (depends on: 7)
Create a measurable feedback loop so leadership can see whether the program is working and where to intervene.
**Primary KPIs (tracked weekly, reported monthly)**
- **MTTD** (Median Time to Detect): target <5 minutes (from 22).
- **MTTA** (Median Time to Acknowledge): target <5 minutes.
- **MTTM** (Median Time to Mitigate): target <60 minutes for SEV1 (from 3 h 10 min), <4 hours for SEV2.
- **Customer-first detection rate**: target <10% (from 40%).
- **SLA compliance**: maintain 99.95%; track monthly SLA credit payouts, target <$100K/quarter (from ~$325K/quarter).
- **Alert signal-to-noise ratio**: target >80% actionable (from ~15%).
- **Alert volume**: target <600/month (from 3,400).
- **Postmortem completion rate**: 100% for SEV1/SEV2 within 5 business days.
- **Action-item completion rate**: >90% within 30 days (from ~17%).
- **On-call health**: pages per engineer per week (target <5), TOIL usage, on-call satisfaction survey score.
- **Unowned incident duration**: target 0 incidents with >15 minutes without an IC.
**Dashboards**
- Real-time operational dashboard (Grafana): current incidents, active alerts, on-call roster, SLA error-budget burn.
- Weekly leadership dashboard (auto-generated): KPI trends, open action items, alert-noise report, on-call load distribution.
- Quarterly IMC scorecard per team.
**Review Cadence**
- **Weekly**: IMO publishes KPI snapshot to `#inc-updates`.
- **Bi-weekly IMC**: review open incidents, overdue action items, alert-quality exceptions, and on-call load.
- **Monthly executive review**: CTO presents KPI trends, SLA credit cost, and risk register to the CEO.
- **Quarterly incident-management review**: deep-dive into trends, training gaps, tooling needs, and process improvements. Output fed into the next quarter's roadmap.
9. Consolidate Tooling and Build the Incident Management Platform (depends on: 2, 3)
Replace six alerting tools and ad-hoc status-page updates with a single, integrated incident management stack.
**Target Tool Architecture**
- **Single alerting and on-call platform** (e.g., PagerDuty or Opsgenie): ingest all alerts, apply severity routing, manage on-call schedules, handle escalations, and send pages. Retire the other five tools within 6 months.
- **Observability consolidation**: standardize on one APM/metrics stack (e.g., Datadog or Grafana Cloud) for all 180 Kubernetes services across both AWS regions. Ensure the shared PostgreSQL ledger has dedicated dashboards.
- **Status page**: a dedicated, branded status page (e.g., Statuspage.io or Instatus) with API integration for auto-updates. Subscribe all 2,100 customers.
- **Incident coordination tool**: integrate incident-management workflows into Slack (auto-create channels, invite responders, post templates) and a dedicated incident record system (e.g., Jira Service Management, incident.io, or Rootly) for timelines, postmortems, and action-item tracking.
- **Runbook repository**: a central wiki (Confluence or Notion) with mandatory runbooks for every alert. No alert goes live without a linked runbook.
**Implementation Tasks**
- Migrate all 28 teams' alert rules into the single platform in three waves (highest-volume teams first).
- Build the severity-based routing rules and escalation policies per S3 and S6.
- Automate status-page updates triggered by severity declaration.
- Build the synthetic-transaction monitor and SLO-based alerting per S6.
- Integrate Jira for automatic action-item ticket creation from postmortems.
- Decommission legacy tools only after all teams have completed training on the new stack.
- Budget: allocate $150K–$250K/year for licensing, plus engineering time for migration.
10. Prepare for the SOC 2 Type II Audit (depends on: 7, 8, 9)
Ensure the incident management process produces the evidence the auditor will need, well before the audit window opens in eight months.
**SOC 2 Requirements to Address (CC7.3, CC7.4, CC7.5)**
- Documented incident response procedures (the severity taxonomy, role definitions, communication templates).
- Evidence of incident detection, response, and recovery for every SEV1/SEV2 incident during the audit period.
- Postmortem records with action-item tracking.
- On-call schedules, training records, and escalation evidence.
- Status-page update logs and customer notification records.
- Regulator notification logs (if any).
**Preparation Tasks**
- The IMO maintains a **SOC 2 evidence folder**: every incident record, postmortem, action-item ticket, status-page update, and training completion certificate is stored and indexed.
- Conduct a **mock SOC 2 audit** at month 5: an internal or external auditor reviews the incident management process end-to-end and identifies gaps.
- Remediate mock-audit findings before month 7.
- Ensure the incident management tool retains all records for at least 12 months (the SOC 2 Type II observation window).
- Document the **chain of custody** for incident records: who accessed, modified, or closed each record.
- Prepare a **narrative document** describing the incident management process, roles, and controls for the auditor.
- Coordinate with the compliance lead to align incident management evidence with the broader SOC 2 scope (access controls, change management, etc.).
11. Design and Deliver Training, Runbooks, and Change Management (depends on: 4, 5, 9)
Equip all 260 engineers, 28 team leads, account managers, and support staff with the knowledge and muscle memory to execute the new process.
**Training Tracks**
- **All 260 engineers** (2-hour session): severity taxonomy, how to acknowledge a page, how to join an incident bridge, how to hand off to an IC, and how to write a postmortem contribution. Delivered in team-level sessions over 4 weeks.
- **IC pool (56+ engineers)** (8-hour certification): incident command techniques, severity declaration, escalation decision-making, bridge facilitation, and blameless postmortem facilitation. Includes two tabletop exercises. Certification valid for 12 months, renewed annually.
- **CL pool (28+ staff)** (4-hour session): status-page writing, customer communication templates, regulator notification triggers, and AM briefing protocol.
- **Account managers and support staff** (1-hour session): how to read the status page, how to escalate a customer report into an incident, and what information to collect.
- **SRE Tier 2** (16-hour onboarding): Kubernetes and PostgreSQL ledger deep-dive, cross-region failover runbooks, and escalation authority.
**Runbooks**
- Every alert must have a runbook before it is routed to on-call. The IMO provides a runbook template and audits compliance weekly.
- Priority runbooks to write first: shared PostgreSQL ledger failover, Kubernetes cluster degradation, payment-processing pipeline failure, cross-region failover, and ledger data-integrity check.
- Runbooks are peer-reviewed and version-controlled.
**Change Management for Adoption**
- Address the 'carrying a pager for other teams' concern directly: publish an FAQ explaining the three-tier model, the SRE Tier 2 absorbing cross-team infrastructure, the compensation package, and the TOIL policy.
- Run **office hours** weekly for the first 8 weeks where any engineer can ask questions or raise concerns.
- Identify **team champions**: one engineer per team who volunteers as an early adopter and peer mentor.
- Publish a **weekly 'incident management newsletter'** during rollout: what changed, what improved, KPI trends, and success stories.
- Make on-call participation a documented expectation in job descriptions and performance reviews for production-facing roles.
12. Execute Phased Rollout, Tabletop Exercises, and Continuous Improvement (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11)
Introduce the process in three waves so teams are not overwhelmed, then validate with exercises and iterate continuously.
**Phase 1 – Weeks 1–6: Foundation**
- Publish the severity taxonomy, role definitions, and communication protocols (S3, S4, S6).
- Launch the single alerting platform for the 12 teams already on-call; begin migration for the other 16.
- Activate the SRE Tier 2 rotation for the shared PostgreSQL ledger and cross-cutting infrastructure.
- Deploy the status page and test the API integration.
- Begin IC and CL training (first cohort of 20 ICs, 10 CLs).
- Publish the on-call compensation package; HR integrates stipends into payroll.
- Write the top 10 priority runbooks.
**Phase 2 – Weeks 7–14: Expansion**
- All 28 teams live on the single alerting platform; legacy tools in read-only mode.
- All 28 teams on the on-call rotation schedule (the 16 new teams in shadow mode for the first 4 weeks).
- Second and third IC/CL training cohorts completed.
- First **tabletop exercise**: simulate a SEV1 ledger corruption scenario with all roles, test the escalation path, status-page updates, and AM briefings. Debrief and fix gaps.
- Postmortem template and Jira integration live; all new incidents use the standard process.
- Alert-tuning sprint: each team reduces its alert volume by 50%.
**Phase 3 – Weeks 15–24: Optimization**
- All teams fully live; legacy alerting tools decommissioned.
- Second **tabletop exercise**: simulate a SEV1 cross-region failure with regulator notification.
- First quarterly IMC review with full KPI dashboard.
- Mock SOC 2 audit (month 5) and remediation.
- Retrospective on the rollout: survey all 260 engineers for feedback, adjust compensation or rotation rules if needed.
- Establish the **continuous improvement cadence**: quarterly process review, annual severity-taxonomy review, and annual on-call compensation benchmarking.
**Ongoing Governance**
- The IMC owns the process document and approves changes.
- The IMO tracks all KPIs and reports to the CTO monthly.
- Any process change requires IMC approval and a 2-week notice period before enforcement.
- Annual external benchmarking against peer B2B payments platforms.
13. Establish Ongoing Governance, Annual Review, and Audit Readiness Cycle (depends on: 12)
Embed incident management as a permanent organizational capability, not a one-time project.
- **Annual process review**: the IMC reviews the severity taxonomy, role definitions, on-call structure, and compensation against industry benchmarks and internal KPIs. Update as needed.
- **Bi-annual tabletop exercises**: one SEV1 infrastructure scenario, one SEV1 data-breach/regulator scenario. Rotate the IC and CL assignments so everyone gets practice.
- **Quarterly alert-quality audit**: the IMO reviews alert volumes, noise rates, and runbook coverage across all 28 teams.
- **On-call health survey**: quarterly anonymous survey measuring burnout, fairness, and compensation satisfaction. Results reviewed by the IMC.
- **SOC 2 readiness cycle**: begin evidence collection immediately after each audit ends. The IMO maintains a rolling evidence folder. Mock audit at month 5 of every 12-month audit cycle.
- **Postmortem maturity tracking**: track the action-item completion rate monthly. If it drops below 80%, the VP Engineering intervenes.
- **Incident management maturity model**: adopt a 5-level maturity model (ad-hoc → defined → managed → optimized → predictive). Assess annually. Target: Level 3 within 12 months, Level 4 within 24 months.
- **Budget review**: annually review on-call compensation, tooling costs, and training budget against the reduction in SLA credits and incident frequency.
Previous Proposal 4 (ID: 31fbbc6c-8c88-4abb-9c6d-b67706d9da75, Agent: grok4.6_initial_4, LLM: xai/grok-4.6):
Estimated Complexity: high
Success Metrics: - Median time to detect customer-impacting incidents ≤ 5 minutes within 6 months of go-live.
- Share of SEV-1/SEV-2 incidents first detected by customers ≤ 5% (from 40%).
- Median time to mitigate SEV-1/SEV-2 ≤ 45 minutes (from 3 h 10 min).
- Named Incident Commander assigned within 5 minutes for ≥ 95% of SEV-1/SEV-2.
- First status-page update within policy time for ≥ 95% of SEV-1/SEV-2.
- SLA credits down ≥ 80% versus the trailing $1.3M within 12 months.
- Paging volume ≤ 500 per month and noise ≤ 15% (from 3,400 and 85%).
- 100% of production services have a named owning team and a paging policy.
- Postmortems filed within 5 business days for 100% of SEV-1/SEV-2; action-item close rate ≥ 80% within 30 days.
- 24×7 IC and critical-path coverage with zero unfilled shifts per quarter.
- Paid on-call live for every rotation before that rotation pages humans.
- SOC 2 Type II incident-response controls evidenced for ≥ 5 months before the auditor's report.
- On-call pulse: ≥ 70% of engineers agree rotations are fair and limited to their services.
Steps (30):
1. Secure executive mandate and budget
Get a written CEO/CTO mandate that incident command is a company process, not a team hobby.
The mandate must state that **paid on-call** is required for production ownership. "No pager for other teams' code" is solved by named ownership, not by refusing coverage.
- Approve budget for tooling, stipends, training, and a dedicated program lead for six months.
- Name an executive sponsor (CTO or VP Engineering) who will chair the weekly incident review.
- Tie the clock to SOC 2 Type II: the process must be live in about 10 weeks so ~6 months of evidence remain.
- Commit that the CEO will hear about outages from this process, not from customers.
2. Form the working group and decision rights (depends on: 1)
Stand up a small group that can decide. Do not form a 28-team committee.
**Core seats:** SRE/platform lead, payments/ledger engineering manager, support lead, legal/compliance, HR, one rotating team EM, and a program manager.
- Meet twice a week for 10 weeks, then weekly.
- RACI: the group proposes; the sponsor decides in 48 hours; teams implement.
- Publish one Slack channel and one source-of-truth doc on day one.
- Time-box design to four weeks. Ship v1 rather than wait for consensus.
3. Inventory services, owners, and on-call gaps (depends on: 1)
Build a living catalog of all ~180 services: owning team, criticality, current on-call, alert sources, and runbook link.
Walk the last 31 customer-impacting incidents and the two events where **nobody was in charge**. Record who detected, who led, time to mitigate, and which alerts fired.
- Tag each service as critical-path, customer-visible, or internal.
- List the 16 teams with no on-call and every orphan service with no owner.
- Map all six alerting tools and the 3,400 monthly alerts onto services.
- Flag the shared PostgreSQL ledger and two-region failover as named special cases.
4. Map regulatory and contractual notification duties (depends on: 1)
Legal and compliance list every duty an incident can trigger. Do not invent clocks that violate a contract.
Cover SOC 2 CC7, customer MSA/SLA credit terms, money-transmitter rules, NYDFS 23 NYCRR 500 if applicable, PCI if in scope, and breach clocks.
- Extract **notification timings** from the largest customer contracts (status page, named AM, written notice).
- Define when legal, regulators, insurers, or the board must be told.
- Feed these clocks into severity triggers and the communications playbook.
5. Approve paid on-call and incident pay (depends on: 1, 4)
Unpaid on-call is why 16 teams refuse the pager and why nights are uncovered. Fix the money before asking for coverage.
HR, legal, and finance design a New York–compliant package: weekly stipend for primary and secondary, extra stipend for company IC and comms, and **after-hours incident pay** or comp time.
- Treat exempt vs non-exempt staff explicitly under NY wage-hour rules.
- Put stipend in the next pay cycle after policy publish, not "later."
- Cap consecutive night weeks. Fund hiring if a team cannot rotate fairly (minimum six people for 24x7 primary plus secondary).
- Publish the package before any new rotation starts. This is the main answer to pager pushback.
6. Ratify four severity levels and their triggers (depends on: 3, 4)
Adopt a business-impact scale. Engineers do not invent severity in the moment.
**SEV-1:** material payments failure; ledger down or inconsistent; security or customer-data incident; both regions impaired; or many customers already in SLA-credit territory.
**SEV-2:** degraded payments or a contracted feature down for multiple customers; SLA at risk.
**SEV-3:** narrow or single-customer impact with a workaround; no fleet-wide SLA risk.
**SEV-4:** no customer impact; ticket only.
- SEV-1 pages IC, comms, scribe, owning SMEs, and an exec; war room in 5 minutes; status page in 10; AM outreach in 20.
- SEV-2 pages IC and owning SMEs; comms may be the IC; status page in 15 minutes; updates every 30 minutes.
- SEV-3 pages the owning team only; customer notice only if that customer is affected.
- Anyone may declare. Only the IC may downgrade. When unsure, start high.
7. Define incident roles and their authority (depends on: 6)
Four roles. Separate coordination from debugging so "who is in charge" cannot stall for an hour again.
- **Incident Commander:** owns severity, the room, the clock, and the next action. Does not write code. May page anyone, freeze deploys, and invoke failover. Staffed from a company-wide trained pool, not from the failing team.
- **Communications Lead:** status page, customers, AMs, execs, regulators. Speaks only from IC-approved facts.
- **Scribe:** timeline in the incident tool. Required for SEV-1 and SEV-2.
- **SME responders:** the owning team's on-call. They mitigate. They do not run the room.
Publish a one-page authority card. The IC stays in charge even if a VP joins.
8. Design 24x7 coverage without 28 night rotations (depends on: 3, 5, 7)
Do not put 28 teams on 24x7. That is what engineers are rejecting.
Use **three layers** so people page for their own code, plus a trained commander.
- Layer A — company IC and SEV-1 comms: 24x7; about 24 trained people; week-long primary and secondary.
- Layer B — critical-path team on-call (ledger, payments processing, auth, API edge, platform/Kubernetes, data stores): 24x7 primary plus secondary.
- Layer C — all other teams: business-hours on-call; after hours the IC pages the team EM, who has a written escalation list.
Platform on-call is the safety net for unknown-owner pages, never the permanent owner. Every service must have a named team within 60 days or be scheduled to shut off.
9. Set rotation, handoff, and load rules (depends on: 8)
Write mechanical rules so rotations are fair and load is visible.
Primary week, then secondary week, then at least two weeks off. No one holds primary on two rotations at once.
- Handoff is a 30-minute overlap covering open incidents, silenced alerts, and upcoming changes.
- Page-load SLO: p50 ≤ 4 pages per 12-hour night shift; p95 ≤ 10. A breach opens an alert-quality action.
- Require a shadow week before a first IC shift or a first critical-path rotation.
- Swaps live in the paging tool. Managers own coverage gaps, not the last person on the roster.
10. Write detection and escalation paths (depends on: 6, 9)
Customers currently detect 40% of incidents and median time to detect is 22 minutes. That is the first failure mode.
Detection path: synthetic full-payment probes in both regions, SLO burn-rate alerts, support-to-incident intake, and one customer callback path that can create a SEV.
- A page must be acked in **5 minutes** or it auto-escalates to secondary, then IC, then the EM, then the VP.
- Support may declare SEV-2 or higher without engineering permission.
- If ownership is unclear for 10 minutes, the IC keeps the incident and assigns a temporary owner. Never wait.
- An exec bridge auto-opens for every SEV-1 at T+15 minutes.
11. Write internal, customer, and regulator communications (depends on: 4, 6, 7)
Stop "whoever is around" from writing the status page. Comms follow the clock, not convenience.
Timings from declaration:
- Internal war room: immediate. Exec summary for SEV-1/2 at 15 minutes, then every 30 minutes.
- **Public status page:** SEV-1 in 10 minutes, SEV-2 in 15. Updates at least every 30 minutes until resolve. Templates only. No speculation.
- Account managers get an affected-customer list and a script at T+20 minutes for SEV-1/2.
- Resolve notice and credit assessment within one business day.
- Comms pages legal on SEV-1 security, ledger integrity, or any outage that will breach contractual notice. Legal owns outbound regulatory letters; the IC owns facts.
12. Standardize blameless postmortems and action tracking (depends on: 6)
A written postmortem is mandatory for every SEV-1 and SEV-2 within 5 business days. SEV-3 if the IC or EM requests it.
Use one template: timeline, customer impact (volume, duration, credits), detection gap, what went well, what did not, process-focused five whys, and numbered actions with owner and due date.
- Review is **blameless** and scheduled. The IC attends. The exec sponsor reads every SEV-1.
- Actions live in one tracker, not in the doc. No action without an owner and a date. Default due date 14 days; 30 days max unless architecture work with a milestone.
- Close rate is a published metric. The old 11-of-64 pattern is a process failure.
13. Set alert quality rules that make paging acceptable (depends on: 3, 6)
3,400 alerts a month and 85% noise is why on-call feels like punishment. Pages are a product with a quality bar.
A page (not a ticket) must map to a customer-facing SLO or a hard dependency of one. It must have an owner team, a runbook link, and a default severity. It must be actionable at 3 a.m. by the person who is paged.
- Ban parallel paging from six tools. One paging policy: symptom-based; burn-rate preferred over raw thresholds.
- Every team gets a monthly noise budget. Exceeding it is a sprint task, not heroics.
- A human may silence a flapping alert only with a linked ticket.
14. Publish Incident Management Policy v1 (depends on: 5, 6, 7, 8, 9, 10, 11, 12, 13)
Collapse the design into a short policy people will open during an outage.
Ten pages or fewer, plus one-page cards for severity, roles, and comms timings. Host it where the incident tool can link it.
- Include the compensation summary and the rule: you are **not on-call for other teams' services**.
- Version it. v1 is mandatory from the pilot start date.
- Legal, HR, and the exec sponsor sign. Announce in all-hands, not only in Slack.
15. Implement a single incident command tool (depends on: 7, 10, 11)
Put one tool in the path that creates the room, pages roles from severity, records the timeline, and prompts status-page updates.
Requirements: Slack (or equivalent) incident bot, severity in one click, role assignment, stakeholder groups, and timeline export for postmortems and auditors.
- Integrate with the pager so IC, comms, and SME pages are automatic.
- Retain artifacts at least one year for SOC 2.
- Ad-hoc Zoom/Slack threads are no longer the system of record.
16. Consolidate six alerting tools onto one pager (depends on: 9, 13)
Pick one paging product. Connect existing monitors to it. Migrate **pages** first, tickets second.
- Inventory every page-producing rule. Delete or downgrade the noisy majority in S23.
- Route by service label → owning team schedule → the escalation policy from S10.
- IC and comms schedules live in the same product.
- Set a hard date after which pages outside the chosen tool are not valid on-call obligations.
17. Operationalize the status page and AM path (depends on: 11, 15)
Put the status page behind the Comms Lead role. Use templates for investigating, identified, mitigating, and resolved.
Subscribe AMs and customers whose contracts require it. Generate the affected-customer list from the incident (tenant, region, payment method).
- Dry-run a SEV-2 update before the pilot goes live.
- Record every public update in the incident timeline for the audit.
- Define partial vs full-outage wording so impact cannot be understated.
18. Stand up one postmortem repo and action board (depends on: 12)
Create the template, the filing location, and one Jira/Linear board with states: open, in progress, blocked, done, and won't-do (with exec reason).
Wire the incident tool so a SEV-1/2 automatically opens a draft postmortem and action tickets.
- Each week the program manager reports actions past due to the exec sponsor.
- If it is not in the board, it does not exist.
19. Assign owners and write critical-path runbooks (depends on: 3, 9)
Close the ownership gaps that cause "pager for other people's code."
Every production service gets a team in the catalog. Unowned services get an owner in 30 days or a decommission date.
- Write runbooks for ledger Postgres, regional failover, payments API, auth, and the Kubernetes control plane: symptoms, dashboards, mitigate vs escalate, customer impact.
- Link runbooks from alerts. If there is no runbook, the alert cannot page at night unless the EM accepts the gap in writing.
20. Detect payments failures before customers do (depends on: 3, 13)
Build or finish end-to-end synthetics: create payment, ledger write, webhook, both regions, both critical payment methods.
Alert on SLO burn, not on a single 500. Page SEV-2 or SEV-1 from these probes. This is the fastest lever on the 22-minute MTTD and the 40% customer-detected rate.
- Add ledger lag, replication, disk, and failover-readiness as first-class pages to the ledger team.
- Every postmortem asks first: **why did a customer see this first?**
21. Train the first cadre of ICs, comms leads, and scribes (depends on: 7, 14, 15)
Train about 24 ICs and 12 comms leads before the pilot. Classroom plus a recorded shadow of a simulated SEV-1.
Curriculum: severity, authority card, tool, comms timings, when to call legal, how to run a room of 20, how to hand off at 2 a.m.
- Certification: pass a tabletop. No certificate, no rotation.
- Recertify yearly and after any SEV-1 where process failed.
- Managers of ICs protect calendar time. This is part of the job.
22. Pilot on the payments critical path for six weeks (depends on: 5, 14, 15, 16, 17, 18, 19, 21)
Go live with policy, tool, paid rotations, and IC coverage for ledger, payments, API, platform, and support intake.
Keep old paths as backup for one week, then cut over. Real incidents use the new process only.
- Staff the program lead in every SEV-2+ as coach, not as secret IC.
- Collect friction daily. Fix tooling and wording in 48 hours.
- Expansion gate: named IC in under 5 minutes, first status update on time, no unpaid pages, postmortem filed.
23. Cut alert noise with a forced burn-down (depends on: 13, 16)
Give every team a numbered list of their noisiest alerts. Move noise from 85% to **under 15%**, and monthly pages from 3,400 toward 500.
Each sprint, critical-path teams must delete, debounce, or convert to ticket a fixed quota. Platform provides burn-rate and grouping libraries.
- Publish a weekly noise leaderboard. Shame systems, not people.
- After eight weeks, any page without a runbook or with >30% false pages in 14 days is auto-downgraded until fixed.
24. Roll remaining teams onto the model by risk (depends on: 22)
After the pilot gate, add teams in waves of four to five every two weeks. Highest customer-impact first.
Layer C teams get business-hours schedules and the EM night list. Do not surprise anyone with a pager.
- Each wave: ownership confirmed, alerts routed, runbooks for paging alerts, paid rotation in HR, one tabletop.
- Finish all 28 teams at least five months before the SOC 2 report date so the observation window covers the company.
- Orphan services still unowned at wave end are escalated to the sponsor for shutdown or reassignment.
25. Run tabletops and multi-region game days (depends on: 21, 22)
Schedule a monthly tabletop: SEV-1 ledger, SEV-1 region loss, SEV-2 degraded payments, customer-detected incident, and a "who is in charge" chaos drill.
Quarterly game day: fail a region or a ledger replica in staging or a controlled production drill.
- Include support, AMs, legal, and an exec. Process fails if only engineers show up.
- Capture actions on the same board as real postmortems.
- Use results as SOC 2 evidence that IR is tested.
26. Resolve ownership fights and pager culture (depends on: 5, 8, 22)
Treat pushback as design input, not defiance. Repeat the contract in office hours: you carry a pager for **your** services; IC is coordination; nights are paid; noise is a defect.
- EMs who cannot staff a fair rotation get headcount or have services reassigned. Do not run two-person 24x7.
- Publicly close the two historical "nobody in charge" incidents with what would be different now.
- Pulse-survey on-call at 60 and 120 days. If load or fairness is red, stop expansion until fixed.
27. Launch metrics, weekly review, and error budgets (depends on: 15, 22)
The CEO email exists because there was no operating rhythm. Stand up a dashboard and use it.
Track incident count by SEV, MTTD, MTTA, MTTM, customer-detected %, time-to-IC, time-to-status, pages per shift, noise %, SLA credits, postmortem SLA, and action close rate.
- **Weekly 30-minute review** chaired by the exec sponsor. Last week's incidents, open actions, alert SLO breaches. Mandatory for EMs of teams that had SEV-2+.
- Error budgets on critical SLOs: burn too fast and the team stops features and pays down reliability.
- Monthly report to the CEO so customer-first news does not recur.
28. Collect SOC 2 Type II evidence as a side effect (depends on: 4, 14, 22, 27)
Do not create a parallel "audit process." The real process is the evidence.
From pilot start, retain policy versions, on-call schedules, incident records, status-page history, postmortems, action tickets, training attendance, tabletop notes, and access reviews of the incident and pager tools.
- Map each artifact to SOC 2 CC7.2–CC7.5. Compliance owns the matrix; engineering owns completeness.
- Internal audit or a consultant runs readiness reviews at week 12 and week 20, not the week before the auditor.
29. Inspect and adapt the process after 90 days (depends on: 22, 27)
After 90 days live, revise v1 using data, not opinions.
Look at MTTD/MTTM, uncovered shifts, IC burnout, missed status updates, action close rate, and survey results.
- Change severity definitions if teams inflate or deflate them.
- Adjust Layer B vs C membership from real pages.
- Cut steps nobody follows. Add only what the last 90 days proved missing.
- Freeze v2 as the SOC 2 described process unless a defect is found.
30. Rehearse the audit and freeze the observation window (depends on: 24, 28, 29)
Six to eight weeks before the auditor, run a mock interview with an IC, a random engineer, support, and compliance.
Show one SEV-1 and one SEV-2 end-to-end: detect, page, command, comms, postmortem, closed actions.
- Fix evidence gaps immediately.
- Freeze process wording for the rest of the observation window; log exceptions.
- Brief the CEO and customer success with the metrics so the "outages we hear about from clients" story is retired.
Previous Proposal 5 (ID: 53260d57-1d36-433f-83b4-7205e06869b2, Agent: deepseek-v4-pro_initial_5, LLM: deepseek/deepseek-v4-pro):
Estimated Complexity: high
Success Metrics: - Customer-detected incidents decrease from 40% to less than 15% within six months.
- Median time to detect (MTTD) is under 5 minutes for SEV1 and SEV2 incidents.
- Median time to mitigate (MTTM) is under 60 minutes for SEV1 and under 2 hours for SEV2.
- Alert noise decreases from 85% to below 10% within three months.
- 100% of SEV1 and SEV2 incidents have a completed blameless postmortem within 5 business days.
- 100% of postmortem action items are tracked with an owner and due date; 90% are completed on time.
- 24x7 on-call coverage achieved across all 28 teams with no unpaid on-call.
- 99% of pages are acknowledged within 5 minutes.
- SLA credits paid reduce by at least 50% over the next 12 months.
- SOC 2 readiness: all incident response controls are documented, tested, and evidence is produced by month 7.
Steps (14):
1. Baseline current incident response and align stakeholders
Collect data from the last 12 months of incidents, all six alert tools, on-call practices, and team interviews. Identify gaps against the target incident management process and secure executive sponsorship.
- Gather incident timeline, detection source, mitigation time, customer impact, SLA credit, and postmortem status for all 31 incidents.
- Survey 28 teams on on-call burden, alert quality, and operational pain.
- Map current tools, escalation paths, and communication workflows.
- Create baseline metrics and a stakeholder map with executive sponsor and audit owner.
2. Define severity levels and response triggers (depends on: 1)
Define a four-level severity scale with objective business-impact criteria so any engineer can classify an incident consistently.
- SEV1: widespread transaction processing outage, data breach, security incident, or severe SLA breach; triggers full incident command, executive notification, and a 5-minute status page update.
- SEV2: major feature outage, significant degradation without workaround, or customer financial risk; triggers incident commander, full communications role, and status page updates.
- SEV3: partial impairment with workaround or limited customer impact; triggers on-call response, internal communication, and optional status page update.
- SEV4: minor or internal issue, no customer impact; handled during business hours through ticketing.
- Include an escalation matrix showing who can declare, downgrade, and invoke regulatory or legal involvement.
3. Define incident roles and decision authority (depends on: 1, 2)
Define incident roles, responsibilities, and decision authority using RACI to remove ambiguity about who is in charge.
- Incident Commander: owns the incident, declares severity, and coordinates resolution.
- Communications Lead: owns internal and external messaging, status page updates, and account manager notifications.
- Scribe: maintains timeline, incident log, and postmortem notes.
- Subject-matter responders: diagnose and fix the incident; may come from multiple teams.
- Executive sponsor: optional for SEV1; customer liaison: handles account managers.
- Define decision rights for severity declaration, escalation, rollback, customer communications, and incident closure.
4. Design 24x7 staffing model across 28 teams (depends on: 3)
Design 24x7 coverage across 28 teams without overloading engineers. Use service-based on-call plus a central incident command pool.
- Each service or domain team assigns primary and secondary on-call for its own services.
- Create central incident commander, communications, and scribe rotations staffed from a trained incident response guild across all teams; use follow-the-sun between the two AWS regions and time zones.
- Define escalation layers: service on-call to team lead or manager to service owner to executive.
- Define handoff times, shadow shifts, and load balancing; target at most one week of on-call per engineer per month.
- Bridge the current 12-team paid on-call to 28-team paid coverage; no team remains uncovered.
5. Define on-call rotations, compensation, and alert quality rules (depends on: 4)
Define sustainable rotations, pay, and rules that eliminate noisy pages.
- Rotations: weekly or biweekly, at least one primary and one secondary, with 12-hour shifts where possible or 24-hour for low-volume services.
- Compensation: monthly on-call stipend for all on-call engineers, additional incident response bonus for after-hours work, and time off in lieu; align with market rates.
- Alert quality rules: every page must be actionable, have a runbook link, specify a service owner, include severity, and be based on SLO burn or known failure signals; no dashboard-only alerts.
- Noise budget: reject or downgrade non-actionable alerts; all pages must go to on-call only after suppression and deduplication.
- Weekly alert review removes the top noisy alerts.
6. Design detection, escalation, and alert routing (depends on: 2, 4, 5)
Define how incidents are detected, routed, and escalated so nothing waits on a human to notice.
- Consolidate the six alert tools into one alerting and paging platform with routing by service, severity, and tags.
- Detection sources: infrastructure metrics, application synthetic transactions, log-based anomalies, business transaction SLI monitoring, and customer-reported issues through support or account managers.
- Routing: alert is paged to service on-call within 30 seconds; primary must acknowledge within 5 minutes; if no ack, page secondary then on-call manager.
- Escalation timeouts: unresolved SEV1 escalates to service owner at 15 minutes and to leadership at 30 minutes; any engineer can escalate to the incident commander.
- Define customer-reported incident intake and classification in the same tool.
7. Define internal and external communication protocols (depends on: 2, 3)
Define communication channels, templates, and timing for internal, customer, and regulator audiences.
- Internal: dedicated incident Slack channel, internal status page mirror, and war room bridge for SEV1; incident commander and communications lead own these channels.
- Status page: SEV1 post within 5 minutes, updates every 30 minutes or on material change, resolution within 60 minutes of mitigation; SEV2 post within 15 minutes, updates hourly; SEV3 optional.
- Account managers: SEV1 and SEV2 notify account managers within 15 minutes with an approved customer-facing description and expected impact.
- Regulators: legal or compliance determines notification for data breaches, security incidents, funds availability issues, or regulatory reportable events; criteria and timing follow legal and regulatory requirements; communications lead coordinates.
- Use pre-approved message templates and an approval chain; no ad-hoc wording.
8. Define postmortem policy and action tracking (depends on: 3)
Define mandatory blameless postmortems and action tracking.
- Mandatory for all SEV1 and SEV2 incidents, and any SEV3 that breaches SLA or is customer-detected.
- Format: impact, timeline, root causes, contributing factors, detection and response gaps, what worked well, and action items.
- Blameless: focus on system and process causes, not individual blame; use trained facilitators.
- Ownership: each action has an owner, due date, and tracking ID in a single backlog.
- Review postmortems at the weekly incident review; track action closure; expect 100% completion.
- Complete postmortems within 5 business days for SEV1 and SEV2 incidents.
9. Define metrics, dashboards, and review cadence (depends on: 2, 3, 8)
Define metrics and review cadence to measure process health.
- Metrics: MTTD, MTTM, customer detected percentage, alert noise percentage, on-call response time, on-call load, SLA credits paid, and postmortem action completion.
- Dashboards: real-time operational dashboard for on-call engineers and management.
- Weekly incident review: review all SEV1 and SEV2 incidents, action items, and noisy alerts.
- Monthly trends with leadership; quarterly review against SLOs and audit controls.
- Success thresholds: MTTD under 5 minutes, MTTM under 60 minutes for SEV1, customer detected under 15%, and alert noise under 10%.
10. Configure incident tooling and integrations (depends on: 5, 6, 7, 8, 9)
Implement and integrate the tools that automate the defined process.
- Aggregate alerts from the existing six tools into PagerDuty, Opsgenie, or a similar platform.
- Configure on-call schedules, escalation policies, and paging targeted at service owners.
- Integrate status page API for automated or one-click updates.
- Add Slack commands to declare incidents, start war rooms, assign roles, and post status updates.
- Integrate runbook and service catalog access; create postmortem templates in Jira or Notion with action item tracking.
- Ensure audit trails and role assignments are logged for SOC 2.
11. Pilot with 2-3 volunteer teams and iterate (depends on: 10)
Run a controlled pilot before full rollout to validate and refine the process.
- Select 2-3 volunteer teams with representative services and on-call patterns.
- Run the new severity, roles, on-call, alerting, and communication process for 2 weeks.
- Track metrics and gather feedback from on-call engineers, incident commanders, and communications leads.
- Iterate severity thresholds, alert rules, templates, and runbooks based on findings.
- Exit criteria: no SEV1 without a declared incident commander, alert noise below target, and positive on-call survey results.
12. Train and certify all 28 teams (depends on: 11)
Prepare all 28 teams so the process is not a paper exercise.
- Create role-specific training: all engineers, incident commanders, communications leads, scribes, subject-matter responders, and managers.
- Deliver interactive workshops and online modules.
- Run at least two game-day drills per quarter simulating SEV1 and SEV2 incidents, testing escalation, communication, and postmortem.
- Certify incident commanders and communications leads; only certified staff can hold those roles in rotations.
- Include scenario-specific training for account managers and legal or regulatory handling.
13. Full rollout and enforce process (depends on: 12)
Move every team onto the new process and retire old habits.
- Migrate all 28 teams into consolidated on-call schedules and alert routing.
- Retire old alerting tools or methods that bypass the process.
- Activate status page templates and postmortem action tracking globally.
- Monitor the first 30 days with a daily operational review and quick fixes.
- Communicate rollout status to leadership and all engineers; provide office hours and support.
14. SOC 2 readiness and continuous improvement (depends on: 13)
Map the process to SOC 2 and keep improving after go-live.
- Map each process component, including severity, roles, detection, communications, postmortems, and metrics, to SOC 2 control requirements.
- Produce runbooks, evidence of training, incident logs, postmortems, and action item tracking.
- Conduct an internal readiness assessment at month 6 and remediate gaps before the audit at month 8.
- Maintain weekly and monthly metrics reviews to drive continuous improvement.
- Refresh on-call compensation and alert quality rules based on operational data.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Program charter, mandate, and funding", "description": "Convert the CEO email into a named program with one owner, a budget, and a deadline earlier than the audit.\n\nThe charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, 28 teams will opt out.\n\n- Appoint a Program Lead (Head of Reliability) with direct CTO and CEO sponsorship.\n- Form a small steering group: CTO, VP Eng, Head of CS/Support, CISO, Legal, Finance, HR. Not a 28-team committee.\n- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.\n- Timeline: operating floor in 7 days; design weeks 1–6; pilot weeks 7–12; rollout weeks 13–24; mock audit week 28; SOC 2 at month 8.\n- Fund tooling ($150–250k/yr), on-call pay (~$600k–1M/yr), 2–3 program FTEs, and reserved engineering capacity. Anchor the ask against $1.3M in credits plus unmeasured incident cost.\n- Freeze baselines now: 31 incidents, MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, $1.3M credits, 3,400 alerts/month at 85% noise, 11 of 64 actions closed.", "dependencies": []}, {"step_id": "S2", "title": "Immediate 7-day operating floor", "description": "Do not wait for tooling, compensation, or the audit. Put a minimum process in place this week so the next outage has a named commander.\n\nStart retaining artifacts on day one. This week becomes the first audit evidence.\n\n- Publish a one-page interim severity guide and a single declaration path through Slack, phone, and pager.\n- Staff interim primary and backup Incident Commander 24x7 from existing on-call veterans and engineering managers. Compensate this duty retroactively.\n- Require a named IC within 10 minutes for every suspected major incident. If nobody claims it, the duty manager is IC.\n- Use one channel, one bridge, one timeline doc, and one naming convention for every major incident.\n- Direct Support to escalate credible customer reports immediately. They do not wait for engineering confirmation.\n- Triage all 53 open historical actions. Close, re-plan, or formally risk-accept. Ledger integrity, payment duplication, regional resilience, security, and detection come first.\n- Hold a daily 15-minute ops review until the permanent process is live.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Forensic baseline of incidents and alerts", "description": "Rebuild the facts before locking design. This is the before picture for the CEO and the auditor.\n\nEvery later design choice should trace to this evidence.\n\n- Re-code each of the 31 incidents: trigger, service, detection source, timestamps, who led, credits paid, root-cause family.\n- Quantify the 40% customer-first detections and name the missing signal in each case.\n- Reconstruct the two nobody-in-charge incidents minute by minute. Use them as the burning-platform story.\n- Audit the six alerting tools: volume per tool and team, top 50 noisy rules, rules with no owner or runbook.\n- Freeze the baseline numbers. Do not let them drift during design.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Listening tour and the fairness contract", "description": "Engineer pushback against carrying a pager for other teams' code is the main delivery risk. Treat it as a design constraint, not an attitude problem.\n\nThe answer that will actually stick is **the deal**: you are paid; you are paged only for services you own; a trained commander runs the room; your postmortem actions get real sprint capacity.\n\n- Interview all 28 teams plus Support, CS, and Sales in two weeks.\n- Separate the objections: unpaid work, nights, unfamiliar code, bad runbooks, fear of blame. Each needs a different fix.\n- Collect informal practices from the 12 teams already on-call. They are the pilot candidates and the veteran pool.\n- Recruit 10–15 credible engineers as a design working group so the process is co-authored.\n- Baseline sentiment on alert trust, on-call willingness, and burnout. Re-measure at 6 and 12 months.", "dependencies": ["S1"]}, {"step_id": "S5", "title": "Service ownership catalog and criticality tiers", "description": "You cannot page the right person across 180 services until each one has a named owner. This is the foundation of fairness, routing, and audit evidence.\n\nBuild a machine-readable catalog as the single source of truth.\n\n- Inventory all 180 services, Kubernetes clusters, AWS accounts, the shared PostgreSQL cluster, queues, partner banks, and customer-facing endpoints.\n- Assign one owning team, named engineering manager, Slack channel, escalation policy, and dependency list per service.\n- Tier 0: money movement, ledger, auth, shared Postgres, regional control plane. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.\n- Map Tier 0/1 services to customer capabilities: initiation, authorization, settlement, reporting, onboarding.\n- Assign coordinated but separate responders for the ledger application and the PostgreSQL platform.\n- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.\n- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.", "dependencies": ["S3"]}, {"step_id": "S6", "title": "Severity scale and declaration rules", "description": "Replace in-the-moment debate with a lookup. Four payments-specific levels, with worked examples drawn from the 31 real incidents.\n\nAnyone may declare. Only the Incident Commander may downgrade, with recorded rationale. When unsure, start high.\n\n- **SEV1**: money movement stopped or incorrect; ledger integrity in doubt; material security or data exposure; both regions impaired; or more than 10% of customers impacted. Pages IC, comms, scribe, SMEs, and exec liaison. Bridge in 5 minutes. Status page in 15 minutes. Mandatory postmortem and regulator assessment.\n- SEV2: severe degradation; settlement window at risk; a strategic customer fully down; SLA breach likely. IC and SMEs paged. Status page in 30 minutes. Mandatory postmortem.\n- SEV3: partial impact with a workaround; no credit exposure. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, or a repeat.\n- SEV4: internal or minor. Ticket only. No page.\n- Auto-escalate: any SEV3 open more than 2 hours, or any incident touching the shared ledger, becomes SEV2. Unknown impact after 15 minutes is raised, not sat on.\n- Nobody is punished for over-declaring. Publish that rule in writing and repeat it.", "dependencies": ["S3", "S5"]}, {"step_id": "S7", "title": "Roles, authority, and ledger dual-control", "description": "Solve nobody-in-charge-for-an-hour by making command explicit, transferable, and logged.\n\nThe Incident Commander owns the incident, not the fix, and **never types** in a production terminal.\n\n- Incident Commander: declares severity, pulls anyone, freezes deploys, invokes failover, authorises spend. One IC at a time. Assumed within 5 minutes and announced in channel.\n- Communications Lead: single voice to customers, status page, account managers, and the exec summary.\n- Scribe: timestamped timeline, decisions, and open questions. Feeds the postmortem and the audit trail. Required for SEV1 and SEV2.\n- Subject-matter responders: diagnose and mitigate only services they own, with access and runbooks.\n- Executive Liaison (SEV1): shields the IC from exec questions; owns regulator and board escalation.\n- The IC may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access rules.\n- Handover is verbal and written, with time and new owner recorded. Roles may combine below SEV2; never at SEV1. The IC stays in charge if a VP joins.", "dependencies": ["S6"]}, {"step_id": "S8", "title": "Three-layer 24x7 coverage model", "description": "Do not put 28 teams on night rotation. That is what engineers are rejecting.\n\nStaff command centrally. Page engineers only for code their team owns.\n\n- Layer A, company command: Duty IC plus secondary, Duty Comms, and a scribe pool. Target 25–35 certified ICs and 15–20 comms leads. Roughly one week per person every 6–8 months. Five-minute ack SLA, then secondary, then on-call director.\n- Layer B, Tier 0/1 domains: group 28 teams into 8–12 product and platform domains (ledger app, Postgres platform, payments orchestration, auth, API edge, Kubernetes/platform, settlement). Primary plus secondary 24x7. Minimum six trained people. Target one week in six, never worse than one in four.\n- Layer C, Tier 2/3: business-hours on-call. After hours the IC pages the EM, who holds a written escalation list.\n- No engineer joins another team's responder pool without training, access, runbooks, shadow shifts, and both teams' acceptance.\n- Platform on-call is the safety net for unknown-owner pages, never the permanent owner.\n- Evaluate follow-the-sun coverage as a 12-month option, not a year-1 dependency.", "dependencies": ["S5", "S7"]}, {"step_id": "S9", "title": "Paid on-call, NY labor compliance, and fatigue rules", "description": "Unpaid on-call is a retention risk and a New York legal exposure. Pay must be live in payroll before any mandatory night rotation starts.\n\nTreat interrupted nights and recovery time as compensable work.\n\n- Weekly stipend by layer, higher for 24x7 Tier 0/1 and Duty IC, lower for business-hours, with holiday and weekend premiums.\n- After-hours call-out pay or TOIL. A paid recovery day after prolonged overnight work, a SEV1, or a qualifying SEV2. Managers cover the next day.\n- HR, Legal, and Finance publish dollar amounts, FLSA exempt/non-exempt treatment, NY wage-hour rules, tax treatment, and payroll timing within 14 days of charter.\n- Load rules: no primary on two rotations; no consecutive primary weeks; no on-call the week after a SEV1 you commanded.\n- A person may declare temporarily unfit after overnight work with no performance penalty.\n- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.\n- If a rotation cannot staff six people, merge domains or hire. Do not run two-person 24x7.\n- Model annual cost against $1.3M in credits and get it as a CFO/board line item.", "dependencies": ["S4", "S8"]}, {"step_id": "S10", "title": "Alert quality contract and page budget", "description": "3,400 alerts a month at 85% noise is why detection takes 22 minutes. Make quality a condition of paging a human.\n\nA page is a product with a quality bar, not a dump of host metrics.\n\n- Every paging alert must have a named owner, customer or SLO impact, runbook, tested threshold, severity mapping, dashboard, expected action, and dedup key. Fail any of these and it becomes a ticket or is deleted.\n- Page on customer symptoms: SLO burn, error budget, settlement-queue age. Cause-based CPU and memory alerts become dashboards or tickets.\n- Set a **page budget** of at most two out-of-hours pages per person per week. A breach triggers a mandatory tuning sprint and blocks new paging alerts for that team.\n- Auto-quarantine alerts that fire more than five times a month with no action, or that have more than 70% no-action acks. Never silently disable without compensating detection and a recorded decision.\n- Run new alerts in shadow for seven days unless an emergency exception is approved.\n- Monthly per-team kill, tune, or keep review. Target under 500 pages a month and actionability above 75% within six months.", "dependencies": ["S3", "S5"]}, {"step_id": "S11", "title": "Customer-journey detection and ledger assurance", "description": "Stop customers being the monitoring system. Detect on money-path outcomes, not host metrics.\n\nEvery postmortem will ask why a customer saw it first.\n\n- Define SLOs per Tier 0/1 capability: initiation success, auth latency, settlement timeliness, API availability, reporting freshness. Set an internal target stricter than 99.95%.\n- Run synthetic full-lifecycle payments from outside the platform, in both regions, every 60 seconds. Include a small real-value canary where legally feasible.\n- Add ledger assurance: continuous double-entry reconciliation, replication lag, failover-readiness, unexpected balances, duplicate identifiers, settlement-window countdown.\n- Add per-customer anomaly detection for the top 100 accounts.\n- Auto-create a triage incident within five minutes when support tickets or AM reports match impact keywords.", "dependencies": ["S5", "S10"]}, {"step_id": "S12", "title": "Single incident and paging platform", "description": "Collapse six alerting tools into one queue, one timeline, and one audit record.\n\nThe platform must still work if one AWS region is down. Verify SMS and phone fallback, plus an offline runbook.\n\n- Select one paging product, one incident layer, and one hosted status page.\n- One Slack command declares the incident, creates the channel and bridge, pages the Duty IC, sets severity, and starts the clock.\n- Ingest the six existing tools first. Deduplicate and route. Retire a legacy path only after named owners and two weeks of verified operation.\n- Auto-capture timestamps, acknowledgements, role assignments, severity changes, comms, mitigated time, and resolved time. Export for SOC 2.\n- Integrate the service catalog, Jira, Salesforce or CS tooling for affected-customer lists, and Zoom or Slack Huddle.\n- Test paging, escalation, status publication, and conference access every week.\n- After a team's wave, old paging paths are disabled, not left as a fallback.", "dependencies": ["S6", "S8", "S10"]}, {"step_id": "S13", "title": "Detection-to-command escalation path", "description": "Write the single path from something looks wrong to someone is in charge. Target: a commander in under five minutes, any hour.\n\nIf nobody claims IC in five minutes, the platform assigns the Duty IC. The assignee can hand over, not decline.\n\n- Converge every entry point on the same declare command: alert, engineer, support, account manager, partner bank, SEV hotline.\n- SEV1: page the owning primary immediately; secondary at 5 minutes unacked; domain manager and Duty IC at 10; exec liaison at 15.\n- SEV2: primary ack in 10 minutes; IC assigned in 15.\n- Human acknowledgement is required. Delivery to a device does not count.\n- The IC can page any team's on-call, with a 10-minute ack obligation. This reciprocity makes single-team ownership viable.\n- Unowned alerts go to Layer A command, then the missing owner record is a control defect.\n- Pre-authorise regional failover, ledger read-only mode, and partner-bank notice so the IC does not wait for an executive. Dual-control still applies to ledger writes.", "dependencies": ["S7", "S8", "S12"]}, {"step_id": "S14", "title": "Live execution and major-incident playbooks", "description": "Limit customer and financial harm before proving root cause. One procedure from the first minute to handback.\n\nA service cannot page at night until it meets the readiness bar.\n\n- Open channel, bridge, record, and timeline immediately for SEV1 and SEV2. The IC states severity, known impact, hypothesis, objective, roles, and next update time.\n- Freeze unrelated production changes during SEV1. Record exceptions the IC approves.\n- Prefer reversible mitigation: rollback, feature flag, traffic isolation, rate limit, partner reroute.\n- Write playbooks first for Postgres ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank outage, settlement-window breach, suspected security compromise, and suspected duplicate payments.\n- Guard against split-brain, replay, and duplication during regional or database recovery. Reconcile and process backlog before calling the incident resolved.\n- Mitigation means customer impact ended. Resolution means stable, backlog processed, and ledger reconciled.\n- Formal IC handover after 4 hours. Reopen if impact recurs during the stability window.\n- Readiness bar per Tier 0/1 service: diagram, dependencies, dashboard, runbook, rollback, kill switch, escalation contacts, RPO/RTO. Untested runbooks are marked stale.", "dependencies": ["S7", "S13"]}, {"step_id": "S15", "title": "Internal, customer, and regulatory communications", "description": "Replace whoever is around with a timed, owned, pre-approved process. The Communications Lead is the single author.\n\nState impact and the next update time. Never speculate on cause.\n\n- Internal: one working channel, one bridge, one read-only broadcast for execs, Support, and Sales. SEV1 update every 30 minutes even if unchanged; SEV2 every 60 minutes.\n- Executives ask questions only of the Executive Liaison. Publish this as a signed exec behaviour rule.\n- Status page: SEV1 within 15 minutes, SEV2 within 30 minutes; then 30/60-minute updates; resolve notice within 30 minutes of mitigation. Templates pre-approved with Legal.\n- Top 100 accounts: named AM contact within 30 minutes of SEV1, with a briefing pack from Comms. Long tail gets the status page plus email or webhook.\n- Customer-facing summary within 5 business days for SEV1.\n- Mandatory regulatory checkpoint on every SEV1 and every security SEV2, within 2 hours, recorded even when not reportable. Cover NYDFS Part 500 72-hour clock, state breach laws, GLBA/FTC, PCI if in scope, sponsor-bank and card-network windows, FinCEN/OFAC if relevant.\n- Legal owns outbound regulatory letters. The IC owns facts. Security or Legal may limit public detail during an active threat, with the reason recorded.\n- Encode bespoke customer-contract notice SLAs into account tiering.", "dependencies": ["S6", "S7", "S12"]}, {"step_id": "S16", "title": "SLA credit and financial-impact workflow", "description": "Link incidents to money so severity, credits, and investment stay consistent. Finance should not learn about outages from invoices.\n\nMake credit calculation an output of the incident record, not a negotiation.\n\n- Agree availability measurement per contract and component with Legal and Finance.\n- Auto-compute affected minutes per customer from the incident record and capability telemetry. Produce a proposed credit schedule within 5 business days of resolution.\n- Posture: proactive credits for top-tier accounts; claims-based for the rest. Document the approval chain.\n- Track credits per incident and root-cause family. Quarterly report which reliability investments would have prevented which credits.\n- Target: cut credits from $1.3M to under $400k in 12 months. Use that delta as the ongoing business case.", "dependencies": ["S6", "S15"]}, {"step_id": "S17", "title": "Blameless postmortems and action tracking", "description": "Eleven of 64 actions closed is the clearest process failure. Make learning mandatory and actions as binding as customer commitments.\n\nClosing a ticket without evidence of effectiveness does not close the action.\n\n- Mandatory for every SEV1 and SEV2, every customer-first detection, every incident over 2 hours, every repeat of a known cause, and every ledger near-miss.\n- Draft within 3 business days, review within 5, published internally within 10. The IC owns the draft. The owning EM is accountable.\n- One template: timeline, customer and financial impact, why detection was late, why mitigation took that long, contributing factors, what went well, actions.\n- Blameless in writing. No individual named as a cause. HR and management commit that postmortems are never used in performance reviews.\n- Every action gets a named person, priority, due date, Jira ticket, and verification method. P0 (prevents SEV1 recurrence) due in 30 days and committed into the next sprint before roadmap work. P1 in 60 days. P2 in 90 days.\n- Teams reserve 15–20% of sprint capacity for reliability and incident actions.\n- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.\n- Weekly Incident Review Board with engineering directors. Target 90% of P0/P1 actions closed on time within two quarters.", "dependencies": ["S6", "S12"]}, {"step_id": "S18", "title": "Training, certification, and commander academy", "description": "Command is a skill, not a title. Nobody holds independent duty untrained. Training happens in paid working time.\n\nThe certification register is an audit artefact.\n\n- All employees, 1 hour: recognise impact, declare, find the channel and status page.\n- Responder, half day: severity, escalation, runbooks, comms hygiene. Mandatory before joining a rotation.\n- Scribe, 2 hours: timeline discipline. The entry point.\n- Incident Commander, two days plus shadowing: command presence, decisions under uncertainty, severity, handover, exec management. Two shadowed incidents and one simulated SEV1 before certification.\n- Communications Lead, one day: status writing, customer tiering, legal boundaries, regulator triggers.\n- Certification valid 12 months, renewed via simulation.\n- New joiners shadow two shifts and do not hold primary in their first 90 days.\n- Appoint a champion in each of the 28 teams.", "dependencies": ["S7", "S13", "S15", "S17"]}, {"step_id": "S19", "title": "Simulations and game days", "description": "Rehearse before the next real SEV1. Use a standing calendar, not a one-off exercise.\n\nDo not inject uncontrolled changes into the production ledger.\n\n- Monthly 60-minute tabletop per engineering group, using a real incident from the 31.\n- Quarterly full-scale game day: regional failover, ledger replica promotion, dependency failure. Whole role structure, timed.\n- Twice-yearly unannounced paging drill, including nights, to measure real acknowledgement times.\n- One security incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.\n- Also exercise status-page failure and loss of the primary chat or pager.\n- Use replicas, staging, or tightly governed tests for ledger scenarios.\n- Every exercise produces actions in the same tracker as real incidents.\n- Complete two cross-company exercises before the SOC 2 audit.", "dependencies": ["S14", "S15", "S18"]}, {"step_id": "S20", "title": "Critical-path pilot", "description": "Prove the process on the highest-risk surface with willing teams before asking 28 teams to adopt it.\n\nPublish a one-page result company-wide. That result is the adoption argument.\n\n- Six-week Wave 0: ledger app, Postgres platform, payments orchestration, Kubernetes/platform, API gateway, Support intake, plus two of the 12 teams already on-call.\n- Activate the full stack: severity scale, Duty IC, one pager, page budget, status page, mandatory postmortems, paid on-call.\n- Parallel-run old paths for one week, then cut over. The program lead coaches every SEV2+ and does not secretly take command.\n- Weekly retro. Expect 20–40 process defects and fix them in the standard before rollout.\n- Exit gate: MTTD under 10 minutes on pilot services; IC assigned within 5 minutes in 95% of incidents; page volume down 50%; all postmortems on time; no unpaid pages; sentiment not worse.", "dependencies": ["S9", "S11", "S12", "S18", "S19"]}, {"step_id": "S21", "title": "Metrics, reviews, and error budgets", "description": "Instrument the process itself. Do not reward hiding incidents or suppressing pages.\n\nEvery metric has a target and a named owner. Dashboards are public inside the company.\n\n- Response: MTTD, time to declare, time to IC, MTTA, MTTM, MTTR, incidents by severity, percent customer-first. Report median and p90, by tier, journey, and region.\n- Quality: pages per person per week, actionability, budget breaches, postmortem on-time rate, action closure and age, missed acknowledgements.\n- Business: credits, 99.95% per capability, error-budget burn, failed-payment volume, reconciliation breaks, repeat-incident rate.\n- People: rotation size, frequency, after-hours pages, recovery days, sentiment, attrition among on-call staff.\n- Cadence: weekly Incident Review Board; monthly Reliability Review; quarterly exec and board review; annual policy review.\n- Error budgets on Tier 0/1 SLOs. Burn too fast and the team pauses features to pay down reliability.\n- Reconcile dashboards monthly against a sample of incident records and customer cases so missing incidents cannot hide.", "dependencies": ["S12", "S17", "S20"]}, {"step_id": "S22", "title": "Wave rollout with readiness gates", "description": "Roll out in four waves by criticality, every three weeks. Gates keep the standard credible. A missed gate is rescheduled, not waived.\n\nFinish all 28 teams by week 24 so roughly three months of operating evidence remain before audit fieldwork.\n\n- Wave 1: remaining Tier 0. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.\n- Per-team kit: catalog complete, alerts migrated and inside budget, runbooks at the readiness bar, rotation of 6+ certified responders, one IC nominee, one drill passed, paid rotation in payroll.\n- Named coach for three weeks. The director signs the gate.\n- Freeze legacy paging per wave.\n- Publish a live adoption scoreboard.\n- Any production service that cannot provide sustainable ownership gets an executive-reviewed deadline, compensating control, and expiry date. No indefinite verbal exceptions.", "dependencies": ["S20", "S21"]}, {"step_id": "S23", "title": "Change management, incentives, and the deal", "description": "Run this from day one in parallel with design. Engineers will judge fairness. Executives will judge visible results.\n\nRepeat the deal until it is muscle memory: paid on-call, paged only for what you own, trained commander, real sprint capacity for actions.\n\n- Launch via CTO all-hands, per-team roadshows, a one-page laptop card, an internal wiki, and a Slack help channel with a 4-hour answer SLA.\n- Put incident-response contribution into promotion criteria. Award best postmortem and biggest noise cut quarterly. Thank people publicly after every SEV1.\n- Put adoption, alert hygiene, action closure, and on-call load fairness into every engineering manager's quarterly objectives.\n- Write an exception path for engineers who cannot do nights because of caring responsibilities or health, covered by stipended volunteers.\n- Prohibit retaliation for good-faith declaration or escalation.\n- Pulse-survey at 60 and 120 days. If fairness or load is red, pause expansion until fixed.\n- Publicly close the two historical nobody-in-charge incidents with what would be different now.", "dependencies": ["S1", "S4", "S9"]}, {"step_id": "S24", "title": "SOC 2 evidence, mock audit, and sustainability", "description": "Evidence is a by-product of doing the work, not a reconstruction the week before the auditor. Protect the process after SOC 2 is signed.\n\nDocumented exceptions beat a claim of perfection. Never edit historical records to look clean.\n\n- Map the process with Compliance and the auditor to CC7.2–CC7.4, CC2.2/CC2.3, CC5, and availability A1.2. Confirm interpretation early.\n- Maintain versioned, signed policies for incident response, severity, on-call, communications, and postmortems, reviewed annually.\n- Automate evidence: incident records, paging and ack logs, status history, postmortem library, action closure, training register, drill records, and reportability decisions including not reportable.\n- Internal dry-run at month 6 on a 15-incident evidence walkthrough. Mock audit at month 7. Fix with weeks to spare.\n- Assign permanent owners for policy, pager, status page, catalog, training, and metrics.\n- Year-two roadmap: follow-the-sun cell, self-healing for the top recurring causes, error-budget release gates, and blast-radius reduction on the shared ledger.\n- Quarterly board summary of severe incidents, credits, overdue P0s, and resilience investment so attention does not die after the audit.", "dependencies": ["S17", "S21", "S22"]}], "estimated_complexity": "high", "success_metrics": "- Median time to detect falls from 22 minutes to under 10 minutes by month 3 and under 5 minutes by month 9.\n- Customer-first detection falls from 40% to under 20% by month 3 and under 10% by month 6.\n- Median time to mitigate falls from 3 h 10 min to under 90 minutes by month 4 and under 60 minutes by month 12.\n- A named Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes after week 4.\n- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases by month 3, with required update cadence met in 95% of cases.\n- Monthly pages fall from 3,400 to under 500 within 6 months, with actionability above 75%; out-of-hours pages at or under 2 per person per week.\n- All six legacy alerting tools route through one paging platform by week 16; legacy paging paths disabled per wave.\n- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.\n- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by due date within two quarters; all 53 currently open historical actions triaged within 30 days of S2.\n- SLA credits fall from $1.3M to under $400k in the first 12 months; customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known unaddressed cause under 10%.\n- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar; every Layer B rotation has 6+ certified responders or a time-limited executive exception.\n- Paid on-call policy approved by HR, Legal, and Finance and in payroll before any mandatory night rotation starts.\n- All 28 teams onboarded by week 24 with 24x7 command cover from 25+ certified ICs and 15+ certified Comms Leads.\n- Internal dry-run at month 6 passes a 15-incident evidence walkthrough; month-7 mock audit finds no unowned high-risk control gap; SOC 2 Type II incident-response controls pass with zero exceptions.\n- On-call sentiment improves quarter over quarter; no increase in attrition among engineers on rotation; pulse at 60 and 120 days shows at least 70% agree rotations are fair and limited to services they own.\n- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked actions; two cross-company exercises completed before the audit."}Expanded from 14 thin steps to 26, closing nearly every R0 gap: interim duty officer, forensic baseline, listening tour, escalation ladder, regulatory playbook, credits, runbooks, exercises, pilot, waves and SOC 2 dry-run. The gains come almost entirely from copying P1's R0 structure and P2's SEV0, so it is now complete but the least differentiated plan.
- S1 stands up an interim 24x7 duty officer and single declaration path within 48 hours, compensated retroactively.
- S5 adopts a SEV0 crisis tier for corrupted or duplicated money movement, with payment-pause consideration.
- S13 requires human acknowledgement (delivery does not count) and routes unowned alerts to command as a logged control defect.
- S16 regulatory playbook with NYDFS 72-hour clock and a 24x7 contact matrix, entirely absent in R0.
- S17 SLA credit workflow and S20 readiness bar give the plan the financial and preparedness legs R0 lacked.
- Dropped R0's incoherent 'follow-the-sun between the two AWS regions' idea in favour of a paid NY command rotation (S7).
- S22 pilot depends on S12, S13, S15, S19, S21 but not on S9 (compensation), contradicting its own metric that paid on-call must be live before any mandatory rotation.
- No risk register or contingency step, and no standalone change-management/culture step — the pager-resistance answer is a bullet in S3 and a survey in S26.
- No anti-gaming or metric-reconciliation control, unlike P1 S26, P2 S18 and P4 S21.
- S8 caps rotation at one primary week in four, harsher than the one-in-six of P1, P2 and P4, weakening the fairness argument.
- S5's clock-driven auto-escalation (SEV2 unresolved >1h becomes SEV1) will inflate SEV1 volume when mitigation is already underway; no exception is defined.
- Success metrics are still anchored to 'within 6 months of full rollout' rather than absolute days, so the dates float with any schedule slip.
- Proposal 2 : SEV0 crisis level for integrity, duplication and dual-region loss, with consideration of pausing payments.
- Proposal 2 : Interim command and one declaration path stood up immediately and paid retroactively.
- Proposal 2 : Time-bound acknowledgement ladder, human acknowledgement required, unowned alerts to the command rotation as a control defect.
- Proposal 1 : Stipend figures, regulatory playbook, credit workflow, runbooks and readiness bar — with the exception of the risk register.
- Proposal 1 : Six-week critical-path pilot with numeric exit criteria and four gated waves with coaches.
- Proposal 2 : Mitigation-versus-resolution semantics requiring reconciliation before closure.
- Proposal 4 : A published Policy v1 and a quota-driven noise burn-down campaign with a leaderboard.
- Proposal 2 : Reconciling dashboards against customer cases to detect metric gaming or hidden incidents.
+ Executive mandate, program governance, and interim incident command+ Baseline data and alert estate analysis+ Design paid on-call compensation and fatigue safeguards+ Define acknowledgement and escalation paths+ Standardize customer and status page communications+ Standardize blameless postmortems+ Track postmortem actions with owner and due date+ Create runbooks and service readiness bar+ Train, certify, and simulate incident response+ Pilot on critical services and iterate+ Wave rollout to all teams and decommission legacy pathsProgram charter, executive mandate and fundingForensic baseline of the 31 incidents and the alert estateOn-call compensation, labour compliance and fairness policyDetection-to-escalation path and the five-minute command ruleCustomer communications and status page policyPostmortem standard and blameless review forumAction item ownership and tracking systemRunbooks, major-incident playbooks and the on-call readiness barTraining, certification and the commander academySimulation programme: game days, drills and wheel of misfortunePilot with wave 0 teamsWave rollout across all 28 teams with readiness gatesChange management, incentives and communications campaignProgram risk register and contingency planning
The plan produced
1. Executive mandate, program governance, and interim incident command new
Convert the CEO's outage complaint into a company-level improvement program with a named owner, budget, and deadlines. In the first week, establish an interim command process so no incident remains unowned while the permanent process is designed.
- Appoint Head of Reliability as program owner and CTO/CEO as executive sponsor.
- Form steering group with Engineering, SRE, Support, CS, Legal, Compliance, HR, Finance, Security.
- Approve non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Fund tooling, on-call compensation, training, and 3 dedicated program FTEs; anchor to $1.3M credits.
- Set timeline: design weeks 1-4, tooling and pilot weeks 5-10, rollout weeks 11-20, audit evidence collection from week 8, dry-run month 6.
- Stand up interim 24x7 duty officer and single declaration path within 48 hours, compensated retroactively.
2. Baseline data and alert estate analysis (after 1) new
Re-open the last 31 incidents and profile the current alert estate so every later design decision is evidence-based.
- Re-code each incident: detection source, timestamps, owner, severity, credits, root-cause family.
- Quantify customer-first detections and missing signals.
- Analyze two 'nobody in charge' incidents minute-by-minute.
- Inventory six alerting tools: volume, noise, owner, runbook coverage, top 50 noisy rules.
- Freeze baseline metrics: MTTD 22m, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed.
3. Stakeholder listening and resistance mapping (after 1, 2) from P1 step 3
Treat engineer pushback as design input. Interview all 28 teams plus Support, CS, and Sales to understand real objections and find early adopters.
- Test objection: unpaid work, night load, unfamiliar code, poor runbooks, or blame.
- Document current informal practices from 12 on-call teams.
- Recruit 10-15 credible engineers as co-design working group.
- Survey baseline sentiment on on-call and alerts.
- Publish promise: paid on-call, page only for owned services, trained commander, reserved reliability capacity.
4. Service ownership catalog and criticality tiering (after 2, 3) from P1 step 4
Build a machine-readable service catalog that assigns each of the 180 services a single owning team, escalation policy, and criticality tier.
- Define fields: owner team, manager, Slack channel, escalation policy, dependencies, dashboards, runbooks.
- Tier 0: ledger, money movement, auth, shared PostgreSQL; Tier 1 customer-facing; Tier 2 internal; Tier 3 non-critical.
- Map customer-visible capabilities and dependencies across regions.
- Identify orphan and shared services; force ownership decision within 30 days or schedule decommissioning.
- Publish coverage gaps by tier; Tier 0 gaps become executive escalations.
5. Define severity levels and declaration triggers (after 4)
Adopt a five-level severity model with objective payment-specific triggers, so declaration is a lookup rather than debate. Anyone may declare; only the Incident Commander may downgrade.
- SEV0: unauthorized, lost, duplicated, or corrupted money movement; ledger integrity loss; confirmed data breach; both regions failed. Page all roles, exec, legal, and consider payment pause.
- SEV1: widespread payment failure, no workaround, severe SLA breach, or single-region total loss. Full role activation and public status page.
- SEV2: significant degradation, multiple customers, or workaround available but SLA risk. IC and SMEs paged; status page if customer-visible.
- SEV3: limited impact with workaround; team-led, business-hours response.
- SEV4: internal/no customer impact; ticket only.
- Auto-escalation: unresolved SEV3 >2h becomes SEV2; unresolved SEV2 >1h becomes SEV1; ledger or security always at least SEV2.
- Provide decision tree and 12 worked examples from actual incidents.
6. Define incident roles, decision authority, and handover rules (after 5) from P1 step 6
Codify five roles with written responsibilities and explicit authority, so there is never confusion about who is in charge.
- Incident Commander: owns severity, priorities, cross-team pulls, mitigation decisions; does not code.
- Communications Lead: owns status page, internal updates, account manager briefs, regulator coordination.
- Scribe: maintains timeline, decisions, and evidence for postmortem/audit.
- Subject-Matter Responders: diagnose and remediate only their owned services.
- Executive Liaison (SEV0/SEV1): handles exec communications and external stakeholders.
- Rule: IC identified within 5 minutes and announced in channel; handover announced and logged; roles may combine below SEV2, never at SEV0/SEV1.
7. Design 24x7 incident command and comms staffing model (after 4, 6) from P3 step 4
Create a central, trained incident command rotation instead of making each of 28 teams field its own commander.
- Recruit 24-30 certified ICs and 12-16 Comms Leads from a cross-team volunteer pool with manager approval.
- Weekly rotations, primary and secondary; 5-minute acknowledgement SLA with auto-escalation.
- Scribe pool as entry-level rotation.
- Night coverage: paid command rotation in New York timezone initially; evaluate follow-the-sun coverage later.
- Eligibility: certification required; commanders can leave with 30 days notice.
- Ensure distinct persons for IC, CL, and primary SME on SEV0/SEV1.
8. Design team on-call rotations and cross-team escalation policies (after 4, 6)
Define tiered team on-call obligations so engineers are only paged for services they own, and cross-team pages go through the IC.
- Tier 0/1 services: 24x7 primary+secondary, at least 6 trained responders per rotation, one-week shifts.
- Tier 2/3: business-hours on-call; after-hours escalation list to engineering manager.
- Platform, infrastructure, database: 24x7 due shared ledger and Kubernetes.
- Routing: every page resolves via service catalog to owning team's escalation policy; cross-team pages only by IC.
- Guardrails: max one primary week in four, no consecutive weeks, no on-call in week after leading SEV0/1, protected recovery time after night work.
9. Design paid on-call compensation and fatigue safeguards (after 3, 8) from P2 step 6
Make on-call paid and legally compliant before any new rotation starts, and convert unpaid pager culture into a fair employment condition.
- Weekly stipends: primary 24x7 $800-1,200, secondary 30-50%, business-hours $300-500; holidays premium.
- Out-of-hours incident pay: $150 per night page plus hourly beyond one hour; time-off-in-lieu after overnight work.
- Additional command rotation stipend and SEV0/1 bonus for active responders.
- Verify FLSA/NY wage rules with HR, Legal, Finance; document exemption treatment.
- Base annual budget against current $1.3M credits.
- Track load and trigger staffing review if >2 after-hours pages per person per week sustained.
10. Enforce alert quality standards and noise budget (after 2, 4, 8) from P1 step 10
Replace 3,400 monthly alerts at 85% noise with a contractual paging standard that makes on-call sustainable.
- Every paging alert must have owning team, customer impact statement, runbook link, severity mapping, and tested threshold.
- Page only on customer-visible symptoms or SLO burn; cause-based alerts become tickets or dashboards.
- Set page budget: max 2 out-of-hours pages per person per week; breach triggers mandatory alert-tuning sprint.
- Auto-quarantine alerts with >5 firings/month without action or >70% no-action acknowledgements.
- Target <500 actionable pages/month and >75% actionability within 6 months.
- Weekly per-team alert review, monthly cross-team review.
11. Build detection uplift: synthetics, SLOs, and support intake (after 4, 10) from P1 step 11
Shift detection from host metrics to customer outcomes so the company stops hearing about outages from clients first.
- Define SLOs per Tier 0/1 capability: payment initiation, auth, settlement timeliness, API availability, ledger consistency.
- Deploy external synthetic transactions from both regions every 60 seconds, covering full payment flow and ledger write.
- Add ledger assurance checks: replication lag, double-entry balance, settlement window countdown.
- Top-100 customer anomaly detection to catch single-tenant outages.
- Auto-create triage incident from support tickets or account manager keywords within 5 minutes.
- Track customer-detected-first as a defect and require a postmortem action.
12. Consolidate alerting and incident tooling (after 5, 8, 10, 11) from P2 step 7
Collapse six alerting tools into one integrated paging and incident management platform to create a single system of record for people and audit.
- Select paging/on-call platform and incident management layer (e.g., PagerDuty + incident.io/FireHydrant).
- Implement one-command Slack declaration that auto-creates channel, bridge, pages roles, sets severity, starts timeline.
- Migrate all monitoring sources into the one tool; decommission legacy paging only after two weeks verified.
- Integrate service catalog, status page, Jira action tracking, Salesforce/CS customer lists, and conference bridge.
- Ensure out-of-band paging and offline fallback if a region or chat tool is down.
- Automate evidence capture for SOC2: timestamps, role assignments, severity changes, comms sent.
13. Define acknowledgement and escalation paths (after 6, 8, 12) from P2 step 9
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts at first signal.
- For SEV0/SEV1: page primary on-call; after 5 minutes unacked page secondary; at 10 page manager and duty IC; at 15 page executive officer.
- For SEV2: primary ack within 10 minutes, IC assigned within 15; escalate on miss.
- For SEV3: ack within 30 minutes or create work item.
- Automatic duty IC page for ledger/security, cross-team, or unresolved ownership.
- If impact unknown after 15 minutes, raise severity.
- Cross-team responders summoned by IC have 10-minute acknowledgement obligation.
- Route alerts with no owner to command rotation, then treat missing ownership as control defect.
- Human acknowledgement required; delivery confirmation is not sufficient.
14. Standardize internal communications (after 5, 6) from P2 step 11
Separate the working war room from executive and stakeholder updates, with fixed cadence and pre-approved templates.
- Auto-create incident channel and one read-only broadcast channel for execs, support, sales.
- SEV0: internal update every 15 minutes; SEV1 every 30; SEV2 every 60; SEV3 on state change.
- Template: impact, what we know, what we are doing, ETA/next update, current IC and CL.
- Executives questions through Executive Liaison only; IC is not interrupted.
- Support/CS receives affected-customer list and holding statement within 15 minutes for SEV0/SEV1 and 30 for SEV2.
- Handover protocol for incidents lasting >4 hours: formal IC handover and fatigue check.
15. Standardize customer and status page communications (after 5, 14) from P2 step 11
Replace ad-hoc status page updates with a timed, role-owned, template-driven process, including account manager outreach.
- Status page timing: SEV0 initial post within 10 minutes, SEV1 within 15, SEV2 within 30; updates every 15/30/60 minutes until resolved.
- Resolution notice within 30 minutes of mitigation; customer-facing summary in 5 business days for SEV0/SEV1.
- Pre-approve 12-15 templates with Legal and Comms.
- Tiered outreach: top 100 accounts get direct call/email from AM within 30 minutes of SEV0/SEV1; long tail via subscription.
- Use factual language: state impact and next update; never speculate cause or blame vendor.
- Comms Lead is sole author for customer language.
16. Regulatory, legal, and account manager notification playbook (after 5, 15) from P1 step 16
Build a notification decision tree and contact matrix so legal/regulatory obligations are assessed early and never forgotten.
- Map obligations: NYDFS Part 500 72-hour cybersecurity event notification, state breach laws, GLBA/FTC, PCI, sponsor bank/card network contractual windows, FinCEN/OFAC if relevant.
- Add regulatory assessment checkpoint for every SEV0 and security SEV1 within 2 hours, even if not reportable.
- Maintain 24x7 contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel with named backups.
- Pre-draft notification templates and test quarterly.
- Encode customer-specific notification SLAs from enterprise contracts into customer tiering.
- Account managers receive legal-approved script and affected-customer list.
17. SLA credit workflow and financial impact measurement (after 5, 15) from P1 step 17
Link every incident to money automatically so severity, credits, and prioritization stay consistent and Finance is never surprised.
- Define availability measurement per contract and capability with Legal/Finance.
- Auto-compute affected minutes per customer from incident record and telemetry; generate credit proposal within 5 business days.
- Decide proactive credits for top tier vs claims-based for others; document approval chain.
- Track credits by root-cause family and incident; feed quarterly reliability investment decisions.
- Target reducing annual credits from $1.3M to under $400K in first year.
18. Standardize blameless postmortems (after 5, 6) from P4 step 12
Make postmortems mandatory with fixed deadlines, a single format, and a blameless review forum, replacing 'various formats'.
- Mandatory: all SEV0/SEV1; SEV2 with customer impact, credits, >2h, repeat cause, or detected by customer; near-miss involving ledger.
- Draft within 3 business days, peer review within 5, publish within 10.
- One template: timeline, customer/financial impact, detection gap, response gap, contributing factors, what went well.
- Blameless rules: context-based, no individual blame, never used in performance reviews.
- Weekly Incident Review Board reviews all postmortems and challenges quality.
- Publish searchable postmortem library and quarterly recurring causes report.
19. Track postmortem actions with owner and due date (after 12, 18) new
Fix the 11/64 completion rate by giving every action the same status as customer commitments, with capacity and escalation.
- Every action gets named owner, priority, due date, and Jira ticket auto-created from postmortem.
- P0 actions prevent SEV0 recurrence, due 30 days; P1 60; P2 90.
- Reserve 15-20% team sprint capacity for incident actions.
- Escalation ladder: manager at 7 days overdue, director at 14, CTO dashboard at 30; P0 overdue blocks features.
- Monthly reporting of closure rate in engineering leadership.
- Target 90% closure of P0/P1 on time in two quarters.
20. Create runbooks and service readiness bar (after 4, 8, 11) new
Ensure every service is prepared for 3 a.m. response before it is allowed to page anyone.
- Readiness checklist for Tier 0/1: architecture diagram, dependencies, dashboards, rollback, feature flags, escalation contacts, data-loss statement.
- Write major-incident playbooks: shared PostgreSQL failure, cross-region failover, Kubernetes control plane loss, partner outage, settlement breach, security compromise.
- Prioritize ledger playbooks: documented failover, read-only mode, reconciliation, signed RPO/RTO.
- Runbooks must be tested twice yearly; stale runbooks marked in catalog.
- No paging alerts without readiness sign-off; gap reported to director.
21. Train, certify, and simulate incident response (after 6, 13, 14, 15, 18, 20) new
Build a certification path so 24x7 roles are staffed by people who have practiced, and validate the process through drills.
- Scribe course (2 hrs); Responder (half-day); Communications Lead (1 day); Incident Commander (2 days + shadowing/tabletop).
- Certify ICs and CLs; recertify annually.
- New responders shadow two shifts before primary; no primary in first 90 days.
- Monthly tabletop using real past incidents; quarterly game day including regional failover or ledger scenario.
- Twice-yearly unannounced paging drill; one regulator/legal exercise annually.
- Track drill metrics and action items.
22. Pilot on critical services and iterate (after 12, 13, 15, 19, 21)
Prove the process on the highest-risk surface before full rollout. Run a six-week pilot with tight measurement and public results.
- Select 5-6 teams: payments core, ledger/database, platform/Kubernetes, API gateway, plus two existing on-call teams.
- Activate full stack: severity, roles, command rotation, paid on-call, single tooling, alert budget, status page, postmortems.
- Weekly pilot retro; fix defects within 48 hours.
- Exit criteria: MTTD <10 min, IC assigned within 5 min in 95% incidents, pages down 50%, postmortems on time, positive sentiment.
- Publish one-page pilot result as adoption argument.
23. Wave rollout to all teams and decommission legacy paths (after 20, 22) new
Roll out to all 28 teams in four waves by criticality, with explicit readiness gates and legacy tool shutdown.
- Wave 1 remaining Tier 0; Wave 2 Tier 1; Wave 3 Tier 2; Wave 4 Tier 3/internal.
- Per-team onboarding: catalog complete, alerts pruned, runbooks ready, rotation staffed with 6+ certified responders, one IC candidate, one drill passed.
- Coach per wave 3 weeks.
- Gate signed by director; failures rescheduled, not waived.
- After onboarding, freeze legacy alerting paths; no fallback to old tools.
- Publish adoption scoreboard.
24. Establish metrics, dashboards, and review cadence (after 12, 19, 22)
Make process performance visible with a small set of metrics and fixed review meetings, so the program is owned by data.
- Response: MTTD, time to IC, MTTA, MTTM, MTTR, % customer-detected-first.
- Quality: page volume per person, alert actionability, postmortem on-time, action closure rate.
- Business: SLA credits, availability vs 99.95, repeat incidents.
- People: on-call load, page per engineer, sentiment, attrition.
- Cadence: weekly Incident Review Board, monthly Reliability Review, quarterly Executive/Board review.
- Dashboards self-serve with targets and named owners.
25. SOC 2 readiness and internal dry-run audit (after 19, 23, 24) from P1 step 26
Design evidence as a by-product and test it with an internal walkthrough before the external auditor arrives.
- Map process to SOC2 CC7.3/CC7.4, CC7.2, CC2.2/2.3, CC5, availability criteria.
- Publish versioned policy documents: Incident Response, Severity, On-Call, Communication, Postmortem.
- Automate evidence: incident records with timestamps, role assignments, paging logs, status history, postmortems, action board, training register, drill records.
- Ensure process operates at least 3 months before fieldwork.
- Run internal dry-run month 6; sample 15 incidents; fix gaps with 8 weeks to spare.
- Keep remediation log for process deviations.
26. Continuous improvement, culture, and sustainability (after 23, 25) from P1 step 29
Prevent the process from decaying after the audit by embedding review, feedback, and roadmap ownership.
- Quarterly process retrospective with IC pool and responders.
- Re-baseline metrics every six months; raise targets.
- Year-two roadmap: follow-the-sun coverage, self-healing top 3 causes, error budgets gating releases, blast-radius reduction for shared ledger.
- Annual policy review and certification renewal.
- Quarterly on-call sentiment survey with published actions.
- Board quarterly report on availability, credits, and incident trends.
- Median time to detect falls from 22 minutes to under 5 minutes within 6 months of full rollout.
- Customer-first detection falls from 40% to below 10% within 6 months and below 5% within 12.
- Median time to mitigate SEV0/SEV1 falls from 3h10 to under 60 minutes within 12 months; SEV2 under 2 hours.
- Incident Commander assigned and announced within 5 minutes for 95% of SEV0/SEV1 incidents; zero incidents with unclear command beyond 10 minutes.
- Status page updated within policy time for 95% of SEV0/SEV1/SEV2 (10/15/30 minutes).
- Monthly pages fall from 3,400 to under 500 with actionability above 75%.
- All six legacy alerting tools decommissioned by week 16.
- 100% of SEV0/SEV1 incidents have a blameless postmortem published within 10 business days.
- Postmortem action closure rises from 17% to 90% for P0/P1 actions on time.
- SLA credits fall from $1.3M to under $400K in first 12 months.
- Customer-impacting incidents decline to under 15 per year; repeat root causes under 10%.
- 100% of 180 services have a named owning team and criticality tier.
- All 28 teams onboarded by week 24; every Tier 0/1 team has 24x7 primary+secondary coverage with 6+ certified responders.
- At least 30 certified Incident Commanders and 20 certified Communications Leads active.
- Paid on-call policy is approved and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among on-call engineers.
- Quarterly game days and twice-yearly unannounced paging drills executed on schedule, each with tracked action items.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 223f0a8a-068a-4034-ad20-8e2ff8553040, Agent: opus5_initial_1, LLM: anthropic/claude-opus-5):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 9 months.
- Customer-first detection falls from 40% of incidents to under 10% within 6 months and under 5% within 12.
- Median time to mitigate falls from 3 h 10 min to under 60 minutes within 12 months.
- An Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes.
- Status page updated within 15 minutes of SEV1 declaration and 30 minutes of SEV2 in 95% of cases.
- Monthly alert volume falls from 3,400 to under 500 pages, with actionability above 75%; out-of-hours pages under 2 per person per week.
- All six legacy alerting tools consolidated into one paging platform, legacy paging paths disabled, by week 16.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by their due date.
- SLA credits fall from $1.3M to under $400k in the first 12 months.
- Customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar.
- All 28 teams onboarded by week 24, with 24x7 rotations of 6+ certified responders for every Tier 0/1 team.
- 30+ certified Incident Commanders and 20+ certified Communications Leads, giving 24x7 primary and secondary command cover.
- Paid on-call policy approved by HR, Legal and Finance and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among engineers on rotation.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked action items.
Steps (29):
1. Program charter, executive mandate and funding
Convert the CEO's frustration into a named program with one accountable owner, a budget and a deadline that is earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, the 28 teams will opt out.
- Appoint a single Incident Management Program Lead (Head of Reliability/SRE) with direct exec sponsorship from CTO and CEO.
- Form a steering group: CTO, VP Eng, Head of Support/CS, CISO/Compliance, Legal, Finance (SLA credits), HR (on-call pay).
- Set the non-negotiables: one severity scale, one paging tool, one postmortem format, mandatory action tracking, paid on-call.
- Fix the timeline: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready by week 28 (four weeks of buffer before the audit).
- Approve budget lines: tooling (~$150–250k/yr), on-call compensation (~$600k–1M/yr), 2–3 dedicated program FTEs. Anchor it against $1.3M of credits plus incident cost.
2. Forensic baseline of the 31 incidents and the alert estate (depends on: 1)
Before designing anything, rebuild the facts. Re-open all 31 incidents and profile the 3,400 monthly alerts so every later design decision is evidence-based.
This also creates the "before" picture the exec and the auditor will compare against.
- Re-code each incident: trigger, service, detection source (customer vs monitor), timestamps for detect/acknowledge/declare/mitigate/resolve, who led, credits paid, root cause family.
- Quantify the 40% customer-first detections: which signal was missing in each case.
- Classify the two "nobody in charge" incidents minute by minute; use them as the burning-platform story.
- Audit the six alerting tools: volume per tool, per team, per alert rule; identify the top 50 rules that produce most of the 85% noise; find rules with no owner and no runbook.
- Baseline the numbers formally: MTTD 22 min, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed. Freeze them as the reference line.
3. Stakeholder listening tour and resistance map (depends on: 1)
Engineer pushback against "carrying a pager for other teams' code" is the main delivery risk. Treat it as a design input, not an attitude problem.
Run structured interviews across all 28 teams, plus Support, CS and Sales, in two weeks.
- Test the real objection: is it unpaid work, night sleep, unfamiliar code, poor runbooks, or fear of blame? Each has a different fix.
- Collect current informal practices — the 12 teams already on-call are the pilot candidates and the source of veterans.
- Document the promise that answers the objection: **you are only paged for services your team owns**, plus a trained commander who runs the incident and pulls in others.
- Map influencers and blockers by team; recruit 10–15 credible engineers as a design working group so the process is co-authored, not imposed.
- Survey baseline sentiment (trust in alerts, willingness to be on-call, burnout) to re-measure at 6 and 12 months.
4. Service ownership catalog and criticality tiering (depends on: 2)
You cannot page the right person across 180 services until each service has a named owning team. This is the foundation of both on-call fairness and severity mapping.
Build a machine-readable catalog (Backstage or equivalent) that is the single source of truth for routing.
- One owning team per service, a named engineering manager, a Slack channel, a paging escalation policy, a dependency list.
- Tier services by business impact: Tier 0 (money movement, ledger, auth, shared PostgreSQL cluster), Tier 1 (customer-facing but degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map each Tier 0/1 service to the customer-visible capability it supports (payment initiation, settlement, reporting, onboarding).
- Flag orphan services and cross-team shared components; force an ownership decision for each within 30 days, or schedule decommissioning.
- Publish coverage gaps to the steering group: any Tier 0 service without an owner is an executive escalation.
5. Severity scale and declaration criteria (depends on: 2, 4)
Define a five-level scale with objective, payments-specific triggers so declaration is a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade.
Each level triggers a fixed bundle of response, comms and postmortem obligations.
- **SEV1**: money movement stopped or incorrect, ledger integrity in doubt, data breach, full region loss, >10% of customers impacted. Triggers: immediate 24x7 page of IC + comms + exec, bridge within 5 min, status page within 15 min, mandatory postmortem, regulator assessment.
- **SEV2**: severe degradation, settlement at risk of missing a window, single large/strategic customer fully down, SLA breach likely. Triggers: IC paged, status page within 30 min, mandatory postmortem.
- **SEV3**: partial or workaround-available degradation, no credit exposure. Team-led, business-hours comms, postmortem optional but encouraged.
- **SEV4/5**: minor or internal-only; ticket-tracked, no paging.
- Add auto-escalation rules: any SEV3 open >2 h, or any incident touching the shared ledger cluster, becomes SEV2 automatically. Include a severity decision tree and 12 worked examples drawn from the 31 real incidents.
6. Incident roles, decision authority and handover rules (depends on: 5)
Solve the "nobody in charge for an hour" failure by making command explicit, transferable and logged.
Define five roles with written responsibilities, entry criteria and explicit authority.
- **Incident Commander**: owns the incident, not the fix. Authority to declare severity, pull any engineer, approve customer-impacting mitigations, invoke failover and authorise spend. The IC never types in the terminal.
- **Communications Lead**: owns status page, internal updates, account-manager briefings and the exec summary. Single voice to customers.
- **Scribe**: maintains the timeline, decisions and open questions; feeds the postmortem and the audit evidence trail.
- **Subject-Matter Responders**: engineers from owning teams; they investigate and remediate, and report to the IC.
- **Executive Liaison** (SEV1 only): shields the IC from exec questions and owns regulator/board escalation.
- Rules: the IC role is assumed within 5 minutes of declaration, stated explicitly in the channel ("I am IC"), and any handover is announced and logged. Roles may be combined below SEV2; never at SEV1.
7. 24x7 incident command coverage model (depends on: 6, 4)
Command is staffed by a small, trained, cross-team pool — not by 28 teams individually. This is what makes 24x7 realistic in one New York time zone.
Design a central rotation that scales with a growing certified pool.
- Create a **Duty Incident Commander** rotation of 25–35 certified volunteers (target ~1 per team, plus managers and senior engineers), giving each person roughly one week per 6–8 months.
- Pair with a Duty Comms Lead rotation (Support/CS leads plus engineering managers, ~15–20 people) and a Scribe pool (rotating, lowest barrier, used as the training entry point).
- Coverage: primary + secondary IC at all times; hard 5-minute acknowledgement SLA with automatic failover to secondary, then to the on-call engineering director.
- Night coverage options to evaluate in writing: US-only rotation with paid night stipend now; a Lisbon/Dublin or APAC follow-the-sun cell as a 12-month option; a 24x7 NOC-style triage desk for first-line detection.
- Eligibility: certification required (S21); commanders are volunteers with manager approval and can step out with 30 days' notice.
8. Team on-call structure, rotations and routing rules (depends on: 4, 6)
Rebuild team on-call around the principle that answers the pushback: **you are only paged for code your team owns**.
Apply a tiered obligation so 28 teams are not treated identically.
- Tier 0/1 owning teams (expected 14–18 teams): 24x7 primary + secondary, minimum 6 people per rotation, one-week shifts, handover Wednesday mornings.
- Tier 2/3 teams: business-hours on-call with a best-effort out-of-hours escalation path, no night paging.
- Platform/Infrastructure and Database teams: 24x7, since they own the shared PostgreSQL ledger cluster and the Kubernetes/regional layer.
- Rotations under 6 people are merged across teams or backfilled by hiring; no rotation of fewer than 4 is approved.
- Routing: every page resolves through the service catalog to the owning team's escalation policy; cross-team pages are made by the IC, never by an alert.
- Guardrails: maximum one week in four, no on-call in the week after a SEV1 you led, protected recovery time after any night page, and a per-person page budget (see S10).
9. On-call compensation, labour compliance and fairness policy (depends on: 8, 3)
Unpaid on-call is both a retention risk and a legal exposure in New York. Paying for it is the fastest way to convert resistance into participation.
Design the scheme with HR, Legal, Finance and Payroll, and publish it before asking anyone to sign up.
- Base stipend per week on rotation, differentiated by tier: e.g. $800–1,200 for 24x7 Tier 0/1, $300–500 for business-hours rotations, with premiums for holidays and weekends.
- Per-incident payment for out-of-hours activation (e.g. $150 per night page plus hourly beyond one hour) and guaranteed time-off-in-lieu after night work.
- Separate Duty IC stipend, since command is a distinct and heavier burden.
- Verify FLSA exempt/non-exempt treatment, NY State wage rules and overtime exposure for non-exempt staff; document the legal review.
- Budget and model the annual cost; get board/CFO approval as a line item, benchmarked against $1.3M of credits.
- Add non-cash elements: on-call time counted as delivery load (teams reduce sprint commitment by ~15%), incident leadership recognised in promotion criteria, and a public quarterly report of on-call load per team.
10. Alert quality standard and page budget (depends on: 2, 4)
3,400 alerts a month at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a contractual condition of paging someone.
Publish a standard, then enforce it mechanically.
- Every paging alert must have: a named owning team, a documented customer impact, a runbook link, a tested threshold, and a severity mapping. Alerts failing this are demoted to ticket or deleted.
- Page only on symptoms that affect customers (SLO burn rate, error budget, queue depth against settlement deadlines); cause-based CPU/memory alerts become dashboards or tickets.
- Set a **page budget**: maximum 2 out-of-hours pages per person per week. Breach triggers a mandatory alert-tuning sprint for the owning team and blocks new alert creation.
- Auto-quarantine: any alert that fires more than 5 times a month without action, or has >70% no-action acknowledgements, is silenced automatically and returned to its owner.
- Monthly alert review per team: kill, tune, or keep, with the numbers on screen.
- Target: 3,400 → under 500 pages a month, with actionability above 75% within six months.
11. Detection uplift: SLOs, synthetic journeys and ledger assurance (depends on: 10, 4)
The goal is to stop customers telling you first. Detection must be driven by customer-visible outcomes, not host metrics.
Instrument the money path end to end and alert on it.
- Define SLOs for each Tier 0/1 customer capability: payment initiation success rate, authorisation latency, settlement file timeliness, API availability and reporting freshness. Tie them to the 99.95% contractual SLA with a stricter internal target.
- Deploy synthetic transactions from outside the platform, in both regions, every 60 seconds, covering the full payment lifecycle including a small real-value canary flow where feasible.
- Add ledger assurance checks: continuous double-entry balance reconciliation, replication lag and failover-readiness alarms on the shared PostgreSQL cluster, and settlement-window countdown alerts.
- Build per-customer anomaly detection for the top 100 accounts (volume drop, error spike) so a single-tenant outage is detected before the account manager calls.
- Create an "inbound signal" bridge: any support ticket or account-manager report matching impact keywords auto-creates a triage incident within 5 minutes.
- Track every incident's detection source; make "customer detected first" a reviewed defect with its own follow-up action.
12. Tool consolidation and incident platform implementation (depends on: 5, 7, 8, 10)
Collapse six alerting tools into one paging and incident platform so there is a single queue, a single timeline and a single audit record.
Run a short, time-boxed selection and migrate within the pilot window.
- Select an integrated stack: paging/on-call scheduling plus an incident management layer (e.g. PagerDuty + incident.io/FireHydrant, or a single vendor) and a hosted status page.
- Implement one-command declaration in Slack (`/incident declare`) that creates the channel and bridge, pages the Duty IC, sets severity, opens the timeline and starts the clock.
- Migrate all monitoring sources to route into the one platform; decommission direct paging from the legacy six and block new integrations that bypass it.
- Automate the evidence trail: timestamps, role assignments, severity changes, comms sent, and postmortem linkage exported for SOC 2.
- Integrate with the service catalog for routing, Jira for actions, Salesforce/CS tooling for affected-customer lists, and Zoom/Slack Huddle for the bridge.
- Hard requirement: the platform must work when AWS in one region is down — verify out-of-band paging (SMS/phone) and a printed/offline fallback runbook.
13. Detection-to-escalation path and the five-minute command rule (depends on: 6, 7, 12)
Write the single path from "something looks wrong" to "someone is in charge", and make it impossible to skip.
The design target is detection to commander in under five minutes, any hour.
- Entry points: automated alert, engineer observation, support ticket, account manager, partner bank, customer-facing SEV hotline. All converge on the same declaration command.
- Anyone in the company may declare up to SEV2; nobody is punished for over-declaring. Publish that rule in writing and repeat it.
- Auto-page ladder: Duty IC (5 min) → secondary IC (5 min) → on-call Director (10 min) → CTO. Same ladder for the owning team's responder.
- Cross-team pull: the IC can page any team's on-call directly, with a 10-minute acknowledgement obligation. This is the reciprocal commitment that makes single-team ownership viable.
- Explicit takeover protocol: if no one claims IC within 5 minutes, the platform assigns it and announces it; the assignee cannot decline, only hand over.
- Define standing severity triggers for immediate regional failover, ledger read-only mode and partner-bank notification, with pre-authorised decision rights so the IC does not wait for an executive.
14. Internal communications protocol (depends on: 6, 12)
Standardise the internal channel so responders, executives and support see the same picture without interrupting the IC.
Separate the working channel from the audience channel.
- One incident channel per incident (auto-created), one bridge, and a read-only broadcast channel for executives, Support and Sales.
- Update cadence by severity: SEV1 every 30 minutes even if nothing has changed; SEV2 every 60 minutes; SEV3 at state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, ETA or next update time, current IC and Comms Lead.
- Exec briefing rule: executives ask questions only to the Executive Liaison; the IC is not interrupted. Publish this as a behavioural expectation signed by the exec team.
- Support/CS enablement: a live affected-customer list and a holding statement within 15 minutes of SEV1/SEV2 so the front line is never guessing.
- Handover protocol for incidents beyond 4 hours: formal IC handover checklist, fatigue rule, and staffing of a second shift.
15. Customer communications and status page policy (depends on: 5, 14)
Customers currently learn of outages from their own monitoring and hear from whoever happens to be around. Replace that with a timed, owned, pre-approved process.
The Comms Lead is the single author; templates remove the need to write under pressure.
- Timing commitments: status page posted within 15 minutes of SEV1 declaration and 30 minutes for SEV2; updates every 30/60 minutes; resolution notice within 30 minutes of mitigation; customer-facing summary within 5 business days for SEV1.
- Pre-approve 12–15 templates with Legal and Comms (degradation, delay in settlement, API errors, security event, third-party failure) so nothing needs legal review mid-incident.
- Subscription-based status page with per-component granularity mapped to the customer capabilities from S11, plus an email/webhook/RSS feed.
- Tiered outreach: top 100 accounts get a direct named call or email from their account manager within 30 minutes of SEV1, with a briefing pack from the Comms Lead; long tail gets the status page and a proactive email.
- Rules of language: state impact and next update time, never speculate on cause, never assign blame to a vendor before facts are confirmed.
- Run a quarterly customer-perception check with the top accounts on whether comms were timely and useful.
16. Regulatory, partner and legal notification playbook (depends on: 5, 15)
In payments, some incidents are reportable and the clock starts at detection. Build the assessment into the process so it is never an afterthought.
Work with Legal, Compliance and the CISO to produce a decision tree and contact matrix.
- Map obligations: NYDFS Part 500 (72-hour cybersecurity event notification), state breach laws, GLBA/FTC Safeguards, PCI DSS if card data is in scope, sponsor-bank and card-network contractual notice windows, and any FinCEN/OFAC implications.
- Add a mandatory regulatory-assessment checkpoint to every SEV1 and every security-related SEV2, owned by the Executive Liaison, completed within 2 hours of declaration and recorded even when the answer is "not reportable".
- Build the contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel, with 24x7 numbers and named backups.
- Pre-draft notification letters and hold them under legal privilege review.
- Check customer contracts for bespoke notification SLAs (often 1–4 hours for enterprise accounts) and encode them in the customer tiering.
- Test the playbook once per quarter as part of the simulation programme.
17. SLA credit and financial impact workflow (depends on: 5, 15)
Link incidents to money so severity, credits and prioritisation stay consistent — and so Finance stops being surprised.
Make credit calculation an automated output of the incident record, not a negotiation.
- Define the availability measurement method per contract, per component, and agree it with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and the per-capability telemetry; generate a proposed credit schedule within 5 business days of resolution.
- Decide the posture: proactive credits for the top tier (reputation upside) versus claims-based for the rest; document the approval chain.
- Track credits per incident and per root-cause family; feed a quarterly report showing which reliability investments would have prevented which credits.
- Set a target: reduce credits from $1.3M to under $400k in year one, and use that delta as the ongoing business case.
18. Postmortem standard and blameless review forum (depends on: 5, 6)
Replace "some incidents, various formats" with a mandatory, single-format, blameless process with fixed deadlines.
The discipline is in the deadlines and the forum, not in the template.
- Mandatory for: every SEV1 and SEV2, every incident where a customer detected it first, every incident over 2 hours, every repeat of a known cause, and every near-miss involving the ledger. Optional but templated for SEV3.
- Fixed timeline: draft within 3 business days, peer review within 5, published company-wide within 10. The IC owns delivery; the owning team's manager is accountable.
- One template: timeline, customer and financial impact, detection analysis (why not sooner), response analysis (why mitigation took as long as it did), contributing factors, what went well, action items with owner and due date.
- Blameless rules in writing: describe systems and decisions in the context available at the time; no individual named as a cause; HR and management commit that postmortems are never used in performance reviews.
- Weekly 60-minute Incident Review Board: reviews all postmortems from the prior week, challenges quality, ratifies severity, and approves or rejects action items. Attendance by engineering directors is mandatory.
- Publish a searchable postmortem library and a quarterly "top five recurring causes" analysis.
19. Action item ownership and tracking system (depends on: 18, 12)
11 of 64 actions closed is the single clearest symptom of a process nobody enforces. Give actions the same status as customer commitments.
Track them where engineering work already lives, with visible escalation.
- Every action gets: a named individual owner (not a team), a priority class, a due date and a Jira ticket auto-created from the postmortem.
- Priority classes with hard SLAs: P0 prevents recurrence of a SEV1, due in 30 days; P1 in 60 days; P2 in 90 days. P0s are committed into the next sprint before any roadmap work.
- Capacity rule: teams reserve a standing 15–20% of sprint capacity for reliability and incident actions. Without reserved capacity, the actions will not land.
- Escalation ladder for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Monthly reporting of closure rate by team in the engineering leadership review; include it in manager performance objectives.
- Target: 90% of P0/P1 actions closed on time within two quarters.
20. Runbooks, major-incident playbooks and the on-call readiness bar (depends on: 4, 8)
Nobody can respond well to unfamiliar systems at 3 a.m. without runbooks — and poor runbooks are a real part of the pager resistance.
Define a minimum readiness bar that a service must meet before it is allowed to page anyone.
- Readiness checklist per Tier 0/1 service: current architecture diagram, dependency map, dashboard link, alert-to-runbook mapping, rollback procedure, feature-flag kill switches, escalation contacts, and a data-loss/latency impact statement.
- Write major-incident playbooks for the top failure modes derived from S2: shared PostgreSQL ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank/third-party outage, settlement-window breach, and suspected security compromise.
- Prioritise the shared ledger: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and a clearly stated data-loss tolerance (RPO/RTO) signed off by the exec.
- Runbooks must be tested at least twice a year in a drill; untested runbooks are marked stale in the catalog.
- Enforcement: a service without readiness sign-off cannot create paging alerts, and the gap is reported to its director.
21. Training, certification and the commander academy (depends on: 6, 13, 14, 18)
Command is a skill, not a title. Build a certification path so 24x7 coverage is staffed by people who have practised.
Use a tiered curriculum with real assessment.
- **Scribe** (2 hours): timeline discipline and tooling. The entry point for everyone.
- **Responder** (half a day): severity scale, declaration, escalation, runbook use, comms hygiene. Mandatory for every engineer joining an on-call rotation.
- **Incident Commander** (two days plus shadowing): command presence, delegation, decision-making under uncertainty, severity calls, handover, exec management. Requires two shadowed incidents and one simulated SEV1 before certification.
- **Communications Lead** (one day): status-page writing, customer tiering, legal boundaries, regulator triggers.
- Certification is valid 12 months and renewed via a simulation; the register of certified people is an audit artefact.
- Add on-call onboarding per team: a new joiner shadows two shifts before holding primary, and never holds primary in their first 90 days.
22. Simulation programme: game days, drills and wheel of misfortune (depends on: 21, 12, 20)
The process must be rehearsed before it meets a real SEV1. Simulations also build the commander pool and expose runbook gaps cheaply.
Run a standing calendar rather than one-off exercises.
- Monthly 60-minute tabletop ("wheel of misfortune") per engineering group, using a real past incident from the 31.
- Quarterly full-scale game day in production or a production-like environment: regional failover, ledger replica promotion, dependency failure, with the whole role structure activated and timed.
- Twice-yearly unannounced paging drill to measure real acknowledgement times at night.
- One security-incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Every exercise produces a lightweight postmortem and action items in the same system as real incidents.
- Measure and publish drill metrics: time to IC, time to first status update, time to correct mitigation decision.
23. Pilot with wave 0 teams (depends on: 22, 9, 11)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Run a six-week pilot with tight measurement and a public verdict.
- Select 5–6 teams: core payments, ledger/database, platform/Kubernetes, API gateway, plus two of the 12 teams already on-call.
- Activate the full stack for them: new severity scale, Duty IC rotation, single paging tool, alert budget, status-page policy, mandatory postmortems, paid on-call.
- Hold a weekly pilot retro; expect and document 20–40 process defects, and fix them in the standard before rollout.
- Validate the hard questions: does the 5-minute IC rule hold at 3 a.m.? Do cross-team pulls get answered? Is the severity tree unambiguous?
- Exit criteria: MTTD under 10 minutes for pilot services, IC assigned within 5 minutes in 95% of incidents, page volume down 50%, all postmortems on time, positive on-call sentiment.
- Publish a one-page pilot result to the whole company — this is the main adoption argument for the remaining teams.
24. Metrics, dashboards and the review cadence (depends on: 12, 18, 23)
Instrument the process itself, so improvement is visible and the audit has evidence of monitoring and review.
Define a small set of metrics with owners and a fixed meeting rhythm.
- Response metrics: MTTD, time to declare, time to IC assigned, MTTA, MTTM, MTTR, incidents per month by severity, % detected by customers first.
- Quality metrics: page volume per person per week, alert actionability rate, page budget breaches, postmortem on-time rate, action closure rate and ageing.
- Business metrics: SLA credits paid, availability against 99.95% per capability, error-budget consumption, repeat-incident rate.
- People metrics: on-call load distribution across teams, out-of-hours pages per person, on-call sentiment and attrition among on-call staff.
- Cadence: weekly Incident Review Board (postmortems and actions), monthly Reliability Review (metrics per team, alert hygiene, on-call load), quarterly Executive/Board review (credits, trends, investment asks), annual policy review.
- Every metric gets a target and a named owner; dashboards are self-serve and public inside the company.
25. Wave rollout across all 28 teams with readiness gates (depends on: 23, 24)
Roll out in four waves of six to eight teams, every three weeks, ordered by criticality. Each wave passes an explicit gate rather than a deadline.
Gates keep quality high and make the standard credible.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with 6+ certified responders, one IC candidate nominated, one drill passed.
- Assign each wave a named coach from the program team for three weeks of hands-on support.
- Gate criteria are checked and signed by the director; teams that fail are re-scheduled, not waived.
- Freeze legacy tooling per wave: after onboarding, the old alerting paths are disabled, not left as a fallback.
- Publish a live adoption scoreboard by team so progress is social, not administrative.
26. SOC 2 control mapping, evidence automation and internal dry-run audit (depends on: 18, 19, 25)
Design the process so audit evidence is a by-product of doing the work, then test it before the auditors do.
Engage the auditor early to confirm the interpretation of controls.
- Map the process to the Trust Services Criteria: CC7.3 and CC7.4 (incident identification, response, recovery), CC7.2 (monitoring), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), plus availability criteria A1.2.
- Produce and approve formal policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard — versioned, signed, annually reviewed.
- Automate evidence: incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, postmortem library, action-item closure reports, training and certification register, drill records.
- Confirm the observation window with the auditor and ensure the process is operating for a minimum of three months before fieldwork.
- Run an internal dry-run audit at month six: sample 15 incidents and walk the full evidence chain; fix gaps with 8 weeks to spare.
- Keep a remediation log for any incident where the process was not followed, with the corrective action — auditors respond better to documented exceptions than to a claim of perfection.
27. Change management, incentives and communications campaign (depends on: 3, 9, 23)
Run this in parallel from day one. The process will be judged by engineers on fairness, and by executives on visible results.
Communicate the deal explicitly and repeatedly.
- The deal in one sentence: **you are paid for on-call, you are paged only for what you own, a trained commander runs the incident, and your postmortem actions get real sprint capacity**.
- Launch communications: CTO all-hands, per-team roadshows, a one-page process card for laptops, an internal wiki hub and a Slack support channel with a 4-hour answer SLA.
- Recognition: incident-response contribution in promotion criteria and performance frameworks, quarterly awards for best postmortem and biggest alert-noise reduction, public thanks after every SEV1.
- Manager accountability: adoption, alert hygiene, action closure and on-call load in each engineering manager's quarterly objectives.
- Handle the exceptions: a written path for engineers who cannot do nights (caring responsibilities, health), covered by stipended volunteers elsewhere.
- Track sentiment quarterly and publish the results, including bad news, to keep credibility.
28. Program risk register and contingency planning (depends on: 1)
Name the ways this program fails and pre-commit the response. Review it monthly in the steering group.
The main risks are predictable.
- **Volunteer shortfall for the IC pool**: contingency is to make command a rostered duty for engineering managers and senior engineers until the pool reaches 25.
- **Compensation not approved in time**: fall back to time-off-in-lieu plus a phased stipend, but do not launch mandatory night on-call without some compensation.
- **Tool migration slipping**: keep the single-queue requirement and cut scope on the incident-management layer, not on paging consolidation.
- **Alert pruning causing a missed incident**: prune from paging to ticket first, observe for 30 days, then delete; keep a recovery path.
- **Burnout or attrition among the 12 experienced on-call teams**: monitor load weekly and cap individual page counts.
- **A major SEV1 mid-rollout**: pre-agree that the program lead becomes a full-time responder and the wave schedule slips by one wave, with the steering group informed the same day.
29. Continuous improvement, maturity roadmap and post-audit sustainability (depends on: 24, 25, 26)
Protect against the classic failure: the process decays once the audit is passed. Build the second-year plan before the first year ends.
Set a maturity model and a forward roadmap with owners.
- Quarterly process retrospective with the IC pool: what in the process itself slowed us down, what needs simplifying, is the severity scale calibrated?
- Re-baseline targets every six months; a process that hits all targets is under-ambitious.
- Year-two roadmap candidates: follow-the-sun coverage cell, automated mitigation and self-healing for the top three recurring causes, error-budget policy that gates releases, per-customer real-time impact reporting, and blast-radius reduction for the shared ledger cluster (the largest single structural risk).
- Move from lagging metrics (MTTR) to leading ones (error-budget burn, near-miss rate, drill performance).
- Make the annual policy review, certification renewal and drill calendar permanent calendar items owned by the Head of Reliability, independent of the audit cycle.
- Report to the board quarterly on availability, credits and incident trends so the process keeps executive attention after SOC 2 is signed.
Previous Proposal 2 (ID: 7ee53cee-6357-4ed7-bd99-535e92e8b1e6, Agent: gpt5.6-sol_initial_2, LLM: openai/gpt-5.6-sol):
Estimated Complexity: high
Success Metrics: - Within 7 days, every suspected SEV0–SEV2 has one incident record, one channel, and a named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named incident commander for more than 10 minutes.
- By day 30, 100% of Tier 0 and Tier 1 services have a named owner, primary escalation, secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and subject-matter coverage.
- By day 120, 100% of production services have sustainable ownership and tested escalation paths.
- At least 95% of SEV0 and SEV1 pages are acknowledged within 5 minutes by month 3.
- At least 95% of SEV2 pages are acknowledged within 10 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to 10 minutes by day 90 and 5 minutes by month 6.
- The proportion of incidents first detected by customers falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to below 90 minutes by day 120 and below 60 minutes by month 6.
- At least 95% of qualifying incidents meet their initial customer-communication deadline by month 3.
- At least 95% of published incidents meet their required update cadence by month 3.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without loss of Tier 0 or Tier 1 detection coverage.
- Monthly pages fall from 3,400 to no more than 1,500 by day 90, with alert actionability and missed-detection reviews used as countermeasures against unsafe suppression.
- 100% of new paging alerts satisfy the owner, runbook, dashboard, action, severity, and escalation quality rules by day 60.
- 100% of required SEV0 and SEV1 postmortems are drafted within 3 business days and reviewed within 5 business days by month 2.
- At least 90% of postmortem actions are completed by their approved due dates by month 6.
- All 53 currently open historical actions are triaged within 30 days; all unaccepted high-risk items are completed within 90 days.
- Repeat incidents with the same unaddressed contributing factor decline by at least 50% within 6 months.
- Every primary rotation has at least six trained responders or a documented, time-limited executive exception by day 120.
- No responder is routinely scheduled more frequently than one primary week in six by day 120.
- Two end-to-end cross-company exercises, including regional and ledger scenarios, are completed before the audit, with all critical findings assigned and tracked.
- Monthly availability meets or exceeds the 99.95% contractual target by month 6, with exceptions reviewed at the executive reliability meeting.
- SLA credits decline by at least 50% on an annualized trailing basis by month 8.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (19):
1. Establish ownership, authority, and funding
Launch the program within 48 hours under an executive sponsor. Give one program owner authority to standardize incident management across all 28 teams.
- Name the CTO or equivalent as executive sponsor and a Head of Incident Management or Reliability as directly accountable owner.
- Form a working group with Engineering, SRE or Platform, Product, Support, Customer Success, Communications, Security, Legal, Compliance, Risk, HR, Finance, and Internal Audit.
- Approve authority for an incident commander to stop deployments, roll back releases, disable features, shift traffic, invoke continuity plans, and pause payment processing when integrity is at risk.
- Preserve financial controls. The incident commander may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access requirements.
- Fund paging tools, compensation, training, observability work, exercises, and dedicated reliability capacity.
- Reserve engineering capacity for incident remediation. Start with 10% of capacity and adjust through quarterly risk reviews.
- Record the current baselines: 31 customer-impacting incidents, 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M in credits, 3,400 monthly alerts, 85% noise, and 11 of 64 actions closed.
- Maintain a risk register for staffing gaps, shared-ledger concentration, regional failover, alert coverage, third parties, and audit readiness.
2. Install immediate minimum controls (depends on: 1)
Put an interim process in place during the first seven days. Do not wait for tool consolidation, policy perfection, or the SOC 2 audit.
- Publish a one-page interim severity guide and incident declaration procedure.
- Establish one continuously monitored incident declaration path through chat, telephone, and the paging system.
- Create a standard incident channel, conference bridge, incident document, and event naming convention.
- Staff an interim primary and backup incident commander at all times. Compensate this duty retroactively under the final compensation policy.
- Give trained duty personnel access to the status page, paging system, dashboards, support queue, service catalog, and emergency contacts.
- Require an incident commander to be named within 10 minutes for every suspected major incident.
- Direct Support to escalate credible customer reports immediately rather than waiting for engineering confirmation.
- Triage all 53 open historical postmortem actions. Complete, re-plan, or formally risk-accept the items affecting ledger integrity, payment duplication, regional resilience, security, and detection first.
- Hold a daily 15-minute operational review until permanent controls are working.
3. Create the service and dependency catalog (depends on: 1)
Build a reliable ownership map for all production services and customer journeys. This is the basis for paging, escalation, impact assessment, and audit evidence.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, data stores, queues, external processors, banking partners, and customer-facing endpoints.
- Assign each component a single accountable team, primary responder group, secondary escalation group, engineering manager, and product owner.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 according to financial integrity, customer impact, dependency centrality, and contractual obligations.
- Treat the ledger, payment orchestration, authentication, settlement, reconciliation, and critical shared infrastructure as Tier 0 or Tier 1.
- Map every important customer journey to its service, database, cloud-region, and third-party dependencies.
- Record SLOs, RTOs, RPOs, data classification, dashboards, runbooks, deployment controls, feature flags, and failover methods.
- Assign separate but coordinated responders for the ledger application and the shared PostgreSQL platform.
- Document whether each service is active-active, active-passive, or region-bound. Identify dependencies that make nominal regional redundancy ineffective.
- Make missing ownership or missing runbooks a release-blocking risk for Tier 0 and Tier 1 services.
4. Adopt severity and incident lifecycle standards (depends on: 1)
Approve one impact-based severity model for operational, security, data, and third-party incidents. When evidence is incomplete, start at the higher credible severity and downgrade later.
- SEV0, crisis: Use for actual or credible unauthorized, lost, duplicated, or corrupted movement of money; ledger integrity loss; material security compromise; material data exposure; both-region failure; or an event likely to require crisis or regulatory management. Page all roles immediately, engage executives, Security, Legal, Compliance, and Risk, and consider pausing payment activity.
- SEV1, critical: Use for widespread inability to initiate, process, settle, or reconcile payments; a core journey failing without a viable workaround; material regional impact; fast SLA-budget exhaustion; or an imminent integrity risk. Staff all incident roles, notify the executive duty officer, and publish customer communications.
- SEV2, major: Use for a material customer subset, one or more critical customers, significant degradation with a workaround, partial transaction failure, or a likely contractual impact. Assign an incident commander and subject-matter responders; add communications and scribe roles whenever customers are affected.
- SEV3, minor: Use for localized, low-impact degradation with no financial-integrity, security, regulatory, or material contractual risk. The owning team leads the response and keeps an internal record; external communication is not normally required.
- Base severity on actual or credible impact, not the seniority of the reporter, number of alerts, or presumed complexity of the fix.
- Permit any employee to declare an incident. Only the incident commander may lower severity after recording the evidence and rationale.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of customer impact. Define resolution only after stability, backlog processing, transaction recovery, and required ledger reconciliation are complete.
- Start measurement from the earliest reliable indication of impact, including telemetry, customer reports, and partner notifications.
5. Define roles and sustainable 24x7 staffing (depends on: 3, 4)
Separate command from technical remediation. This allows trained commanders to coordinate any incident without asking engineers to debug code they do not own.
- Incident commander: Owns severity, priorities, role assignment, escalation, decision cadence, mitigation strategy, handoffs, and final closure. One person has command at a time.
- Communications lead: Owns internal notices, status-page updates, account-manager briefs, approved customer language, and coordination with Legal or regulators.
- Scribe: Maintains a timestamped timeline of observations, decisions, commands, owners, and status changes. Automation may assist but does not replace human validation for SEV0 and SEV1.
- Subject-matter responders: Diagnose and mitigate only services or domains for which they have accepted ownership, training, access, and runbooks.
- Executive duty officer: Removes organizational obstacles and approves exceptional business decisions. This role does not take command unless a formal transfer occurs.
- Security, Legal, Compliance, Support, Vendor Management, and Business Continuity join according to predefined triggers.
- Create a company-wide incident-command rotation with at least eight certified primary commanders and eight qualified backups. Use weekly rotations with explicit handoffs.
- Create similarly sustainable communications and scribe pools using Engineering Operations, Support, Customer Operations, and Communications personnel.
- Group service responders into approximately 8–12 coherent product or platform domains rather than creating 28 fragile rotations. Each domain rotation should normally contain at least six trained responders.
- Do not place an engineer into another team's responder pool without training, access, runbooks, shadow shifts, and explicit acceptance by both teams.
- Maintain dedicated database-platform and ledger-application escalation coverage for the shared PostgreSQL environment.
- Require distinct people for commander, communications, and primary technical lead during SEV0 and SEV1 incidents.
- Require a verbal and written handoff when any incident role changes. Record the exact time and the new role owner.
6. Implement compensation and fatigue safeguards (depends on: 5)
End unpaid on-call before expanding coverage. Treat availability, interrupted personal time, and overnight recovery as compensable work.
- Pay a fixed stipend for each primary on-call week and a secondary stipend equal to a defined percentage of the primary amount.
- Pay a higher holiday stipend. Apply overtime and call-out rules to non-exempt employees as required by law.
- Give exempt employees a minimum call-out credit or equivalent paid recovery time for material after-hours work.
- Provide a paid recovery day after prolonged overnight work, a SEV0, or a qualifying SEV1. Managers must arrange daytime coverage rather than expecting normal output.
- Have HR, Finance, and employment counsel publish dollar amounts, tax treatment, eligibility, and payroll procedures within 14 days. Apply the policy consistently across teams and locations.
- Target rotations no more frequent than one week in six. Exceptions require a time-limited staffing plan and executive risk acceptance.
- Avoid consecutive primary and secondary weeks. A person must not be primary for two simultaneous domain rotations.
- Track after-hours pages, sleep interruptions, swaps, missed acknowledgements, and reported burnout by rotation.
- Trigger a staffing or alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
- Permit responders to declare themselves temporarily unfit after overnight work without performance penalty.
7. Consolidate incident and paging tooling (depends on: 4)
Create one operational system of record while migrating safely from the six current alerting tools. Consolidation must reduce ambiguity without creating a monitoring gap.
- Select one enterprise paging and escalation platform and one integrated incident record.
- Initially ingest events from all six tools. Deduplicate, correlate, and route them through the new platform before retiring sources.
- Integrate paging with chat, conference bridges, ticket tracking, the service catalog, observability tools, and the customer status page.
- Automatically capture declaration time, acknowledgements, role assignments, severity changes, messages, decisions, mitigated time, and resolved time.
- Use role-based access, multifactor authentication, break-glass controls, immutable audit logs, and periodic access reviews.
- Provide mobile and telephone fallback paths if chat, identity, or the primary paging tool is unavailable.
- Test paging, escalation, status publication, and conference access every week.
- Retire a legacy alert path only after its signals have named owners, successful end-to-end tests, and at least two weeks of verified operation in the new platform.
8. Improve detection and enforce alert quality (depends on: 3, 7)
Shift detection toward customer journeys, payment outcomes, and ledger integrity. Infrastructure metrics alone will not solve the current customer-first detection problem.
- Instrument payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, and reporting with SLOs and business-level success metrics.
- Run external synthetic transactions and API checks from outside the production boundary and from both AWS regions.
- Monitor transaction failure rates, processing latency, queue age, unprocessed volume, reconciliation breaks, unexpected ledger balances, duplicate identifiers, regional asymmetry, and third-party response quality.
- Correlate application telemetry with Kubernetes, AWS, PostgreSQL, network, deployment, and feature-flag events.
- Route high-priority support cases and credible partner notifications into the same incident declaration path within five minutes.
- Define noise as a page that is duplicate, informational, unactionable, non-production, or requires no timely human action.
- Require every paging alert to name an owner, affected service, urgency, customer or SLO risk, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Send non-urgent conditions to a ticket queue rather than a pager.
- Run new alerts in shadow mode for at least seven days unless an emergency risk exception is approved. Test both firing and recovery behavior.
- Review any alert with less than 50% actionability or more than three firings in seven days within two business days.
- Never silently disable a noisy alert. Verify compensating detection, record the decision, and assign a correction owner first.
- Review alert actionability, false positives, missed detection, and page load with every responder group each month.
9. Codify acknowledgement and escalation paths (depends on: 3, 4, 5, 7, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts when a qualifying signal or customer report enters the system.
- For SEV0 and SEV1, page the owning primary immediately; page the secondary after five unacknowledged minutes; page the domain manager and company incident commander at 10 minutes; and engage the executive duty officer by 15 minutes.
- For SEV2, require primary acknowledgement within 10 minutes and incident-command assignment within 15 minutes. Escalate to the secondary and manager when either target is missed.
- For SEV3, require acknowledgement within 30 minutes when immediate production action is needed. Otherwise create a prioritized work item.
- Automatically page the company incident commander for any credible integrity or security concern, cross-team event, customer-visible Tier 0 failure, regional event, or unresolved ownership question.
- If impact remains unknown after 15 minutes, raise severity rather than waiting for certainty.
- Let the incident commander summon dependency owners, cloud support, database support, payment processors, banking partners, and vendors through maintained escalation contacts.
- Test vendor contacts and premium-support entitlements quarterly.
- Route alerts with no valid owner to the central command rotation, then treat the missing ownership record as a control defect.
- Require human acknowledgement. Delivery to a device or chat channel does not count.
- Record every missed acknowledgement, failed escalation, and manual contact workaround for review.
10. Standardize live incident execution (depends on: 4, 5, 7, 9)
Give responders one concise operating procedure for the first minutes through resolution. Prioritize limiting customer and financial harm before proving a root cause.
- Open a dedicated channel, bridge, incident record, and timeline immediately for SEV0 through SEV2.
- Have the incident commander state severity, known impact, current hypothesis, immediate objective, assigned roles, and next update time.
- Freeze unrelated production changes during SEV0 and SEV1 incidents. Record exceptions approved by the incident commander.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Use pre-approved runbooks for region failover, Kubernetes recovery, PostgreSQL failover, credential rotation, queue recovery, and payment suspension.
- Guard against split-brain, replay, duplication, and out-of-order processing during regional or database recovery.
- Require reconciliation and controlled backlog processing before declaring payment or ledger incidents resolved.
- Keep diagnosis and mitigation workstreams separate when enough responders are available.
- State decisions and owners aloud and in the incident record. Avoid unrecorded direct-message command paths.
- Require stability for a severity-specific observation period before closure. Reopen the incident if impact recurs during that period.
- Conduct an explicit operational handback to the owning team, Support, and Customer Success.
11. Standardize internal, customer, and regulatory communications (depends on: 4, 5, 7, 10)
Communicate known impact early without waiting for a root cause. Use approved facts, acknowledge uncertainty, and give the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish an initial customer status within 15 minutes when disclosure is operationally and legally appropriate, then update every 15 minutes.
- For SEV1, notify internal stakeholders within 15 minutes, publish an initial customer status within 15 minutes, and update at least every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send an initial notice within 30 minutes. Update at least every 60 minutes.
- Do not normally publish SEV3 events. Notify specifically affected customers if contracts or material impact require it.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. State affected capabilities, customer symptoms, workarounds, regions, and next update time.
- Do not speculate about root cause, blame, security scope, recovery time, or data integrity.
- Give account managers a single approved briefing and an affected-customer list. Prohibit contradictory or improvised incident explanations.
- Maintain templates for outages, delays, data-integrity investigation, third-party failure, regional failure, security events, and resolution.
- Issue a resolution notice only after operational recovery and required reconciliation. Provide a customer-facing incident summary within five business days for qualifying events.
- Have Legal and Compliance maintain a jurisdiction, regulator, sponsor-bank, network, cyber-insurer, partner, and contract notification matrix.
- Where applicable, explicitly track the current New York cybersecurity-event notification clock, including the 72-hour requirement, without assuming every incident is reportable.
- Have Legal record the reportability decision, decision time, evidence, approver, deadline, and submission confirmation.
- Allow Security or Legal to limit public detail during an active threat, but require the reason and an alternative stakeholder plan to be recorded.
- Coordinate service-credit calculations and contractual notices with Finance and Customer Success from the same incident record.
12. Make postmortems mandatory and actionable (depends on: 4, 7, 10)
Use postmortems to improve systems and controls, not to assign personal blame. Keep performance or misconduct processes separate from the learning review.
- Require a postmortem for every SEV0 and SEV1.
- Require one for a SEV2 that affected customers, incurred credits, breached an SLO or contract, involved financial or data integrity, repeated a prior failure, exposed a control gap, or lasted more than two hours.
- Permit incident command, Security, Compliance, or the service owner to require a review for a near miss.
- Produce a factual draft within three business days and hold the cross-functional review within five business days.
- Use one template covering executive summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, what worked, what failed, and lessons.
- Include why monitoring did or did not detect the event before customers.
- Avoid a single-root-cause assumption. Examine technical, organizational, process, dependency, testing, and incentive factors.
- Give every action one owner, due date, priority, expected risk reduction, verification method, and linked engineering item.
- Classify actions as containment due within 7 days, corrective work due within 30 days, or strategic work normally due within 90 days.
- Require director approval and documented residual-risk acceptance for overdue high-risk actions.
- Verify effectiveness after implementation. Closing a ticket without evidence does not close the action.
- Publish broadly useful reviews internally. Maintain access-restricted versions for security, privacy, personnel, or legally privileged details.
13. Measure performance and review it routinely (depends on: 7, 8, 11, 12)
Use outcome, process, quality, and human-sustainability measures together. Do not reward teams for suppressing declarations or hiding incidents.
- Measure detection time from first impact to first internal signal, declaration time, acknowledgement time, role-staffing time, mitigation time, resolution time, and recurrence.
- Report both median and 90th percentile. Break results down by severity, service tier, customer journey, region, detection source, and owning domain.
- Track customer-first detection, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, and role conflicts.
- Track availability, error-budget consumption, failed-payment volume, delayed value, reconciliation breaks, impacted customers, contractual breaches, and service credits.
- Track alert volume, actionability, duplicates, after-hours pages, missed pages, pages per responder, and tool-source distribution.
- Track required postmortems completed on time, actions completed by due date, action age, verified effectiveness, and repeat contributing factors.
- Track rotation size, on-call frequency, swaps, recovery days, attrition signals, and quarterly responder sentiment.
- Hold a weekly operational review for recent incidents, overdue actions, alert problems, and upcoming risk.
- Hold a monthly executive reliability review covering trends, investment decisions, accepted risks, and SLA exposure.
- Hold a quarterly resilience and control review with Security, Compliance, Risk, Internal Audit, and Product leadership.
- Use team scorecards to direct investment and assistance, not individual performance penalties.
- Reconcile dashboard data against a monthly sample of incident records and customer cases to detect metric gaming or missing incidents.
14. Train and certify participants (depends on: 4, 5, 9, 10, 11, 12)
Train people before assigning full independent duty. Use paid working time for training, shadowing, exercises, and certification.
- Train all employees to recognize impact, declare an incident, and find the incident channel and status page.
- Train engineers and Support on severity, escalation, customer-report handling, evidence preservation, and financial-integrity precautions.
- Certify incident commanders through instruction, tabletop exercises, shadow incidents, and observed command performance.
- Train communications leads in status writing, contractual communications, regulator escalation, and avoiding unsupported claims.
- Train scribes in timestamping, decision capture, evidence hygiene, and separating fact from hypothesis.
- Require subject-matter responders to demonstrate dashboard, runbook, rollback, failover, and access competence for their assigned domain.
- Add training to new-hire onboarding and repeat role-specific certification annually.
- Appoint an incident-management champion in each of the 28 teams to collect feedback and support local adoption.
- Conduct listening sessions focused on pager fairness, cross-team boundaries, tooling friction, and psychological safety.
- Publish that command duty is process coordination, not responsibility for understanding or repairing another team's code.
15. Pilot and expand the on-call model (depends on: 3, 5, 6, 8, 9, 14)
Pilot the model on the highest-risk customer journeys before expanding it. Correct staffing, alert, access, and compensation defects at each stage gate.
- Start with the ledger, payment orchestration, Kubernetes platform, PostgreSQL platform, authentication, settlement, and Support intake.
- Run the command, communications, and domain rotations in parallel with existing paths for two weeks.
- Verify primary and secondary coverage, handoffs, access, runbooks, paging, conference access, status publication, and compensation processing.
- Require at least two shadow shifts before independent primary duty.
- Review every pilot page within one business day for routing accuracy, actionability, responder load, and missing context.
- Expand by customer journey and dependency domain, not by arbitrary team order.
- Provide central first-line triage if useful, but keep technical remediation with the accepted service owner.
- Do not use contractors or a managed service as the sole incident commander or sole owner of payment and ledger remediation.
- Permit temporary shared domain rotations only after service owners document training, access, runbooks, and escalation boundaries.
- Set an executive-reviewed deadline and remediation plan for any production service that cannot provide sustainable 24x7 ownership.
16. Exercise regional, ledger, and communication failures (depends on: 10, 11, 14, 15)
Validate the process under realistic conditions before relying on it. Begin in tabletop and staging environments, then use controlled production tests where risk permits.
- Run a company-wide incident-command tabletop within 30 days of policy approval.
- Exercise loss of one AWS region, Kubernetes control-plane degradation, shared PostgreSQL failure, payment-processor failure, queue backlog, credential compromise, and suspected duplicate payments.
- Exercise a simultaneous operational and security event to test command boundaries and disclosure control.
- Exercise status-page failure and loss of the primary chat or paging provider.
- Exercise overnight staffing, role handoff, executive escalation, account-manager messaging, and a potential regulator-notification decision.
- Validate backups, restore procedures, RPO, RTO, failover prerequisites, and post-recovery reconciliation.
- Do not inject uncontrolled changes into the production ledger. Use replicas, staging, simulations, or tightly governed production tests.
- Record exercise observations as tracked actions under the same ownership and due-date rules as incident actions.
- Run at least one domain exercise per quarter and two cross-company exercises before the SOC 2 audit.
17. Execute a time-boxed enterprise rollout (depends on: 2, 6, 7, 11, 12, 13, 15)
Use fixed implementation waves so the audit deadline does not become the start date. Report progress weekly and escalate missed stage gates as business risks.
- Days 0–7: Establish governance, interim command coverage, one declaration path, provisional severity, and daily operational reviews.
- By day 14: Approve the core policy, role definitions, communications timings, compensation design, and historical-action triage.
- By day 30: Complete Tier 0 ownership, certify the first command roster, begin the on-call pilot, enable standard incident records, and run the first tabletop.
- By day 60: Provide 24x7 coverage for all Tier 0 and Tier 1 customer journeys, integrate the six alert sources, and enforce postmortem tracking.
- By day 90: Migrate critical paging, implement customer-journey detection, complete status and regulatory playbooks, and materially reduce alert noise.
- By day 120: Assign sustainable ownership and escalation for every production service and complete the first controlled regional or continuity exercise.
- By day 180: Complete tool retirement decisions, verify action closure, rerun weak scenarios, and demonstrate improving detection and mitigation trends.
- In month 7: Conduct a mock audit and executive readiness review, leaving at least one month to correct evidence or operating defects.
- Use exception records with owners, expiry dates, compensating controls, and executive approval. Do not allow indefinite verbal exceptions.
18. Build SOC 2 evidence as the process operates (depends on: 1)
Design evidence collection at the start rather than reconstructing it before the audit. Demonstrate both control design and sustained operation.
- Map the incident process to applicable SOC 2 criteria with Compliance and the auditor, including detection, response, communication, change management, access, availability, and corrective action.
- Maintain approved, version-controlled policies, procedures, severity definitions, role descriptions, and exception records.
- Preserve rotation schedules, compensation activation, training attendance, certification, paging tests, access reviews, and exercise results.
- Preserve incident declarations, timestamps, role assignments, communications, decisions, status updates, postmortems, and corrective-action evidence.
- Record regulatory and contractual notification assessments, including decisions that no notification was required.
- Define retention, confidentiality, legal-hold, and access requirements for operational and security records.
- Sample evidence monthly and trace incidents from initial signal through action verification.
- Have Internal Audit or an independent control owner test the process in months 4 and 6.
- Correct control failures through tracked actions rather than editing historical records.
- Conduct the formal mock audit in month 7 using the same evidence populations expected for the external audit.
19. Sustain accountability and continuous improvement (depends on: 13, 17, 18)
Make incident management an operating discipline rather than an audit project. Keep policy, staffing, tools, and investment aligned with changing customer and system risk.
- Assign permanent owners for the incident policy, paging platform, status page, service catalog, training program, and metrics.
- Review severity thresholds, communication timings, compensation, and staffing at least annually and after material incidents.
- Use incident trends to prioritize architectural work on the shared ledger, regional independence, deployment safety, dependency isolation, and graceful degradation.
- Review repeat incidents and repeat contributing factors quarterly. Require executive action when remediation repeatedly loses priority.
- Survey responders quarterly and publish actions addressing fatigue, fairness, psychological safety, and tool friction.
- Recognize effective incident leadership, early declaration, useful postmortems, and preventive work.
- Prohibit retaliation for good-faith incident declaration or escalation.
- Provide the board or risk committee a quarterly summary of severe incidents, SLA exposure, regulatory events, overdue high-risk actions, and resilience investment.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
Previous Proposal 3 (ID: c051663c-fdf6-4324-bfbe-a5f42ed62eed, Agent: qwen3.8-max_initial_3, LLM: alibaba/qwen3.8-max):
Estimated Complexity: high
Success Metrics: - **MTTD reduced from 22 minutes to <5 minutes** within 6 months of full rollout.
- **Customer-first detection rate reduced from 40% to <10%** within 6 months.
- **MTTM for SEV1 incidents reduced from 3 h 10 min to <60 minutes** within 9 months.
- **Monthly SLA credit payouts reduced from ~$325K to <$100K per quarter** within 12 months.
- **Alert volume reduced from 3,400/month to <600 actionable alerts/month** within 6 months; signal-to-noise ratio >80%.
- **Postmortem completion rate: 100% of SEV1/SEV2 incidents** have a blameless postmortem within 5 business days.
- **Postmortem action-item completion rate >90% within 30 days** of the postmortem (up from ~17%).
- **Zero incidents with >15 minutes of unowned command** (down from 2 incidents with >1 hour).
- **100% on-call coverage**: all 28 teams staffed with primary + secondary on-call 24×7 within 14 weeks.
- **On-call compensation adopted**: 100% of on-call engineers receiving stipends and page pay; on-call satisfaction score ≥4/5 in quarterly survey.
- **Status-page first-update within 15 minutes for SEV1 and 30 minutes for SEV2**, 100% compliance.
- **SOC 2 Type II audit passed** at month 8 with zero incident-response findings.
- **All 260 engineers trained**; 56+ certified ICs and 28+ certified CLs active within 14 weeks.
- **Six alerting tools consolidated to one** within 6 months; legacy tools decommissioned.
- **Regulator notification process tested**: at least one tabletop exercise includes a NY DFS / FinCEN notification drill, and the legal/compliance playbook is documented and approved.
- **Quarterly IMC reviews held consistently** with published KPI dashboards and action-item tracking.
- **On-call participation resistance resolved**: <10% of engineers report 'unwilling to participate' in the 6-month pulse survey (baseline to be measured in S1).
Steps (13):
1. Assess Current State and Baseline Metrics
Build an evidence-based picture of the current incident management reality before designing anything new.
Collect and catalog the last 12 months of incident data: all 31 customer-impacting incidents, 3,400 monthly alerts, on-call coverage gaps across the 28 teams, and the 11 of 64 closed postmortem action items. Interview one lead from each of the 28 teams to surface pain points, political concerns (the 'carrying a pager for other teams' pushback), and tool sprawl.
Deliverables to produce:
- **Alert inventory**: which of the six alert tools feed which teams, alert volume per team, noise rate per tool, and overlap between tools.
- **Incident timeline analysis**: median detection-to-notification-to-mitigation-to-resolution times, who detected first (internal vs. customer), who mitigated, and where handoff gaps occurred.
- **On-call coverage map**: which 12 teams have on-call, which 16 do not, rotation length, compensation status, and escalation paths (or lack thereof).
- **Postmortem audit**: format variance, action-item tracking gaps, and the two incidents with no clear owner for over an hour.
- **Tooling and integration audit**: Kubernetes observability stack, six alerting tools, status-page provider, communication channels (Slack, email, phone), and any existing runbooks.
- **Compliance gap analysis**: SOC 2 Type II CC7.3/CC7.4 requirements vs. current practice, with a risk register for the eight-month window.
- **Peer benchmarking**: incident management practices at 3–4 comparable B2B fintech platforms (e.g., Plaid, Stripe, Adyen) for severity scales, on-call comp, and MTTR targets.
2. Secure Executive Sponsorship and Form the IM Governance Body (depends on: 1)
Anchor the program with visible, top-down authority so that 28 teams adopt changes they did not individually request.
The CEO email about 'outages we hear about from clients' is a ready-made mandate. Convert it into a formal sponsorship structure.
- Appoint an **executive sponsor** (CTO or VP Engineering) who owns the program end-to-end and reports to the CEO monthly.
- Create an **Incident Management Office (IMO)**: one dedicated senior incident management lead, one tooling/platform engineer, and one part-time data analyst.
- Establish an **Incident Management Council (IMC)**: one engineering manager from each of the 28 teams, plus the VP of Customer Success, a compliance lead, and a security lead. The IMC meets bi-weekly during rollout, monthly thereafter.
- Draft and circulate an **executive mandate memo** that states: incident response is a shared operational obligation, not a per-team favor; participation in on-call rotations is a condition of employment for production-facing roles; and the program is not optional pending the SOC 2 audit.
- Allocate a dedicated budget line for on-call compensation, tooling consolidation, status-page licensing, training, and external facilitation.
3. Define Severity Levels and Automatic Triggers (depends on: 1)
Replace the current ad-hoc triage with a five-level severity taxonomy that every engineer, support agent, and account manager can apply in under 60 seconds.
- **SEV1 – Critical**: Ledger data corruption or loss, complete payment processing halt, confirmed data breach affecting customer PII or funds, or regulatory reporting breach. Triggers: automatic all-hands page to on-call, CTO + CEO paged within 10 minutes, dedicated bridge call within 5 minutes, status-page update within 15 minutes, regulator notification assessment within 1 hour, customer comms within 30 minutes.
- **SEV2 – High**: Payment processing degraded >30% throughput or >5% error rate, single-region failover failure, ledger read-only mode, or any condition likely to breach the 99.95% SLA within the current window. Triggers: primary + secondary on-call paged, incident commander assigned within 10 minutes, bridge call within 15 minutes, status-page update within 30 minutes, VP Engineering notified within 20 minutes.
- **SEV3 – Medium**: Non-critical service degradation affecting <30% of customers, single non-ledger microservice outage with working fallback, or elevated latency above SLA threshold but below halt. Triggers: primary on-call paged, IM notified within 30 minutes, status-page update within 1 hour if customer-visible, daily standup update.
- **SEV4 – Low**: Degraded internal tooling, minor UX bug with workaround, non-customer-facing alert. Triggers: next-business-day response, ticket created, no page unless on-call agrees.
- **SEV5 – Informational / Noise**: Cosmetic issues, planned-maintenance notifications, alert misfires logged for tuning. Triggers: no page, logged for weekly alert-quality review.
Define **escalation rules**: any SEV3 unresolved after 4 hours auto-escalates to SEV2; any SEV2 unresolved after 2 hours auto-escalates to SEV1. Severity can be **downgraded** only by the incident commander with IMC notification.
Publish the taxonomy as a one-page decision tree, a Slack slash-command (`/sev`), and an integration into the alerting tool so that every alert carries a suggested severity.
4. Define Incident Roles and Staffing Model (depends on: 3)
Codify four mandatory roles for every SEV1/SEV2 incident and optional roles for SEV3, then solve the 24×7 staffing problem across 28 teams.
**Roles**
- **Incident Commander (IC)**: owns the incident end-to-end, declares severity, assigns tasks, authorizes mitigations, decides when to escalate or stand down. Never writes code during the incident.
- **Communications Lead (CL)**: owns status-page updates, internal Slack channels, account-manager briefings, and regulator notifications. Separate from the IC so the IC can focus on mitigation.
- **Scribe / Timeline Keeper**: logs every decision, action, and timestamp in the incident channel and the incident-management tool. Produces the raw timeline for the postmortem.
- **Subject-Matter Responders (SMRs)**: 1–3 engineers from the owning team(s) who diagnose and fix. For the shared PostgreSQL ledger, a dedicated DBA responder is always required.
**24×7 Staffing via a Three-Tier Follow-the-Sun Model**
- **Tier 1 – Front-line on-call**: Primary + secondary responder per team, paged first. Covers the team's own services.
- **Tier 2 – Platform / SRE on-call**: A dedicated 6-person SRE rotation covering cross-cutting infrastructure: Kubernetes, the shared PostgreSQL ledger, networking, and the two AWS regions. This tier directly addresses the 'carrying a pager for other teams' concern by absorbing infrastructure incidents.
- **Tier 3 – IMC escalation**: Engineering managers and the IMO on-call for multi-team or SEV1 incidents. Provides the IC and CL when no team-level IC is available.
**Follow-the-Sun**: If any engineering hub exists in a second timezone, use it for overnight Tier 1 coverage. If not, partner with a managed on-call service for overnight first-response triage (severity declaration + paging the correct team), reducing 3 a.m. pages for NY-based engineers.
**IC and CL pools**: Nominate at least 2 ICs and 1 CL per team (56 ICs, 28 CLs minimum). ICs are trained and certified before they rotate. For SEV1 incidents, the IC must be a certified IC from the IMC pool, not just 'whoever is around'.
**Ledger-specific rule**: Because the PostgreSQL ledger is shared, a **Ledger Duty Officer** from the SRE Tier 2 is always on the bridge for any incident touching ledger services, regardless of which team owns the failing microservice.
5. Design On-Call Rotations, Compensation, and Alert-Quality Rules (depends on: 4)
Make on-call sustainable, fairly compensated, and free of alert noise so engineers stop resisting participation.
**Rotation Design**
- 7-day rotations, one primary + one secondary per team per week. No engineer is on-call more than one week in four.
- Minimum 48-hour rest between rotations. No on-call during approved PTO.
- All 28 teams participate. Teams without current on-call get a 90-day ramp with a shadow rotation before going live.
- Tier 2 SRE rotation: 6 engineers, one week on / five weeks off, with a dedicated backup.
**Compensation Package**
- **Base on-call stipend**: $500 per week of primary on-call, $250 for secondary, paid regardless of whether pages fire.
- **Page pay**: $75 per acknowledged page outside business hours; $150 if the page leads to active incident work.
- **Time-off-in-lieu (TOIL)**: Any engineer who works >4 hours overnight (00:00–06:00 local) gets a full TOIL day. >8 hours in a single incident gets 1.5 TOIL days.
- **SEV1 bonus**: $300 flat bonus for every engineer who actively works a SEV1 incident, paid within the next pay cycle.
- **Annual on-call cap**: No engineer exceeds 13 weeks of on-call per year. Exceeding the cap triggers a mandatory team-staffing review.
- Budget estimate: ~$420K/year for stipends and page pay across 28 teams; present this to the CFO as a fraction of the $1.3M annual SLA credit cost.
**Alert-Quality Rules (the '85% noise' problem)**
- Every alert must carry: owning team, suggested severity, runbook link, and a 30-day noise score.
- **Alert budget**: each team gets a maximum of 100 actionable alerts per month. Exceeding the budget triggers a mandatory alert-tuning session with the IMO.
- **Noise threshold**: any alert that fires >10 times in 7 days with no human action is auto-flagged for suppression or tuning within 14 days.
- **Alert review cadence**: weekly 30-minute alert-quality review per team; monthly cross-team alert review in the IMC.
- **Sunset rule**: alerts with no runbook are demoted to SEV5 after 30 days and suppressed after 60 days unless a runbook is written.
- Target: reduce monthly alert volume from 3,400 to <600 actionable alerts within 6 months.
6. Build Detection, Escalation, and Communication Paths (depends on: 3, 4)
Eliminate the 22-minute median detection gap and the 40% customer-first-detection rate with layered monitoring and a single escalation spine.
**Detection Layers**
- **Synthetic transactions**: run a payment end-to-end through the full stack (API → service → ledger → confirmation) every 60 seconds from both AWS regions. Alert if latency >2× baseline or any step fails. This catches what per-service metrics miss.
- **Customer-traffic anomaly detection**: monitor API error rates, payment success rates, and latency percentiles per customer cohort. Alert on >2σ deviation.
- **SLO-based alerting**: define SLIs for the 99.95% SLA (availability, latency p99, ledger consistency). Alert when error budget burn rate exceeds threshold, before the SLA actually breaches.
- **Infrastructure health**: Kubernetes node/pod health, PostgreSQL replication lag, disk I/O, and cross-region latency.
- **Support-ticket spike detection**: if >5 customers open tickets about the same symptom within 10 minutes, auto-create a SEV3 candidate.
**Escalation Path**
- Alert fires → PagerDuty routes to Tier 1 primary → 5-min no-ack → Tier 1 secondary → 10-min no-ack → Tier 2 SRE → 15-min no-ack → IMC on-call manager → 20-min no-ack → VP Engineering auto-page.
- Any SEV1 declaration auto-pages the CTO, opens a dedicated Slack channel + Zoom bridge, and notifies the CL.
- **No incident goes unowned for >15 minutes.** If no IC is assigned by minute 15, the IMC on-call manager assumes IC role by default.
**Internal Communications**
- Dedicated Slack channels: `#inc-sev1`, `#inc-sev2`, `#inc-sev3` (auto-created per incident), plus `#inc-updates` for broadcast.
- IC posts a structured update every 15 minutes (SEV1), 30 minutes (SEV2), 1 hour (SEV3) into the incident channel.
- CL posts a summary to `#inc-updates` and notifies relevant engineering managers.
**Customer Communications**
- **Status page**: auto-updated via API. SEV1: first update within 15 minutes, then every 30 minutes until resolved. SEV2: first update within 30 minutes, then every hour. SEV3: within 1 hour if customer-visible.
- **Account managers**: CL briefs AMs via a dedicated Slack channel within 30 minutes (SEV1) or 1 hour (SEV2). AMs contact their top-20 revenue accounts directly.
- **Customer email/SMS**: for SEV1 and SEV2, automated notification to all 2,100 customers via the status-page subscription system within 30 minutes.
- **Regulator notification**: Legal/Compliance assesses within 1 hour whether NY DFS, FinCEN, or card-network notification is required. If yes, file within the regulatory deadline (typically 72 hours for NY DFS cybersecurity events). Log the decision and filing in the incident record.
**Post-resolution**: CL publishes a 'resolved' update within 30 minutes of mitigation. For SEV1/SEV2, a preliminary customer-facing RCA summary is published within 5 business days.
7. Standardize Postmortems with Tracking and Accountability (depends on: 3, 6)
Fix the '11 of 64 action items closed' problem with a mandatory, uniform, blameless postmortem process backed by engineering-manager accountability.
**When Mandatory**
- All SEV1 and SEV2 incidents: postmortem required within 5 business days.
- SEV3 incidents: postmortem required if >50 customers affected, if the incident lasted >4 hours, or if it was customer-detected.
- SEV4/SEV5: optional, but any recurring SEV4 (≥3 times in 30 days) triggers a mandatory review.
**Format (single template, enforced by tooling)**
- Incident summary (severity, duration, customers affected, revenue impact, SLA credit exposure).
- Timeline (auto-generated from scribe notes + alert timestamps).
- Detection analysis: how was it detected, why did it take X minutes, could it have been faster.
- Root cause analysis using 5-Whys or fault-tree, not blame.
- Contributing factors (process, tooling, staffing, knowledge gaps).
- Impact quantification: customers affected, transactions failed, SLA credits triggered.
- Action items: each with a **named owner**, **due date**, **priority**, and **ticket in Jira**.
- Lessons learned and what went well.
**Blameless Review Meeting**
- Held within 5 business days, facilitated by the IMO or a trained facilitator (never the IC of that incident).
- All responders, the CL, relevant engineering managers, and an IMC representative attend.
- Ground rules: focus on system and process failures, not individual mistakes. The facilitator enforces this.
- Meeting recorded; notes published to the engineering-wide wiki within 48 hours.
**Action-Item Tracking and Accountability**
- Every action item is created as a Jira ticket with a due date and a named owner.
- **Engineering managers are accountable**: action-item completion is a standing agenda item in the bi-weekly IMC meeting. Any action item >7 days overdue is escalated to the VP Engineering.
- **Completion gate**: no team may close its postmortem until 100% of its action items have Jira tickets. Postmortem is 'closed' only when all tickets are resolved.
- **Quarterly audit**: the IMO audits action-item completion rates and reports to the IMC and the executive sponsor. Target: >90% completion within 30 days of the postmortem.
- Link postmortem quality and action-item completion to team health scores and engineering-manager performance reviews.
8. Define KPIs, Dashboards, and Governance Reviews (depends on: 7)
Create a measurable feedback loop so leadership can see whether the program is working and where to intervene.
**Primary KPIs (tracked weekly, reported monthly)**
- **MTTD** (Median Time to Detect): target <5 minutes (from 22).
- **MTTA** (Median Time to Acknowledge): target <5 minutes.
- **MTTM** (Median Time to Mitigate): target <60 minutes for SEV1 (from 3 h 10 min), <4 hours for SEV2.
- **Customer-first detection rate**: target <10% (from 40%).
- **SLA compliance**: maintain 99.95%; track monthly SLA credit payouts, target <$100K/quarter (from ~$325K/quarter).
- **Alert signal-to-noise ratio**: target >80% actionable (from ~15%).
- **Alert volume**: target <600/month (from 3,400).
- **Postmortem completion rate**: 100% for SEV1/SEV2 within 5 business days.
- **Action-item completion rate**: >90% within 30 days (from ~17%).
- **On-call health**: pages per engineer per week (target <5), TOIL usage, on-call satisfaction survey score.
- **Unowned incident duration**: target 0 incidents with >15 minutes without an IC.
**Dashboards**
- Real-time operational dashboard (Grafana): current incidents, active alerts, on-call roster, SLA error-budget burn.
- Weekly leadership dashboard (auto-generated): KPI trends, open action items, alert-noise report, on-call load distribution.
- Quarterly IMC scorecard per team.
**Review Cadence**
- **Weekly**: IMO publishes KPI snapshot to `#inc-updates`.
- **Bi-weekly IMC**: review open incidents, overdue action items, alert-quality exceptions, and on-call load.
- **Monthly executive review**: CTO presents KPI trends, SLA credit cost, and risk register to the CEO.
- **Quarterly incident-management review**: deep-dive into trends, training gaps, tooling needs, and process improvements. Output fed into the next quarter's roadmap.
9. Consolidate Tooling and Build the Incident Management Platform (depends on: 2, 3)
Replace six alerting tools and ad-hoc status-page updates with a single, integrated incident management stack.
**Target Tool Architecture**
- **Single alerting and on-call platform** (e.g., PagerDuty or Opsgenie): ingest all alerts, apply severity routing, manage on-call schedules, handle escalations, and send pages. Retire the other five tools within 6 months.
- **Observability consolidation**: standardize on one APM/metrics stack (e.g., Datadog or Grafana Cloud) for all 180 Kubernetes services across both AWS regions. Ensure the shared PostgreSQL ledger has dedicated dashboards.
- **Status page**: a dedicated, branded status page (e.g., Statuspage.io or Instatus) with API integration for auto-updates. Subscribe all 2,100 customers.
- **Incident coordination tool**: integrate incident-management workflows into Slack (auto-create channels, invite responders, post templates) and a dedicated incident record system (e.g., Jira Service Management, incident.io, or Rootly) for timelines, postmortems, and action-item tracking.
- **Runbook repository**: a central wiki (Confluence or Notion) with mandatory runbooks for every alert. No alert goes live without a linked runbook.
**Implementation Tasks**
- Migrate all 28 teams' alert rules into the single platform in three waves (highest-volume teams first).
- Build the severity-based routing rules and escalation policies per S3 and S6.
- Automate status-page updates triggered by severity declaration.
- Build the synthetic-transaction monitor and SLO-based alerting per S6.
- Integrate Jira for automatic action-item ticket creation from postmortems.
- Decommission legacy tools only after all teams have completed training on the new stack.
- Budget: allocate $150K–$250K/year for licensing, plus engineering time for migration.
10. Prepare for the SOC 2 Type II Audit (depends on: 7, 8, 9)
Ensure the incident management process produces the evidence the auditor will need, well before the audit window opens in eight months.
**SOC 2 Requirements to Address (CC7.3, CC7.4, CC7.5)**
- Documented incident response procedures (the severity taxonomy, role definitions, communication templates).
- Evidence of incident detection, response, and recovery for every SEV1/SEV2 incident during the audit period.
- Postmortem records with action-item tracking.
- On-call schedules, training records, and escalation evidence.
- Status-page update logs and customer notification records.
- Regulator notification logs (if any).
**Preparation Tasks**
- The IMO maintains a **SOC 2 evidence folder**: every incident record, postmortem, action-item ticket, status-page update, and training completion certificate is stored and indexed.
- Conduct a **mock SOC 2 audit** at month 5: an internal or external auditor reviews the incident management process end-to-end and identifies gaps.
- Remediate mock-audit findings before month 7.
- Ensure the incident management tool retains all records for at least 12 months (the SOC 2 Type II observation window).
- Document the **chain of custody** for incident records: who accessed, modified, or closed each record.
- Prepare a **narrative document** describing the incident management process, roles, and controls for the auditor.
- Coordinate with the compliance lead to align incident management evidence with the broader SOC 2 scope (access controls, change management, etc.).
11. Design and Deliver Training, Runbooks, and Change Management (depends on: 4, 5, 9)
Equip all 260 engineers, 28 team leads, account managers, and support staff with the knowledge and muscle memory to execute the new process.
**Training Tracks**
- **All 260 engineers** (2-hour session): severity taxonomy, how to acknowledge a page, how to join an incident bridge, how to hand off to an IC, and how to write a postmortem contribution. Delivered in team-level sessions over 4 weeks.
- **IC pool (56+ engineers)** (8-hour certification): incident command techniques, severity declaration, escalation decision-making, bridge facilitation, and blameless postmortem facilitation. Includes two tabletop exercises. Certification valid for 12 months, renewed annually.
- **CL pool (28+ staff)** (4-hour session): status-page writing, customer communication templates, regulator notification triggers, and AM briefing protocol.
- **Account managers and support staff** (1-hour session): how to read the status page, how to escalate a customer report into an incident, and what information to collect.
- **SRE Tier 2** (16-hour onboarding): Kubernetes and PostgreSQL ledger deep-dive, cross-region failover runbooks, and escalation authority.
**Runbooks**
- Every alert must have a runbook before it is routed to on-call. The IMO provides a runbook template and audits compliance weekly.
- Priority runbooks to write first: shared PostgreSQL ledger failover, Kubernetes cluster degradation, payment-processing pipeline failure, cross-region failover, and ledger data-integrity check.
- Runbooks are peer-reviewed and version-controlled.
**Change Management for Adoption**
- Address the 'carrying a pager for other teams' concern directly: publish an FAQ explaining the three-tier model, the SRE Tier 2 absorbing cross-team infrastructure, the compensation package, and the TOIL policy.
- Run **office hours** weekly for the first 8 weeks where any engineer can ask questions or raise concerns.
- Identify **team champions**: one engineer per team who volunteers as an early adopter and peer mentor.
- Publish a **weekly 'incident management newsletter'** during rollout: what changed, what improved, KPI trends, and success stories.
- Make on-call participation a documented expectation in job descriptions and performance reviews for production-facing roles.
12. Execute Phased Rollout, Tabletop Exercises, and Continuous Improvement (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11)
Introduce the process in three waves so teams are not overwhelmed, then validate with exercises and iterate continuously.
**Phase 1 – Weeks 1–6: Foundation**
- Publish the severity taxonomy, role definitions, and communication protocols (S3, S4, S6).
- Launch the single alerting platform for the 12 teams already on-call; begin migration for the other 16.
- Activate the SRE Tier 2 rotation for the shared PostgreSQL ledger and cross-cutting infrastructure.
- Deploy the status page and test the API integration.
- Begin IC and CL training (first cohort of 20 ICs, 10 CLs).
- Publish the on-call compensation package; HR integrates stipends into payroll.
- Write the top 10 priority runbooks.
**Phase 2 – Weeks 7–14: Expansion**
- All 28 teams live on the single alerting platform; legacy tools in read-only mode.
- All 28 teams on the on-call rotation schedule (the 16 new teams in shadow mode for the first 4 weeks).
- Second and third IC/CL training cohorts completed.
- First **tabletop exercise**: simulate a SEV1 ledger corruption scenario with all roles, test the escalation path, status-page updates, and AM briefings. Debrief and fix gaps.
- Postmortem template and Jira integration live; all new incidents use the standard process.
- Alert-tuning sprint: each team reduces its alert volume by 50%.
**Phase 3 – Weeks 15–24: Optimization**
- All teams fully live; legacy alerting tools decommissioned.
- Second **tabletop exercise**: simulate a SEV1 cross-region failure with regulator notification.
- First quarterly IMC review with full KPI dashboard.
- Mock SOC 2 audit (month 5) and remediation.
- Retrospective on the rollout: survey all 260 engineers for feedback, adjust compensation or rotation rules if needed.
- Establish the **continuous improvement cadence**: quarterly process review, annual severity-taxonomy review, and annual on-call compensation benchmarking.
**Ongoing Governance**
- The IMC owns the process document and approves changes.
- The IMO tracks all KPIs and reports to the CTO monthly.
- Any process change requires IMC approval and a 2-week notice period before enforcement.
- Annual external benchmarking against peer B2B payments platforms.
13. Establish Ongoing Governance, Annual Review, and Audit Readiness Cycle (depends on: 12)
Embed incident management as a permanent organizational capability, not a one-time project.
- **Annual process review**: the IMC reviews the severity taxonomy, role definitions, on-call structure, and compensation against industry benchmarks and internal KPIs. Update as needed.
- **Bi-annual tabletop exercises**: one SEV1 infrastructure scenario, one SEV1 data-breach/regulator scenario. Rotate the IC and CL assignments so everyone gets practice.
- **Quarterly alert-quality audit**: the IMO reviews alert volumes, noise rates, and runbook coverage across all 28 teams.
- **On-call health survey**: quarterly anonymous survey measuring burnout, fairness, and compensation satisfaction. Results reviewed by the IMC.
- **SOC 2 readiness cycle**: begin evidence collection immediately after each audit ends. The IMO maintains a rolling evidence folder. Mock audit at month 5 of every 12-month audit cycle.
- **Postmortem maturity tracking**: track the action-item completion rate monthly. If it drops below 80%, the VP Engineering intervenes.
- **Incident management maturity model**: adopt a 5-level maturity model (ad-hoc → defined → managed → optimized → predictive). Assess annually. Target: Level 3 within 12 months, Level 4 within 24 months.
- **Budget review**: annually review on-call compensation, tooling costs, and training budget against the reduction in SLA credits and incident frequency.
Previous Proposal 4 (ID: 31fbbc6c-8c88-4abb-9c6d-b67706d9da75, Agent: grok4.6_initial_4, LLM: xai/grok-4.6):
Estimated Complexity: high
Success Metrics: - Median time to detect customer-impacting incidents ≤ 5 minutes within 6 months of go-live.
- Share of SEV-1/SEV-2 incidents first detected by customers ≤ 5% (from 40%).
- Median time to mitigate SEV-1/SEV-2 ≤ 45 minutes (from 3 h 10 min).
- Named Incident Commander assigned within 5 minutes for ≥ 95% of SEV-1/SEV-2.
- First status-page update within policy time for ≥ 95% of SEV-1/SEV-2.
- SLA credits down ≥ 80% versus the trailing $1.3M within 12 months.
- Paging volume ≤ 500 per month and noise ≤ 15% (from 3,400 and 85%).
- 100% of production services have a named owning team and a paging policy.
- Postmortems filed within 5 business days for 100% of SEV-1/SEV-2; action-item close rate ≥ 80% within 30 days.
- 24×7 IC and critical-path coverage with zero unfilled shifts per quarter.
- Paid on-call live for every rotation before that rotation pages humans.
- SOC 2 Type II incident-response controls evidenced for ≥ 5 months before the auditor's report.
- On-call pulse: ≥ 70% of engineers agree rotations are fair and limited to their services.
Steps (30):
1. Secure executive mandate and budget
Get a written CEO/CTO mandate that incident command is a company process, not a team hobby.
The mandate must state that **paid on-call** is required for production ownership. "No pager for other teams' code" is solved by named ownership, not by refusing coverage.
- Approve budget for tooling, stipends, training, and a dedicated program lead for six months.
- Name an executive sponsor (CTO or VP Engineering) who will chair the weekly incident review.
- Tie the clock to SOC 2 Type II: the process must be live in about 10 weeks so ~6 months of evidence remain.
- Commit that the CEO will hear about outages from this process, not from customers.
2. Form the working group and decision rights (depends on: 1)
Stand up a small group that can decide. Do not form a 28-team committee.
**Core seats:** SRE/platform lead, payments/ledger engineering manager, support lead, legal/compliance, HR, one rotating team EM, and a program manager.
- Meet twice a week for 10 weeks, then weekly.
- RACI: the group proposes; the sponsor decides in 48 hours; teams implement.
- Publish one Slack channel and one source-of-truth doc on day one.
- Time-box design to four weeks. Ship v1 rather than wait for consensus.
3. Inventory services, owners, and on-call gaps (depends on: 1)
Build a living catalog of all ~180 services: owning team, criticality, current on-call, alert sources, and runbook link.
Walk the last 31 customer-impacting incidents and the two events where **nobody was in charge**. Record who detected, who led, time to mitigate, and which alerts fired.
- Tag each service as critical-path, customer-visible, or internal.
- List the 16 teams with no on-call and every orphan service with no owner.
- Map all six alerting tools and the 3,400 monthly alerts onto services.
- Flag the shared PostgreSQL ledger and two-region failover as named special cases.
4. Map regulatory and contractual notification duties (depends on: 1)
Legal and compliance list every duty an incident can trigger. Do not invent clocks that violate a contract.
Cover SOC 2 CC7, customer MSA/SLA credit terms, money-transmitter rules, NYDFS 23 NYCRR 500 if applicable, PCI if in scope, and breach clocks.
- Extract **notification timings** from the largest customer contracts (status page, named AM, written notice).
- Define when legal, regulators, insurers, or the board must be told.
- Feed these clocks into severity triggers and the communications playbook.
5. Approve paid on-call and incident pay (depends on: 1, 4)
Unpaid on-call is why 16 teams refuse the pager and why nights are uncovered. Fix the money before asking for coverage.
HR, legal, and finance design a New York–compliant package: weekly stipend for primary and secondary, extra stipend for company IC and comms, and **after-hours incident pay** or comp time.
- Treat exempt vs non-exempt staff explicitly under NY wage-hour rules.
- Put stipend in the next pay cycle after policy publish, not "later."
- Cap consecutive night weeks. Fund hiring if a team cannot rotate fairly (minimum six people for 24x7 primary plus secondary).
- Publish the package before any new rotation starts. This is the main answer to pager pushback.
6. Ratify four severity levels and their triggers (depends on: 3, 4)
Adopt a business-impact scale. Engineers do not invent severity in the moment.
**SEV-1:** material payments failure; ledger down or inconsistent; security or customer-data incident; both regions impaired; or many customers already in SLA-credit territory.
**SEV-2:** degraded payments or a contracted feature down for multiple customers; SLA at risk.
**SEV-3:** narrow or single-customer impact with a workaround; no fleet-wide SLA risk.
**SEV-4:** no customer impact; ticket only.
- SEV-1 pages IC, comms, scribe, owning SMEs, and an exec; war room in 5 minutes; status page in 10; AM outreach in 20.
- SEV-2 pages IC and owning SMEs; comms may be the IC; status page in 15 minutes; updates every 30 minutes.
- SEV-3 pages the owning team only; customer notice only if that customer is affected.
- Anyone may declare. Only the IC may downgrade. When unsure, start high.
7. Define incident roles and their authority (depends on: 6)
Four roles. Separate coordination from debugging so "who is in charge" cannot stall for an hour again.
- **Incident Commander:** owns severity, the room, the clock, and the next action. Does not write code. May page anyone, freeze deploys, and invoke failover. Staffed from a company-wide trained pool, not from the failing team.
- **Communications Lead:** status page, customers, AMs, execs, regulators. Speaks only from IC-approved facts.
- **Scribe:** timeline in the incident tool. Required for SEV-1 and SEV-2.
- **SME responders:** the owning team's on-call. They mitigate. They do not run the room.
Publish a one-page authority card. The IC stays in charge even if a VP joins.
8. Design 24x7 coverage without 28 night rotations (depends on: 3, 5, 7)
Do not put 28 teams on 24x7. That is what engineers are rejecting.
Use **three layers** so people page for their own code, plus a trained commander.
- Layer A — company IC and SEV-1 comms: 24x7; about 24 trained people; week-long primary and secondary.
- Layer B — critical-path team on-call (ledger, payments processing, auth, API edge, platform/Kubernetes, data stores): 24x7 primary plus secondary.
- Layer C — all other teams: business-hours on-call; after hours the IC pages the team EM, who has a written escalation list.
Platform on-call is the safety net for unknown-owner pages, never the permanent owner. Every service must have a named team within 60 days or be scheduled to shut off.
9. Set rotation, handoff, and load rules (depends on: 8)
Write mechanical rules so rotations are fair and load is visible.
Primary week, then secondary week, then at least two weeks off. No one holds primary on two rotations at once.
- Handoff is a 30-minute overlap covering open incidents, silenced alerts, and upcoming changes.
- Page-load SLO: p50 ≤ 4 pages per 12-hour night shift; p95 ≤ 10. A breach opens an alert-quality action.
- Require a shadow week before a first IC shift or a first critical-path rotation.
- Swaps live in the paging tool. Managers own coverage gaps, not the last person on the roster.
10. Write detection and escalation paths (depends on: 6, 9)
Customers currently detect 40% of incidents and median time to detect is 22 minutes. That is the first failure mode.
Detection path: synthetic full-payment probes in both regions, SLO burn-rate alerts, support-to-incident intake, and one customer callback path that can create a SEV.
- A page must be acked in **5 minutes** or it auto-escalates to secondary, then IC, then the EM, then the VP.
- Support may declare SEV-2 or higher without engineering permission.
- If ownership is unclear for 10 minutes, the IC keeps the incident and assigns a temporary owner. Never wait.
- An exec bridge auto-opens for every SEV-1 at T+15 minutes.
11. Write internal, customer, and regulator communications (depends on: 4, 6, 7)
Stop "whoever is around" from writing the status page. Comms follow the clock, not convenience.
Timings from declaration:
- Internal war room: immediate. Exec summary for SEV-1/2 at 15 minutes, then every 30 minutes.
- **Public status page:** SEV-1 in 10 minutes, SEV-2 in 15. Updates at least every 30 minutes until resolve. Templates only. No speculation.
- Account managers get an affected-customer list and a script at T+20 minutes for SEV-1/2.
- Resolve notice and credit assessment within one business day.
- Comms pages legal on SEV-1 security, ledger integrity, or any outage that will breach contractual notice. Legal owns outbound regulatory letters; the IC owns facts.
12. Standardize blameless postmortems and action tracking (depends on: 6)
A written postmortem is mandatory for every SEV-1 and SEV-2 within 5 business days. SEV-3 if the IC or EM requests it.
Use one template: timeline, customer impact (volume, duration, credits), detection gap, what went well, what did not, process-focused five whys, and numbered actions with owner and due date.
- Review is **blameless** and scheduled. The IC attends. The exec sponsor reads every SEV-1.
- Actions live in one tracker, not in the doc. No action without an owner and a date. Default due date 14 days; 30 days max unless architecture work with a milestone.
- Close rate is a published metric. The old 11-of-64 pattern is a process failure.
13. Set alert quality rules that make paging acceptable (depends on: 3, 6)
3,400 alerts a month and 85% noise is why on-call feels like punishment. Pages are a product with a quality bar.
A page (not a ticket) must map to a customer-facing SLO or a hard dependency of one. It must have an owner team, a runbook link, and a default severity. It must be actionable at 3 a.m. by the person who is paged.
- Ban parallel paging from six tools. One paging policy: symptom-based; burn-rate preferred over raw thresholds.
- Every team gets a monthly noise budget. Exceeding it is a sprint task, not heroics.
- A human may silence a flapping alert only with a linked ticket.
14. Publish Incident Management Policy v1 (depends on: 5, 6, 7, 8, 9, 10, 11, 12, 13)
Collapse the design into a short policy people will open during an outage.
Ten pages or fewer, plus one-page cards for severity, roles, and comms timings. Host it where the incident tool can link it.
- Include the compensation summary and the rule: you are **not on-call for other teams' services**.
- Version it. v1 is mandatory from the pilot start date.
- Legal, HR, and the exec sponsor sign. Announce in all-hands, not only in Slack.
15. Implement a single incident command tool (depends on: 7, 10, 11)
Put one tool in the path that creates the room, pages roles from severity, records the timeline, and prompts status-page updates.
Requirements: Slack (or equivalent) incident bot, severity in one click, role assignment, stakeholder groups, and timeline export for postmortems and auditors.
- Integrate with the pager so IC, comms, and SME pages are automatic.
- Retain artifacts at least one year for SOC 2.
- Ad-hoc Zoom/Slack threads are no longer the system of record.
16. Consolidate six alerting tools onto one pager (depends on: 9, 13)
Pick one paging product. Connect existing monitors to it. Migrate **pages** first, tickets second.
- Inventory every page-producing rule. Delete or downgrade the noisy majority in S23.
- Route by service label → owning team schedule → the escalation policy from S10.
- IC and comms schedules live in the same product.
- Set a hard date after which pages outside the chosen tool are not valid on-call obligations.
17. Operationalize the status page and AM path (depends on: 11, 15)
Put the status page behind the Comms Lead role. Use templates for investigating, identified, mitigating, and resolved.
Subscribe AMs and customers whose contracts require it. Generate the affected-customer list from the incident (tenant, region, payment method).
- Dry-run a SEV-2 update before the pilot goes live.
- Record every public update in the incident timeline for the audit.
- Define partial vs full-outage wording so impact cannot be understated.
18. Stand up one postmortem repo and action board (depends on: 12)
Create the template, the filing location, and one Jira/Linear board with states: open, in progress, blocked, done, and won't-do (with exec reason).
Wire the incident tool so a SEV-1/2 automatically opens a draft postmortem and action tickets.
- Each week the program manager reports actions past due to the exec sponsor.
- If it is not in the board, it does not exist.
19. Assign owners and write critical-path runbooks (depends on: 3, 9)
Close the ownership gaps that cause "pager for other people's code."
Every production service gets a team in the catalog. Unowned services get an owner in 30 days or a decommission date.
- Write runbooks for ledger Postgres, regional failover, payments API, auth, and the Kubernetes control plane: symptoms, dashboards, mitigate vs escalate, customer impact.
- Link runbooks from alerts. If there is no runbook, the alert cannot page at night unless the EM accepts the gap in writing.
20. Detect payments failures before customers do (depends on: 3, 13)
Build or finish end-to-end synthetics: create payment, ledger write, webhook, both regions, both critical payment methods.
Alert on SLO burn, not on a single 500. Page SEV-2 or SEV-1 from these probes. This is the fastest lever on the 22-minute MTTD and the 40% customer-detected rate.
- Add ledger lag, replication, disk, and failover-readiness as first-class pages to the ledger team.
- Every postmortem asks first: **why did a customer see this first?**
21. Train the first cadre of ICs, comms leads, and scribes (depends on: 7, 14, 15)
Train about 24 ICs and 12 comms leads before the pilot. Classroom plus a recorded shadow of a simulated SEV-1.
Curriculum: severity, authority card, tool, comms timings, when to call legal, how to run a room of 20, how to hand off at 2 a.m.
- Certification: pass a tabletop. No certificate, no rotation.
- Recertify yearly and after any SEV-1 where process failed.
- Managers of ICs protect calendar time. This is part of the job.
22. Pilot on the payments critical path for six weeks (depends on: 5, 14, 15, 16, 17, 18, 19, 21)
Go live with policy, tool, paid rotations, and IC coverage for ledger, payments, API, platform, and support intake.
Keep old paths as backup for one week, then cut over. Real incidents use the new process only.
- Staff the program lead in every SEV-2+ as coach, not as secret IC.
- Collect friction daily. Fix tooling and wording in 48 hours.
- Expansion gate: named IC in under 5 minutes, first status update on time, no unpaid pages, postmortem filed.
23. Cut alert noise with a forced burn-down (depends on: 13, 16)
Give every team a numbered list of their noisiest alerts. Move noise from 85% to **under 15%**, and monthly pages from 3,400 toward 500.
Each sprint, critical-path teams must delete, debounce, or convert to ticket a fixed quota. Platform provides burn-rate and grouping libraries.
- Publish a weekly noise leaderboard. Shame systems, not people.
- After eight weeks, any page without a runbook or with >30% false pages in 14 days is auto-downgraded until fixed.
24. Roll remaining teams onto the model by risk (depends on: 22)
After the pilot gate, add teams in waves of four to five every two weeks. Highest customer-impact first.
Layer C teams get business-hours schedules and the EM night list. Do not surprise anyone with a pager.
- Each wave: ownership confirmed, alerts routed, runbooks for paging alerts, paid rotation in HR, one tabletop.
- Finish all 28 teams at least five months before the SOC 2 report date so the observation window covers the company.
- Orphan services still unowned at wave end are escalated to the sponsor for shutdown or reassignment.
25. Run tabletops and multi-region game days (depends on: 21, 22)
Schedule a monthly tabletop: SEV-1 ledger, SEV-1 region loss, SEV-2 degraded payments, customer-detected incident, and a "who is in charge" chaos drill.
Quarterly game day: fail a region or a ledger replica in staging or a controlled production drill.
- Include support, AMs, legal, and an exec. Process fails if only engineers show up.
- Capture actions on the same board as real postmortems.
- Use results as SOC 2 evidence that IR is tested.
26. Resolve ownership fights and pager culture (depends on: 5, 8, 22)
Treat pushback as design input, not defiance. Repeat the contract in office hours: you carry a pager for **your** services; IC is coordination; nights are paid; noise is a defect.
- EMs who cannot staff a fair rotation get headcount or have services reassigned. Do not run two-person 24x7.
- Publicly close the two historical "nobody in charge" incidents with what would be different now.
- Pulse-survey on-call at 60 and 120 days. If load or fairness is red, stop expansion until fixed.
27. Launch metrics, weekly review, and error budgets (depends on: 15, 22)
The CEO email exists because there was no operating rhythm. Stand up a dashboard and use it.
Track incident count by SEV, MTTD, MTTA, MTTM, customer-detected %, time-to-IC, time-to-status, pages per shift, noise %, SLA credits, postmortem SLA, and action close rate.
- **Weekly 30-minute review** chaired by the exec sponsor. Last week's incidents, open actions, alert SLO breaches. Mandatory for EMs of teams that had SEV-2+.
- Error budgets on critical SLOs: burn too fast and the team stops features and pays down reliability.
- Monthly report to the CEO so customer-first news does not recur.
28. Collect SOC 2 Type II evidence as a side effect (depends on: 4, 14, 22, 27)
Do not create a parallel "audit process." The real process is the evidence.
From pilot start, retain policy versions, on-call schedules, incident records, status-page history, postmortems, action tickets, training attendance, tabletop notes, and access reviews of the incident and pager tools.
- Map each artifact to SOC 2 CC7.2–CC7.5. Compliance owns the matrix; engineering owns completeness.
- Internal audit or a consultant runs readiness reviews at week 12 and week 20, not the week before the auditor.
29. Inspect and adapt the process after 90 days (depends on: 22, 27)
After 90 days live, revise v1 using data, not opinions.
Look at MTTD/MTTM, uncovered shifts, IC burnout, missed status updates, action close rate, and survey results.
- Change severity definitions if teams inflate or deflate them.
- Adjust Layer B vs C membership from real pages.
- Cut steps nobody follows. Add only what the last 90 days proved missing.
- Freeze v2 as the SOC 2 described process unless a defect is found.
30. Rehearse the audit and freeze the observation window (depends on: 24, 28, 29)
Six to eight weeks before the auditor, run a mock interview with an IC, a random engineer, support, and compliance.
Show one SEV-1 and one SEV-2 end-to-end: detect, page, command, comms, postmortem, closed actions.
- Fix evidence gaps immediately.
- Freeze process wording for the rest of the observation window; log exceptions.
- Brief the CEO and customer success with the metrics so the "outages we hear about from clients" story is retired.
Previous Proposal 5 (ID: 53260d57-1d36-433f-83b4-7205e06869b2, Agent: deepseek-v4-pro_initial_5, LLM: deepseek/deepseek-v4-pro):
Estimated Complexity: high
Success Metrics: - Customer-detected incidents decrease from 40% to less than 15% within six months.
- Median time to detect (MTTD) is under 5 minutes for SEV1 and SEV2 incidents.
- Median time to mitigate (MTTM) is under 60 minutes for SEV1 and under 2 hours for SEV2.
- Alert noise decreases from 85% to below 10% within three months.
- 100% of SEV1 and SEV2 incidents have a completed blameless postmortem within 5 business days.
- 100% of postmortem action items are tracked with an owner and due date; 90% are completed on time.
- 24x7 on-call coverage achieved across all 28 teams with no unpaid on-call.
- 99% of pages are acknowledged within 5 minutes.
- SLA credits paid reduce by at least 50% over the next 12 months.
- SOC 2 readiness: all incident response controls are documented, tested, and evidence is produced by month 7.
Steps (14):
1. Baseline current incident response and align stakeholders
Collect data from the last 12 months of incidents, all six alert tools, on-call practices, and team interviews. Identify gaps against the target incident management process and secure executive sponsorship.
- Gather incident timeline, detection source, mitigation time, customer impact, SLA credit, and postmortem status for all 31 incidents.
- Survey 28 teams on on-call burden, alert quality, and operational pain.
- Map current tools, escalation paths, and communication workflows.
- Create baseline metrics and a stakeholder map with executive sponsor and audit owner.
2. Define severity levels and response triggers (depends on: 1)
Define a four-level severity scale with objective business-impact criteria so any engineer can classify an incident consistently.
- SEV1: widespread transaction processing outage, data breach, security incident, or severe SLA breach; triggers full incident command, executive notification, and a 5-minute status page update.
- SEV2: major feature outage, significant degradation without workaround, or customer financial risk; triggers incident commander, full communications role, and status page updates.
- SEV3: partial impairment with workaround or limited customer impact; triggers on-call response, internal communication, and optional status page update.
- SEV4: minor or internal issue, no customer impact; handled during business hours through ticketing.
- Include an escalation matrix showing who can declare, downgrade, and invoke regulatory or legal involvement.
3. Define incident roles and decision authority (depends on: 1, 2)
Define incident roles, responsibilities, and decision authority using RACI to remove ambiguity about who is in charge.
- Incident Commander: owns the incident, declares severity, and coordinates resolution.
- Communications Lead: owns internal and external messaging, status page updates, and account manager notifications.
- Scribe: maintains timeline, incident log, and postmortem notes.
- Subject-matter responders: diagnose and fix the incident; may come from multiple teams.
- Executive sponsor: optional for SEV1; customer liaison: handles account managers.
- Define decision rights for severity declaration, escalation, rollback, customer communications, and incident closure.
4. Design 24x7 staffing model across 28 teams (depends on: 3)
Design 24x7 coverage across 28 teams without overloading engineers. Use service-based on-call plus a central incident command pool.
- Each service or domain team assigns primary and secondary on-call for its own services.
- Create central incident commander, communications, and scribe rotations staffed from a trained incident response guild across all teams; use follow-the-sun between the two AWS regions and time zones.
- Define escalation layers: service on-call to team lead or manager to service owner to executive.
- Define handoff times, shadow shifts, and load balancing; target at most one week of on-call per engineer per month.
- Bridge the current 12-team paid on-call to 28-team paid coverage; no team remains uncovered.
5. Define on-call rotations, compensation, and alert quality rules (depends on: 4)
Define sustainable rotations, pay, and rules that eliminate noisy pages.
- Rotations: weekly or biweekly, at least one primary and one secondary, with 12-hour shifts where possible or 24-hour for low-volume services.
- Compensation: monthly on-call stipend for all on-call engineers, additional incident response bonus for after-hours work, and time off in lieu; align with market rates.
- Alert quality rules: every page must be actionable, have a runbook link, specify a service owner, include severity, and be based on SLO burn or known failure signals; no dashboard-only alerts.
- Noise budget: reject or downgrade non-actionable alerts; all pages must go to on-call only after suppression and deduplication.
- Weekly alert review removes the top noisy alerts.
6. Design detection, escalation, and alert routing (depends on: 2, 4, 5)
Define how incidents are detected, routed, and escalated so nothing waits on a human to notice.
- Consolidate the six alert tools into one alerting and paging platform with routing by service, severity, and tags.
- Detection sources: infrastructure metrics, application synthetic transactions, log-based anomalies, business transaction SLI monitoring, and customer-reported issues through support or account managers.
- Routing: alert is paged to service on-call within 30 seconds; primary must acknowledge within 5 minutes; if no ack, page secondary then on-call manager.
- Escalation timeouts: unresolved SEV1 escalates to service owner at 15 minutes and to leadership at 30 minutes; any engineer can escalate to the incident commander.
- Define customer-reported incident intake and classification in the same tool.
7. Define internal and external communication protocols (depends on: 2, 3)
Define communication channels, templates, and timing for internal, customer, and regulator audiences.
- Internal: dedicated incident Slack channel, internal status page mirror, and war room bridge for SEV1; incident commander and communications lead own these channels.
- Status page: SEV1 post within 5 minutes, updates every 30 minutes or on material change, resolution within 60 minutes of mitigation; SEV2 post within 15 minutes, updates hourly; SEV3 optional.
- Account managers: SEV1 and SEV2 notify account managers within 15 minutes with an approved customer-facing description and expected impact.
- Regulators: legal or compliance determines notification for data breaches, security incidents, funds availability issues, or regulatory reportable events; criteria and timing follow legal and regulatory requirements; communications lead coordinates.
- Use pre-approved message templates and an approval chain; no ad-hoc wording.
8. Define postmortem policy and action tracking (depends on: 3)
Define mandatory blameless postmortems and action tracking.
- Mandatory for all SEV1 and SEV2 incidents, and any SEV3 that breaches SLA or is customer-detected.
- Format: impact, timeline, root causes, contributing factors, detection and response gaps, what worked well, and action items.
- Blameless: focus on system and process causes, not individual blame; use trained facilitators.
- Ownership: each action has an owner, due date, and tracking ID in a single backlog.
- Review postmortems at the weekly incident review; track action closure; expect 100% completion.
- Complete postmortems within 5 business days for SEV1 and SEV2 incidents.
9. Define metrics, dashboards, and review cadence (depends on: 2, 3, 8)
Define metrics and review cadence to measure process health.
- Metrics: MTTD, MTTM, customer detected percentage, alert noise percentage, on-call response time, on-call load, SLA credits paid, and postmortem action completion.
- Dashboards: real-time operational dashboard for on-call engineers and management.
- Weekly incident review: review all SEV1 and SEV2 incidents, action items, and noisy alerts.
- Monthly trends with leadership; quarterly review against SLOs and audit controls.
- Success thresholds: MTTD under 5 minutes, MTTM under 60 minutes for SEV1, customer detected under 15%, and alert noise under 10%.
10. Configure incident tooling and integrations (depends on: 5, 6, 7, 8, 9)
Implement and integrate the tools that automate the defined process.
- Aggregate alerts from the existing six tools into PagerDuty, Opsgenie, or a similar platform.
- Configure on-call schedules, escalation policies, and paging targeted at service owners.
- Integrate status page API for automated or one-click updates.
- Add Slack commands to declare incidents, start war rooms, assign roles, and post status updates.
- Integrate runbook and service catalog access; create postmortem templates in Jira or Notion with action item tracking.
- Ensure audit trails and role assignments are logged for SOC 2.
11. Pilot with 2-3 volunteer teams and iterate (depends on: 10)
Run a controlled pilot before full rollout to validate and refine the process.
- Select 2-3 volunteer teams with representative services and on-call patterns.
- Run the new severity, roles, on-call, alerting, and communication process for 2 weeks.
- Track metrics and gather feedback from on-call engineers, incident commanders, and communications leads.
- Iterate severity thresholds, alert rules, templates, and runbooks based on findings.
- Exit criteria: no SEV1 without a declared incident commander, alert noise below target, and positive on-call survey results.
12. Train and certify all 28 teams (depends on: 11)
Prepare all 28 teams so the process is not a paper exercise.
- Create role-specific training: all engineers, incident commanders, communications leads, scribes, subject-matter responders, and managers.
- Deliver interactive workshops and online modules.
- Run at least two game-day drills per quarter simulating SEV1 and SEV2 incidents, testing escalation, communication, and postmortem.
- Certify incident commanders and communications leads; only certified staff can hold those roles in rotations.
- Include scenario-specific training for account managers and legal or regulatory handling.
13. Full rollout and enforce process (depends on: 12)
Move every team onto the new process and retire old habits.
- Migrate all 28 teams into consolidated on-call schedules and alert routing.
- Retire old alerting tools or methods that bypass the process.
- Activate status page templates and postmortem action tracking globally.
- Monitor the first 30 days with a daily operational review and quick fixes.
- Communicate rollout status to leadership and all engineers; provide office hours and support.
14. SOC 2 readiness and continuous improvement (depends on: 13)
Map the process to SOC 2 and keep improving after go-live.
- Map each process component, including severity, roles, detection, communications, postmortems, and metrics, to SOC 2 control requirements.
- Produce runbooks, evidence of training, incident logs, postmortems, and action item tracking.
- Conduct an internal readiness assessment at month 6 and remediate gaps before the audit at month 8.
- Maintain weekly and monthly metrics reviews to drive continuous improvement.
- Refresh on-call compensation and alert quality rules based on operational data.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
Answer with a single JSON object, and nothing else, that conforms to this JSON schema:
{"description": "The structure every proposing agent must answer with, in every round.", "properties": {"steps": {"items": {"description": "One step of a proposed plan. Dependencies are step_ids of the same proposal.", "properties": {"step_id": {"title": "Step Id", "type": "string"}, "title": {"title": "Title", "type": "string"}, "description": {"title": "Description", "type": "string"}, "dependencies": {"items": {"type": "string"}, "title": "Dependencies", "type": "array"}}, "required": ["step_id", "title", "description", "dependencies"], "title": "Step", "type": "object", "additionalProperties": false}, "title": "Steps", "type": "array"}, "estimated_complexity": {"title": "Estimated Complexity", "type": "string"}, "success_metrics": {"title": "Success Metrics", "type": "string"}}, "required": ["steps", "estimated_complexity", "success_metrics"], "title": "ProposalContent", "type": "object", "additionalProperties": false}
{"steps": [{"step_id": "S1", "title": "Executive mandate, program governance, and interim incident command", "description": "Convert the CEO's outage complaint into a company-level improvement program with a named owner, budget, and deadlines. In the first week, establish an interim command process so no incident remains unowned while the permanent process is designed.\n\n- Appoint Head of Reliability as program owner and CTO/CEO as executive sponsor.\n- Form steering group with Engineering, SRE, Support, CS, Legal, Compliance, HR, Finance, Security.\n- Approve non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.\n- Fund tooling, on-call compensation, training, and 3 dedicated program FTEs; anchor to $1.3M credits.\n- Set timeline: design weeks 1-4, tooling and pilot weeks 5-10, rollout weeks 11-20, audit evidence collection from week 8, dry-run month 6.\n- Stand up interim 24x7 duty officer and single declaration path within 48 hours, compensated retroactively.", "dependencies": []}, {"step_id": "S2", "title": "Baseline data and alert estate analysis", "description": "Re-open the last 31 incidents and profile the current alert estate so every later design decision is evidence-based.\n\n- Re-code each incident: detection source, timestamps, owner, severity, credits, root-cause family.\n- Quantify customer-first detections and missing signals.\n- Analyze two 'nobody in charge' incidents minute-by-minute.\n- Inventory six alerting tools: volume, noise, owner, runbook coverage, top 50 noisy rules.\n- Freeze baseline metrics: MTTD 22m, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Stakeholder listening and resistance mapping", "description": "Treat engineer pushback as design input. Interview all 28 teams plus Support, CS, and Sales to understand real objections and find early adopters.\n\n- Test objection: unpaid work, night load, unfamiliar code, poor runbooks, or blame.\n- Document current informal practices from 12 on-call teams.\n- Recruit 10-15 credible engineers as co-design working group.\n- Survey baseline sentiment on on-call and alerts.\n- Publish promise: paid on-call, page only for owned services, trained commander, reserved reliability capacity.", "dependencies": ["S1", "S2"]}, {"step_id": "S4", "title": "Service ownership catalog and criticality tiering", "description": "Build a machine-readable service catalog that assigns each of the 180 services a single owning team, escalation policy, and criticality tier.\n\n- Define fields: owner team, manager, Slack channel, escalation policy, dependencies, dashboards, runbooks.\n- Tier 0: ledger, money movement, auth, shared PostgreSQL; Tier 1 customer-facing; Tier 2 internal; Tier 3 non-critical.\n- Map customer-visible capabilities and dependencies across regions.\n- Identify orphan and shared services; force ownership decision within 30 days or schedule decommissioning.\n- Publish coverage gaps by tier; Tier 0 gaps become executive escalations.", "dependencies": ["S2", "S3"]}, {"step_id": "S5", "title": "Define severity levels and declaration triggers", "description": "Adopt a five-level severity model with objective payment-specific triggers, so declaration is a lookup rather than debate. Anyone may declare; only the Incident Commander may downgrade.\n\n- SEV0: unauthorized, lost, duplicated, or corrupted money movement; ledger integrity loss; confirmed data breach; both regions failed. Page all roles, exec, legal, and consider payment pause.\n- SEV1: widespread payment failure, no workaround, severe SLA breach, or single-region total loss. Full role activation and public status page.\n- SEV2: significant degradation, multiple customers, or workaround available but SLA risk. IC and SMEs paged; status page if customer-visible.\n- SEV3: limited impact with workaround; team-led, business-hours response.\n- SEV4: internal/no customer impact; ticket only.\n- Auto-escalation: unresolved SEV3 >2h becomes SEV2; unresolved SEV2 >1h becomes SEV1; ledger or security always at least SEV2.\n- Provide decision tree and 12 worked examples from actual incidents.", "dependencies": ["S4"]}, {"step_id": "S6", "title": "Define incident roles, decision authority, and handover rules", "description": "Codify five roles with written responsibilities and explicit authority, so there is never confusion about who is in charge.\n\n- Incident Commander: owns severity, priorities, cross-team pulls, mitigation decisions; does not code.\n- Communications Lead: owns status page, internal updates, account manager briefs, regulator coordination.\n- Scribe: maintains timeline, decisions, and evidence for postmortem/audit.\n- Subject-Matter Responders: diagnose and remediate only their owned services.\n- Executive Liaison (SEV0/SEV1): handles exec communications and external stakeholders.\n- Rule: IC identified within 5 minutes and announced in channel; handover announced and logged; roles may combine below SEV2, never at SEV0/SEV1.", "dependencies": ["S5"]}, {"step_id": "S7", "title": "Design 24x7 incident command and comms staffing model", "description": "Create a central, trained incident command rotation instead of making each of 28 teams field its own commander.\n\n- Recruit 24-30 certified ICs and 12-16 Comms Leads from a cross-team volunteer pool with manager approval.\n- Weekly rotations, primary and secondary; 5-minute acknowledgement SLA with auto-escalation.\n- Scribe pool as entry-level rotation.\n- Night coverage: paid command rotation in New York timezone initially; evaluate follow-the-sun coverage later.\n- Eligibility: certification required; commanders can leave with 30 days notice.\n- Ensure distinct persons for IC, CL, and primary SME on SEV0/SEV1.", "dependencies": ["S6", "S4"]}, {"step_id": "S8", "title": "Design team on-call rotations and cross-team escalation policies", "description": "Define tiered team on-call obligations so engineers are only paged for services they own, and cross-team pages go through the IC.\n\n- Tier 0/1 services: 24x7 primary+secondary, at least 6 trained responders per rotation, one-week shifts.\n- Tier 2/3: business-hours on-call; after-hours escalation list to engineering manager.\n- Platform, infrastructure, database: 24x7 due shared ledger and Kubernetes.\n- Routing: every page resolves via service catalog to owning team's escalation policy; cross-team pages only by IC.\n- Guardrails: max one primary week in four, no consecutive weeks, no on-call in week after leading SEV0/1, protected recovery time after night work.", "dependencies": ["S4", "S6"]}, {"step_id": "S9", "title": "Design paid on-call compensation and fatigue safeguards", "description": "Make on-call paid and legally compliant before any new rotation starts, and convert unpaid pager culture into a fair employment condition.\n\n- Weekly stipends: primary 24x7 $800-1,200, secondary 30-50%, business-hours $300-500; holidays premium.\n- Out-of-hours incident pay: $150 per night page plus hourly beyond one hour; time-off-in-lieu after overnight work.\n- Additional command rotation stipend and SEV0/1 bonus for active responders.\n- Verify FLSA/NY wage rules with HR, Legal, Finance; document exemption treatment.\n- Base annual budget against current $1.3M credits.\n- Track load and trigger staffing review if >2 after-hours pages per person per week sustained.", "dependencies": ["S8", "S3"]}, {"step_id": "S10", "title": "Enforce alert quality standards and noise budget", "description": "Replace 3,400 monthly alerts at 85% noise with a contractual paging standard that makes on-call sustainable.\n\n- Every paging alert must have owning team, customer impact statement, runbook link, severity mapping, and tested threshold.\n- Page only on customer-visible symptoms or SLO burn; cause-based alerts become tickets or dashboards.\n- Set page budget: max 2 out-of-hours pages per person per week; breach triggers mandatory alert-tuning sprint.\n- Auto-quarantine alerts with >5 firings/month without action or >70% no-action acknowledgements.\n- Target <500 actionable pages/month and >75% actionability within 6 months.\n- Weekly per-team alert review, monthly cross-team review.", "dependencies": ["S2", "S4", "S8"]}, {"step_id": "S11", "title": "Build detection uplift: synthetics, SLOs, and support intake", "description": "Shift detection from host metrics to customer outcomes so the company stops hearing about outages from clients first.\n\n- Define SLOs per Tier 0/1 capability: payment initiation, auth, settlement timeliness, API availability, ledger consistency.\n- Deploy external synthetic transactions from both regions every 60 seconds, covering full payment flow and ledger write.\n- Add ledger assurance checks: replication lag, double-entry balance, settlement window countdown.\n- Top-100 customer anomaly detection to catch single-tenant outages.\n- Auto-create triage incident from support tickets or account manager keywords within 5 minutes.\n- Track customer-detected-first as a defect and require a postmortem action.", "dependencies": ["S4", "S10"]}, {"step_id": "S12", "title": "Consolidate alerting and incident tooling", "description": "Collapse six alerting tools into one integrated paging and incident management platform to create a single system of record for people and audit.\n\n- Select paging/on-call platform and incident management layer (e.g., PagerDuty + incident.io/FireHydrant).\n- Implement one-command Slack declaration that auto-creates channel, bridge, pages roles, sets severity, starts timeline.\n- Migrate all monitoring sources into the one tool; decommission legacy paging only after two weeks verified.\n- Integrate service catalog, status page, Jira action tracking, Salesforce/CS customer lists, and conference bridge.\n- Ensure out-of-band paging and offline fallback if a region or chat tool is down.\n- Automate evidence capture for SOC2: timestamps, role assignments, severity changes, comms sent.", "dependencies": ["S5", "S8", "S10", "S11"]}, {"step_id": "S13", "title": "Define acknowledgement and escalation paths", "description": "Make escalation automatic, time-bound, and independent of personal contacts. The clock starts at first signal.\n\n- For SEV0/SEV1: page primary on-call; after 5 minutes unacked page secondary; at 10 page manager and duty IC; at 15 page executive officer.\n- For SEV2: primary ack within 10 minutes, IC assigned within 15; escalate on miss.\n- For SEV3: ack within 30 minutes or create work item.\n- Automatic duty IC page for ledger/security, cross-team, or unresolved ownership.\n- If impact unknown after 15 minutes, raise severity.\n- Cross-team responders summoned by IC have 10-minute acknowledgement obligation.\n- Route alerts with no owner to command rotation, then treat missing ownership as control defect.\n- Human acknowledgement required; delivery confirmation is not sufficient.", "dependencies": ["S12", "S6", "S8"]}, {"step_id": "S14", "title": "Standardize internal communications", "description": "Separate the working war room from executive and stakeholder updates, with fixed cadence and pre-approved templates.\n\n- Auto-create incident channel and one read-only broadcast channel for execs, support, sales.\n- SEV0: internal update every 15 minutes; SEV1 every 30; SEV2 every 60; SEV3 on state change.\n- Template: impact, what we know, what we are doing, ETA/next update, current IC and CL.\n- Executives questions through Executive Liaison only; IC is not interrupted.\n- Support/CS receives affected-customer list and holding statement within 15 minutes for SEV0/SEV1 and 30 for SEV2.\n- Handover protocol for incidents lasting >4 hours: formal IC handover and fatigue check.", "dependencies": ["S6", "S5"]}, {"step_id": "S15", "title": "Standardize customer and status page communications", "description": "Replace ad-hoc status page updates with a timed, role-owned, template-driven process, including account manager outreach.\n\n- Status page timing: SEV0 initial post within 10 minutes, SEV1 within 15, SEV2 within 30; updates every 15/30/60 minutes until resolved.\n- Resolution notice within 30 minutes of mitigation; customer-facing summary in 5 business days for SEV0/SEV1.\n- Pre-approve 12-15 templates with Legal and Comms.\n- Tiered outreach: top 100 accounts get direct call/email from AM within 30 minutes of SEV0/SEV1; long tail via subscription.\n- Use factual language: state impact and next update; never speculate cause or blame vendor.\n- Comms Lead is sole author for customer language.", "dependencies": ["S5", "S14"]}, {"step_id": "S16", "title": "Regulatory, legal, and account manager notification playbook", "description": "Build a notification decision tree and contact matrix so legal/regulatory obligations are assessed early and never forgotten.\n\n- Map obligations: NYDFS Part 500 72-hour cybersecurity event notification, state breach laws, GLBA/FTC, PCI, sponsor bank/card network contractual windows, FinCEN/OFAC if relevant.\n- Add regulatory assessment checkpoint for every SEV0 and security SEV1 within 2 hours, even if not reportable.\n- Maintain 24x7 contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel with named backups.\n- Pre-draft notification templates and test quarterly.\n- Encode customer-specific notification SLAs from enterprise contracts into customer tiering.\n- Account managers receive legal-approved script and affected-customer list.", "dependencies": ["S5", "S15"]}, {"step_id": "S17", "title": "SLA credit workflow and financial impact measurement", "description": "Link every incident to money automatically so severity, credits, and prioritization stay consistent and Finance is never surprised.\n\n- Define availability measurement per contract and capability with Legal/Finance.\n- Auto-compute affected minutes per customer from incident record and telemetry; generate credit proposal within 5 business days.\n- Decide proactive credits for top tier vs claims-based for others; document approval chain.\n- Track credits by root-cause family and incident; feed quarterly reliability investment decisions.\n- Target reducing annual credits from $1.3M to under $400K in first year.", "dependencies": ["S5", "S15"]}, {"step_id": "S18", "title": "Standardize blameless postmortems", "description": "Make postmortems mandatory with fixed deadlines, a single format, and a blameless review forum, replacing 'various formats'.\n\n- Mandatory: all SEV0/SEV1; SEV2 with customer impact, credits, >2h, repeat cause, or detected by customer; near-miss involving ledger.\n- Draft within 3 business days, peer review within 5, publish within 10.\n- One template: timeline, customer/financial impact, detection gap, response gap, contributing factors, what went well.\n- Blameless rules: context-based, no individual blame, never used in performance reviews.\n- Weekly Incident Review Board reviews all postmortems and challenges quality.\n- Publish searchable postmortem library and quarterly recurring causes report.", "dependencies": ["S5", "S6"]}, {"step_id": "S19", "title": "Track postmortem actions with owner and due date", "description": "Fix the 11/64 completion rate by giving every action the same status as customer commitments, with capacity and escalation.\n\n- Every action gets named owner, priority, due date, and Jira ticket auto-created from postmortem.\n- P0 actions prevent SEV0 recurrence, due 30 days; P1 60; P2 90.\n- Reserve 15-20% team sprint capacity for incident actions.\n- Escalation ladder: manager at 7 days overdue, director at 14, CTO dashboard at 30; P0 overdue blocks features.\n- Monthly reporting of closure rate in engineering leadership.\n- Target 90% closure of P0/P1 on time in two quarters.", "dependencies": ["S18", "S12"]}, {"step_id": "S20", "title": "Create runbooks and service readiness bar", "description": "Ensure every service is prepared for 3 a.m. response before it is allowed to page anyone.\n\n- Readiness checklist for Tier 0/1: architecture diagram, dependencies, dashboards, rollback, feature flags, escalation contacts, data-loss statement.\n- Write major-incident playbooks: shared PostgreSQL failure, cross-region failover, Kubernetes control plane loss, partner outage, settlement breach, security compromise.\n- Prioritize ledger playbooks: documented failover, read-only mode, reconciliation, signed RPO/RTO.\n- Runbooks must be tested twice yearly; stale runbooks marked in catalog.\n- No paging alerts without readiness sign-off; gap reported to director.", "dependencies": ["S4", "S8", "S11"]}, {"step_id": "S21", "title": "Train, certify, and simulate incident response", "description": "Build a certification path so 24x7 roles are staffed by people who have practiced, and validate the process through drills.\n\n- Scribe course (2 hrs); Responder (half-day); Communications Lead (1 day); Incident Commander (2 days + shadowing/tabletop).\n- Certify ICs and CLs; recertify annually.\n- New responders shadow two shifts before primary; no primary in first 90 days.\n- Monthly tabletop using real past incidents; quarterly game day including regional failover or ledger scenario.\n- Twice-yearly unannounced paging drill; one regulator/legal exercise annually.\n- Track drill metrics and action items.", "dependencies": ["S6", "S13", "S14", "S15", "S18", "S20"]}, {"step_id": "S22", "title": "Pilot on critical services and iterate", "description": "Prove the process on the highest-risk surface before full rollout. Run a six-week pilot with tight measurement and public results.\n\n- Select 5-6 teams: payments core, ledger/database, platform/Kubernetes, API gateway, plus two existing on-call teams.\n- Activate full stack: severity, roles, command rotation, paid on-call, single tooling, alert budget, status page, postmortems.\n- Weekly pilot retro; fix defects within 48 hours.\n- Exit criteria: MTTD <10 min, IC assigned within 5 min in 95% incidents, pages down 50%, postmortems on time, positive sentiment.\n- Publish one-page pilot result as adoption argument.", "dependencies": ["S12", "S13", "S15", "S19", "S21"]}, {"step_id": "S23", "title": "Wave rollout to all teams and decommission legacy paths", "description": "Roll out to all 28 teams in four waves by criticality, with explicit readiness gates and legacy tool shutdown.\n\n- Wave 1 remaining Tier 0; Wave 2 Tier 1; Wave 3 Tier 2; Wave 4 Tier 3/internal.\n- Per-team onboarding: catalog complete, alerts pruned, runbooks ready, rotation staffed with 6+ certified responders, one IC candidate, one drill passed.\n- Coach per wave 3 weeks.\n- Gate signed by director; failures rescheduled, not waived.\n- After onboarding, freeze legacy alerting paths; no fallback to old tools.\n- Publish adoption scoreboard.", "dependencies": ["S22", "S20"]}, {"step_id": "S24", "title": "Establish metrics, dashboards, and review cadence", "description": "Make process performance visible with a small set of metrics and fixed review meetings, so the program is owned by data.\n\n- Response: MTTD, time to IC, MTTA, MTTM, MTTR, % customer-detected-first.\n- Quality: page volume per person, alert actionability, postmortem on-time, action closure rate.\n- Business: SLA credits, availability vs 99.95, repeat incidents.\n- People: on-call load, page per engineer, sentiment, attrition.\n- Cadence: weekly Incident Review Board, monthly Reliability Review, quarterly Executive/Board review.\n- Dashboards self-serve with targets and named owners.", "dependencies": ["S12", "S19", "S22"]}, {"step_id": "S25", "title": "SOC 2 readiness and internal dry-run audit", "description": "Design evidence as a by-product and test it with an internal walkthrough before the external auditor arrives.\n\n- Map process to SOC2 CC7.3/CC7.4, CC7.2, CC2.2/2.3, CC5, availability criteria.\n- Publish versioned policy documents: Incident Response, Severity, On-Call, Communication, Postmortem.\n- Automate evidence: incident records with timestamps, role assignments, paging logs, status history, postmortems, action board, training register, drill records.\n- Ensure process operates at least 3 months before fieldwork.\n- Run internal dry-run month 6; sample 15 incidents; fix gaps with 8 weeks to spare.\n- Keep remediation log for process deviations.", "dependencies": ["S19", "S23", "S24"]}, {"step_id": "S26", "title": "Continuous improvement, culture, and sustainability", "description": "Prevent the process from decaying after the audit by embedding review, feedback, and roadmap ownership.\n\n- Quarterly process retrospective with IC pool and responders.\n- Re-baseline metrics every six months; raise targets.\n- Year-two roadmap: follow-the-sun coverage, self-healing top 3 causes, error budgets gating releases, blast-radius reduction for shared ledger.\n- Annual policy review and certification renewal.\n- Quarterly on-call sentiment survey with published actions.\n- Board quarterly report on availability, credits, and incident trends.", "dependencies": ["S23", "S25"]}], "estimated_complexity": "high", "success_metrics": "- Median time to detect falls from 22 minutes to under 5 minutes within 6 months of full rollout.\n- Customer-first detection falls from 40% to below 10% within 6 months and below 5% within 12.\n- Median time to mitigate SEV0/SEV1 falls from 3h10 to under 60 minutes within 12 months; SEV2 under 2 hours.\n- Incident Commander assigned and announced within 5 minutes for 95% of SEV0/SEV1 incidents; zero incidents with unclear command beyond 10 minutes.\n- Status page updated within policy time for 95% of SEV0/SEV1/SEV2 (10/15/30 minutes).\n- Monthly pages fall from 3,400 to under 500 with actionability above 75%.\n- All six legacy alerting tools decommissioned by week 16.\n- 100% of SEV0/SEV1 incidents have a blameless postmortem published within 10 business days.\n- Postmortem action closure rises from 17% to 90% for P0/P1 actions on time.\n- SLA credits fall from $1.3M to under $400K in first 12 months.\n- Customer-impacting incidents decline to under 15 per year; repeat root causes under 10%.\n- 100% of 180 services have a named owning team and criticality tier.\n- All 28 teams onboarded by week 24; every Tier 0/1 team has 24x7 primary+secondary coverage with 6+ certified responders.\n- At least 30 certified Incident Commanders and 20 certified Communications Leads active.\n- Paid on-call policy is approved and in payroll before any mandatory night rotation starts.\n- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident response controls pass with zero exceptions.\n- On-call sentiment score improves quarter over quarter; no increase in attrition among on-call engineers.\n- Quarterly game days and twice-yearly unannounced paging drills executed on schedule, each with tracked action items."}Round 2 — refinement 2 of 3
All five plans converged onto a near-common template: four-level severity, a central command corps plus 10–12 domain rotations plus business-hours Layer C, paid on-call gated before any mandatory night rotation, a two-page-per-week budget, one paging platform, 3/5/10-day blameless postmortems, wave rollout with gates and a month-7 mock audit. Real differentiation narrowed to severity anchoring (P4's integrity flag, P1's payment anchors), plan weight and audit sequencing, while P3 essentially re-issued P1's round-1 plan.
What still separates them
- Severity anchoring. P4 step 6 attaches a financial-integrity/security flag to any level (SEV0's forcing function without a fifth level); P1 step 6 adds payments anchors (value at risk per hour, settlement-deadline proximity, named accounts); P2 step 6 demotes percentages to guardrails; P3 step 7 and P5 step 6 stay purely qualitative. P1/P4 also tighten the SEV3 postmortem trigger (customer-detected, >2h, repeat), while P3/P5 keep "postmortem on request".
- Audit sequencing. P4 confirms the observation window in week 1 (step 1), P2 in step 3, P3 in step 5; P1 leaves it in step 30, which depends on steps 24/26/28, and cuts independent testing from two rounds to one at month 5. P2/P3/P5 keep tests in months 4 and 6.
- Coverage economics. P1 alone funds an overnight Duty Triage Desk owning the first ten minutes (step 8) but never sizes or budgets it; P4 alone shows the arithmetic (260 engineers sustain 10–12 rotations, not 28, step 8); P2 alone refuses to merge small teams for schedule convenience, scopes the six-responder rule to direct 24x7 rotations (steps 8, 22) and keeps a surge roster for concurrent incidents.
- Weight and pace. P1/P3 run 32 steps with standalone noise campaign, Policy v1 and risk register; P4 compresses to 26 with an 11-bullet comms mega-step (15) and the risk register demoted to last (26); P5 merges regulatory+credits (16) and postmortems+actions (17) and has no policy-publication step; P2 has 26 steps, no risk register at all, and the fastest schedule (Tier 0/1 by week 10, all teams by week 18 versus week 24 elsewhere).
Who took what from whom
- P2 and P5 abandoned their SEV0 tiers for the SEV1–SEV4 scale of P1/P3/P4 (P2 step 6, P5 step 6); P4 instead converted SEV0 into a financial-integrity/security flag (step 6) and P1 into payments anchors (step 6) — the round's one genuine design argument, resolved three different ways.
- P3 imported P1's round-1 plan almost wholesale (21 matching steps: Layer A/B/C 9, noise burn-down 26, Policy v1 24, risk register 31, 90-day inspect 32) and grafted on P2's day-one compliance step (P3 step 5 from P2 step 4) and P2's fake-redundancy check (P3 step 12 from P2 step 12).
- P5 took P1's Layer A/B/C coverage (step 8), readiness bar and ledger blast-radius workstream (step 10), risk register (step 25) and culture step (23), and tightened its own 1-in-4 rotation cap to P1's 1-in-6 (step 9).
- P2 took the standalone listening tour/fairness contract from P1 step 4, P3 step 3 and P4 step 4 (now P2 step 4) and concrete pay bands from P1/P5 step 9; in the other direction P4 took P2's week-1 observation-window confirmation (P4 step 1) and P1 added P2's 99.95% availability metric and an Internal Audit seat (step 1).
- Nobody took P5's named vendor shortlist (PagerDuty/incident.io/FireHydrant, round-1 step 12) — P5 dropped it itself; nobody adopted P2's numeric declaration thresholds (10% / 1–10% failure rates), which P2 also demoted; and only P2 (step 8) still plans for simultaneous incidents with a surge roster.
The calls of this round
Influences: who took what from whom
| Round 2 ↓ · round 1 → | Proposal 1 | Proposal 2 | Proposal 3 | Proposal 4 | Proposal 5 | New steps |
|---|---|---|---|---|---|---|
| Proposal 1 |
kept27 | same titles0 analyst sees+3 / −1 | same titles0 analyst sees+0 / −0 | same titles1 analyst sees+0 / −0 | same titles0 analyst sees+0 / −1 | new4 |
| Proposal 2 |
same titles0 analyst sees+5 / −3 | kept18 | same titles0 analyst sees+0 / −0 | same titles1 analyst sees+0 / −0 | same titles0 analyst sees+0 / −0 | new7 |
| Proposal 3 |
same titles21 analyst sees+1 / −1 | same titles1 analyst sees+2 / −1 | kept5 | same titles5 analyst sees+1 / −0 | same titles0 analyst sees+0 / −0 | new0 |
| Proposal 4 |
same titles6 analyst sees+3 / −2 | same titles2 analyst sees+2 / −1 | same titles0 analyst sees+0 / −0 | kept7 | same titles1 analyst sees+0 / −0 | new10 |
| Proposal 5 |
same titles4 analyst sees+5 / −2 | same titles4 analyst sees+0 / −1 | same titles5 analyst sees+0 / −0 | same titles3 analyst sees+0 / −0 | kept6 | new4 |
Structure is unchanged at 32 steps; the gains are inside steps. Step 8 adds a paid overnight Duty Triage Desk, step 6 adds payments-specific severity anchors, step 11 adds an incident-replay test with "replay coverage" as a leading metric. Independent control testing shrank from two rounds to one.
- Duty Triage Desk (step 8) owns the first ten minutes of any page — the round's only concrete mechanism for cutting night MTTD without waking domain owners, and a direct answer to the pager objection.
- Payments anchors in step 6 (value at risk per hour, settlement-deadline proximity, number of affected named accounts) make severity quantitative without importing arbitrary failure percentages.
- Replay test in step 11 plus the metric "a named detector would have caught ≥90% of the 31 historical incidents, with the expected detection minute" — makes the detection uplift falsifiable rather than aspirational.
- New outcome metrics: availability ≥99.95% per customer journey by month 6; executive-signed ledger RPO/RTO with rehearsed read-only mode by month 6 and board-visible blast-radius milestones (step 21).
- Rollout deadline corrected from five to four months before the audit report date (step 28) — now consistent with a week-24 finish and a month-8 audit.
- Internal Audit added to the decision group (step 1); charter published company-wide on day 2.
- Independent testing cut from months 4 and 6 to a single month-5 test (step 30), where P2/P3/P5 keep two.
- SOC 2 mapping still sits in step 30, which depends on steps 24, 26 and 28 — the auditor conversation is formally gated behind policy publication and full rollout despite the text saying "confirm the observation window early".
- The Duty Triage Desk is unsized and unfunded: no headcount, source pool or line in step 1's budget, although it is a new paid overnight rotation.
- Still the heaviest plan — 32 steps and 24 metric lines — with steps 10 and 27 largely overlapping on alert quality.
- Proposal 2 : Percentage thresholds are declaration guardrails that must never justify under-classifying integrity, settlement or security risk.
- Proposal 2 : Internal Audit on the governing body and independent control testing before fieldwork.
- Proposal 2: Monthly availability ≥99.95% as an outcome metric (also in P3's round-1 metrics).
- Proposal 2 : A standalone day-one step that designs compliance and evidence controls before the process is finalised.
- Proposal 5 : A five-level scale with a separate SEV0 money/security crisis tier.
+ Charter the program: one owner, one mandate, funded, dated before the audit+ Seven-day operating floor so the next outage already has an owner+ Three-layer 24x7 coverage: command corps, domain rotations, triage desk+ Live execution doctrine and payments safety rulesExecutive mandate, single owner, funding and non-negotiablesSeven-day interim command bridge24x7 coverage model: central command corps, local expertiseLive incident execution doctrine
The plan produced
1. Charter the program: one owner, one mandate, funded, dated before the audit new
Turn the CEO email into a chartered company program with a single accountable owner and authority across all 28 teams. Incident response becomes a company operating process, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (after 1) new
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (after 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the written on-call deal (after 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (after 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: Tier 0 (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (after 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- SEV1 (crisis): money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- SEV2 (critical): material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- SEV3 (contained): narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- SEV4: no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring. Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (after 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- Incident Commander: owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- Communications Lead: the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- Subject-Matter Responders: engineers of the owning team. They mitigate; they do not run the room.
- Executive Duty Officer (SEV1): removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation or privileged-access rules.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (after 5, 7) new
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- Layer A — Incident Command corps: ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- Layer B — domain responder rotations: consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- Layer C — everyone else: business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- Duty Triage Desk: a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (after 4, 8)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Hard gate: no mandatory night rotation starts before its compensation is live in payroll. Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (after 3, 5) from P4 step 10
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a page budget of two out-of-hours pages per responder per week. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (after 5, 10)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the replay test: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (after 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (after 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (after 7, 12, 13) new
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- Payments guards: protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
15. Internal communications protocol (after 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (after 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity or blame; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (after 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded even when the answer is "not reportable", with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
18. SLA credit and financial impact workflow (after 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Use the credit delta as the standing business case for on-call pay and reserved reliability capacity.
19. Blameless postmortem standard and Incident Review Board (after 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did?
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, reserved capacity and enforcement (after 12, 19)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
21. Runbooks, readiness bar and ledger blast-radius reduction (after 5, 8)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (after 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days plus shadowing): command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Comms Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (after 12, 21, 22)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced overnight paging drill to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
24. Publish Incident Management Policy v1 and the exception register (after 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (after 9, 11, 12, 21, 22, 24)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (after 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- Response: time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Alert noise burn-down campaign (after 10, 12, 25)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
28. Wave rollout to all 28 teams with readiness gates (after 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. A failed gate is rescheduled, never waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
29. Change management, fairness and pager culture (after 4, 9, 25)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (after 24, 26, 28)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and pre-committed contingencies (after 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (after 26, 28, 30)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to leading indicators: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
- A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 3efb9ff1-28a5-40aa-8f73-9853e91aa095, Agent: opus5_refine_1, LLM: anthropic/claude-opus-5):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (32):
1. Executive mandate, single owner, funding and non-negotiables
Convert the CEO email into a chartered program with one accountable owner and authority over all 28 teams. Incident response becomes a **company operating process**, not a per-team choice.
- Name an executive sponsor (CTO) and one accountable owner (Head of Reliability / Incident Management) with a small permanent office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group with Engineering, SRE/Platform, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours. It is not a 28-person committee.
- Fix the non-negotiables now: one severity scale, one paging tool, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.
- Set the clock deliberately earlier than the audit: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready week 28.
- Approve budget against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises.
2. Seven-day interim command bridge (depends on: 1)
Do not let design work leave the company unprotected for six weeks. Put a crude but real process in place within seven days and improve it later.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster from engineering managers and senior SREs, primary plus backup, 24x7. Pay it retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Tell Support to escalate credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items: complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This becomes both the design input and the "before" picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start / detect / declare / commander-assigned / mitigate / resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. This list drives the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome. Identify the top 50 rules producing most of the noise and every rule with no owner or runbook.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS and Sales in two weeks. Test what the objection actually is: unpaid work, lost sleep, unfamiliar code, missing runbooks, or fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey (fairness, trust in alerts, willingness, burnout) to re-measure at 60, 120 and 365 days.
5. Service catalog, ownership and money-path tiering (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that is the single source of truth for paging, impact and audit.
- One accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not by technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions and third parties, including sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Every orphan service gets an owner in 30 days or a decommission date approved by the sponsor. Tier 0 without an owner is an executive escalation.
6. Severity scale, declaration rules and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- **SEV3 (major, contained):** narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- **SEV4:** no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring; only the commander may downgrade, with the evidence recorded.
- Lifecycle: Detected → Declared → Triaged → **Mitigated** (customer impact ends) → Monitoring → **Resolved** (backlog processed and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Incident roles, authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, roles, cadence and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit spend. Retains command when a VP joins.
- **Communications Lead:** single voice for status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates for SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team; they mitigate, they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance and Vendor Management join on defined triggers.
- Rules: command claimed within 5 minutes, stated in channel ("I am IC"), distinct people for command, comms and technical lead at SEV1/SEV2, and every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation or privileged-access rules.
8. 24x7 coverage model: central command corps, local expertise (depends on: 5, 7)
Do not create 28 night rotations. Centralise coordination in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- **Layer A — Incident Command corps:** ~30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of ~18 from Support, CS and engineering management; a scribe pool used as the training entry point.
- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- **Layer C — everyone else:** business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing, with cost and timeline: a US-only paid night rotation now, a Lisbon or APAC follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
9. On-call compensation, labour compliance and fatigue safeguards (depends on: 8, 4)
Unpaid on-call in New York is both a retention problem and a legal exposure. Pay for it before asking anyone to sign up, and publish the numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 for business-hours rotations, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover rather than expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, tax treatment and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a ~15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
10. Alert quality standard, page budget and noise burn-down (depends on: 3, 5)
3,400 alerts at 85% noise is the reason detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human, and burn the backlog down deliberately rather than by mass silencing.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a **page budget** of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements; return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
- Target 3,400 → under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
11. Detection uplift on the money path (depends on: 10, 5)
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls; five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrating without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack or the identity provider is unavailable. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live incident execution doctrine (depends on: 7, 12, 13)
Give responders one short operating procedure for the first minutes through closure. Priority is limiting customer and financial harm, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; controlled backlog drain; **reconciliation completed before any payment or ledger incident is declared resolved**.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and Customer Success, and reopening if impact recurs.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook and RSS; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack; the long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity or blame; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded **even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review; Legal owns outbound regulatory text, the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
18. SLA credit and financial impact workflow (depends on: 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one, and use the delta as the standing business case for on-call pay and reliability capacity.
19. Blameless postmortem standard and Incident Review Board (depends on: 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first?** and **why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, capacity reservation and enforcement (depends on: 19, 12)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
21. Runbooks, readiness bar and ledger blast-radius reduction (depends on: 5, 8)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the single largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel workstream on blast-radius reduction — tenant or function partitioning, read replicas, and isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (depends on: 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days + shadowing):** command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence for their own domain before independent primary duty; two shadow shifts minimum, never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (depends on: 22, 12, 21)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 11, 12, 21, 22, 9)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect (from first impact), time to declare, time to commander, acknowledgement, mitigate, resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; team scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Wave rollout to all 28 teams with readiness gates (depends on: 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director; failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
28. Alert noise burn-down campaign (depends on: 10, 12, 25)
Run the noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
29. Change management, fairness and pager culture (depends on: 4, 9, 25)
Run this from day one in parallel. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (depends on: 24, 26, 27)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and contingencies (depends on: 1)
Name the ways this programme fails and pre-commit the response. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 26, 27, 30)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 2 (ID: 2e38c07e-bf01-4b70-ba95-0ca1ef39e2d8, Agent: gpt5.6-sol_refine_2, LLM: openai/gpt-5.6-sol):
Estimated Complexity: high
Success Metrics: - By day 7, every suspected SEV0–SEV2 uses one incident record, one coordination channel, and one named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named commander for more than 10 minutes; at least 95% are assigned within 5 minutes.
- By day 30, 100% of Tier 0 services have a named owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, every Tier 0 and Tier 1 customer journey has compensated 24×7 command and subject-matter coverage.
- By day 120, all 180 production services have an owner, response tier, tested escalation path, and appropriate coverage model.
- No mandatory night rotation starts before its compensation, training, access, and staffing controls are active.
- Every critical primary rotation has at least six qualified responders or a documented, expiring executive exception by day 120.
- No responder is routinely scheduled for primary duty more often than one week in six by day 120.
- At least 95% of critical pages are acknowledged within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by day 120 and under 60 minutes by month 8.
- At least 95% of SEV0 and SEV1 customer notices are issued within 15 minutes of declaration by month 3.
- At least 95% of customer-visible SEV2 notices are issued within 30 minutes by month 3.
- At least 95% of incidents meet their required update cadence by month 3.
- Monthly paging volume falls from 3,400 to no more than 1,500 by day 90 and no more than 700 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6 without loss of critical detection coverage.
- All new paging alerts satisfy the owner, impact, action, dashboard, runbook, deduplication, and escalation standard by day 60.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 currently open historical actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date by month 6, with effectiveness evidence.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Two end-to-end cross-company exercises, including regional failure and ledger recovery, are completed before the audit.
- Monthly availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- SLA credits decline by at least 50% on an annualized trailing basis within 12 months.
- Quarterly on-call surveys show improving fairness and sustainability, with at least 75% favorable responses by month 6.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (24):
1. Create the mandate, ownership, and funding
Launch incident management as a company operating program within 48 hours. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable program owner.
- Form a small steering group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Fund two to three implementation staff, paging and incident tooling, observability work, training, exercises, and on-call compensation.
- Reserve 10%–15% of engineering capacity for alert remediation, runbooks, and incident actions.
- Authorize incident commanders to freeze deployments, order rollback, disable features, shift traffic, and invoke continuity plans.
- Preserve financial controls. Incident commanders may coordinate ledger recovery but may not bypass dual approval, privileged-access controls, or reconciliation.
- Set the delivery target: critical controls operational within 60 days, enterprise rollout within 120 days, and a mock audit in month 7.
2. Install an interim process in seven days (depends on: 1)
Do not wait for new tools or the final policy. Put a minimum viable incident process into operation immediately and start collecting evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communications schedule.
- Establish one monitored declaration path through chat, telephone, and the existing paging tools.
- Create a standard incident channel, bridge, incident document, and naming convention.
- Staff temporary primary and backup incident commanders 24×7 from the existing on-call teams and engineering leadership.
- Compensate interim duties retroactively under the final compensation policy.
- Require a named incident commander within 10 minutes for every suspected major incident.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Hold a daily 15-minute control review until the permanent process is live.
3. Build the baseline, ownership catalog, and risk map (depends on: 1)
Establish the facts behind the current failures and assign every production component an owner. Use the resulting catalog as the source for routing, escalation, and audit evidence.
- Reconstruct all 31 customer-impacting incidents, including first impact, detection source, declaration, command assignment, mitigation, resolution, customer communications, and credits.
- Analyze the two incidents with no clear leader and every case detected first by customers.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, databases, queues, regional dependencies, payment processors, and banking partners.
- Assign one accountable team, engineering manager, product owner, primary escalation, secondary escalation, dashboard, and runbook to each service.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 based on financial integrity, customer impact, contractual exposure, and dependency centrality.
- Map critical customer journeys to their application, PostgreSQL, Kubernetes, regional, and third-party dependencies.
- Inventory the six alert sources and all 3,400 monthly alerts by owner, volume, actionability, and duplication.
- Interview representatives from all 28 teams and baseline on-call sentiment, fatigue, and objections.
- Give orphan services an owner or decommissioning decision within 30 days.
4. Design compliance and evidence controls from day one (depends on: 1)
Map the operating process to audit, legal, contractual, and record-retention requirements before finalizing it. Confirm the expected SOC 2 observation period with the auditor immediately.
- Map controls to applicable SOC 2 criteria for monitoring, incident identification, response, recovery, communications, corrective action, access, and availability.
- Define evidence required for declarations, pages, acknowledgements, role assignments, decisions, status updates, postmortems, actions, training, drills, and exceptions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve the Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, and Postmortem Standard.
- Record control exceptions with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing evidence before the audit.
5. Adopt one severity and lifecycle standard (depends on: 2, 3, 4)
Use four impact-based severity levels across operational, security, data, and third-party incidents. Start at the highest credible severity when facts are uncertain, then downgrade with recorded evidence.
- **SEV0 — financial or security crisis:** suspected ledger corruption, unauthorized or duplicated funds movement, material data compromise, both-region loss, or a decision to suspend payment processing. Page all roles immediately; engage executives, Security, Legal, Compliance, and Risk; assess regulatory duties within one hour; use distinct role holders; require a postmortem.
- **SEV1 — critical availability event:** a core payment journey is unavailable, payment failures exceed an initial 10% guardrail for five minutes, regional loss has impaired failover, a settlement deadline is at imminent risk, or rapid error-budget burn makes a material SLA breach likely. Assign all roles; notify internal stakeholders within 10 minutes; publish customer status within 15 minutes; update every 30 minutes; require a postmortem.
- **SEV2 — major bounded event:** approximately 1%–10% of payment attempts fail, a material customer subset or critical customer is down, degradation has a workaround, or contractual impact is likely. Assign an incident commander and responders; add communications and scribe roles for customer impact; publish status within 30 minutes; update every 60 minutes; require a postmortem for customer-visible events.
- **SEV3 — limited event:** localized impact, a safe workaround, and no financial-integrity, security, regulatory, or material contractual risk. The owning team leads; page only if immediate action is necessary; use a ticket otherwise.
- Treat the percentage thresholds as declaration guardrails, not reasons to under-classify integrity, settlement, security, or strategic-customer risk.
- Permit any employee to declare an incident. Only the incident commander may lower severity, with the rationale logged.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of active customer harm. Declare resolution only after stability, backlog recovery, and required ledger reconciliation.
6. Define roles, authority, and handoffs (depends on: 5)
Separate command, communication, recordkeeping, and technical repair. One named person must hold command at every moment of a major incident.
- **Incident Commander:** owns severity, priorities, role assignment, escalation, decision cadence, mitigation coordination, and closure. The commander does not act as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language. Legal or Compliance retains ownership of regulatory submissions.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, and status changes.
- **Subject-Matter Responders:** diagnose and mitigate services for which they have accepted ownership, access, training, and runbooks.
- **Executive Duty Officer:** removes organizational obstacles and approves exceptional business decisions without displacing the incident commander.
- Require separate commander, communications, scribe, and primary technical lead for SEV0 and SEV1.
- For customer-visible SEV2, keep the commander separate from the primary technical responder; communications and scribe may be combined if workload permits.
- Announce every role assignment and transfer in the incident channel. Require verbal and written handoff with current impact, decisions, risks, and next actions.
- Keep executives and account managers out of the technical command path; questions flow through the Executive Duty Officer or Communications Lead.
7. Create sustainable 24×7 coverage across the service estate (depends on: 3, 6)
Use central command coverage and service-domain responder coverage rather than creating 28 fragile night rotations. Engineers remain responsible only for services they own or have formally trained to support.
- Build a company incident-command pool of approximately 18–24 certified senior engineers and managers, with a primary and backup scheduled at all times.
- Build 12–16-person communications and scribe pools using Support, Customer Operations, Engineering Operations, and qualified engineering managers.
- Schedule an Engineering Director or equivalent as the 24×7 Executive Duty Officer.
- Group related service owners into approximately 8–12 coherent responder domains only where members have training, access, and explicit acceptance.
- Require Tier 0 and Tier 1 domains to provide 24×7 primary and secondary responders, normally with at least six trained people in each sustainable rotation.
- Give Tier 2 services business-hours coverage plus a maintained manager escalation path. Treat Tier 3 conditions as tickets unless their impact changes.
- Maintain distinct but coordinated coverage for the ledger application and the PostgreSQL platform.
- Route unknown-owner events to the duty commander and platform triage temporarily. Treat every such event as an ownership control defect.
- If a small team cannot staff a fair rotation, merge coverage only after training or provide headcount, service reassignment, or decommissioning.
8. Approve compensation and fatigue protections (depends on: 7)
End unpaid on-call before expanding mandatory coverage. Publish the policy through HR, Legal, Finance, and Payroll within 14 days.
- Pay fixed weekly stipends for primary and secondary service rotations.
- Pay separate stipends for duty commander, communications, and scribe assignments.
- Provide additional call-out compensation or equivalent paid recovery time for material after-hours work.
- Apply overtime and reporting rules correctly for non-exempt employees under federal and New York requirements.
- Pay higher rates for company holidays and provide a protected recovery day after qualifying overnight work, SEV0 events, or prolonged SEV1 response.
- Target no more than one primary week in six and prohibit simultaneous primary assignments.
- Avoid consecutive primary weeks and make all swaps visible in the paging system.
- Reduce sprint commitments for people carrying primary duty rather than expecting normal delivery capacity.
- Provide a documented accommodation path for health, disability, or caregiving constraints without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
9. Set the service readiness and runbook standard (depends on: 3, 5, 7)
A team cannot respond effectively at night without ownership, access, telemetry, and rehearsed recovery procedures. Apply a formal readiness gate to every Tier 0 and Tier 1 service.
- Require a current architecture diagram, dependency map, dashboards, SLOs, runbooks, rollback method, feature-control method, contacts, and tested escalation path.
- Record RTO, RPO, data-integrity requirements, regional mode, and customer-facing capabilities in the service catalog.
- Link every paging alert to the exact runbook step expected from the responder.
- Prevent new paging alerts for services that fail readiness review. Preserve existing critical detection through a documented exception until a safe replacement exists.
- Write priority playbooks for PostgreSQL failure, ledger-integrity investigation, regional failover, Kubernetes control-plane degradation, payment-processor failure, queue backlog, credential compromise, and payment suspension.
- For the ledger, document read-only or stop-processing modes, failover controls, replay and duplicate protections, backlog handling, and post-recovery reconciliation.
- Require runbook review after material incidents and at least twice per year through exercises.
10. Establish one paging and incident system of record (depends on: 4, 5, 6, 7)
Consolidate paging and incident coordination without requiring an unsafe big-bang replacement of every monitoring system. Monitoring sources may remain specialized, but all human pages must enter one controlled platform.
- Select one enterprise paging platform and one integrated incident record, with chat, telephone, SMS, conference, status-page, ticketing, and service-catalog integrations.
- Initially ingest events from all six current tools, then deduplicate, correlate, and route by service ownership.
- Automate incident channel and bridge creation, role paging, timeline capture, severity changes, communications reminders, and postmortem creation.
- Preserve immutable records of declarations, acknowledgements, role assignments, decisions, and messages.
- Use role-based access, multifactor authentication, break-glass controls, and periodic access reviews.
- Provide telephone and offline fallback procedures for loss of chat, identity, the paging vendor, or an AWS region.
- Test paging and fallback paths weekly.
- Retire a legacy paging route only after its signals have owners, quality review, successful end-to-end tests, and at least two weeks of verified operation in the new path.
11. Enforce alert quality and burn down noise safely (depends on: 3, 10)
Treat paging alerts as production products with owners and quality requirements. Do not reduce noise by silently disabling detection.
- Require every page to identify the service, owner, customer or SLO risk, urgency, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Define an actionable page as one requiring prompt human judgment or intervention that materially reduces customer, financial, security, or contractual risk.
- Route informational, capacity-planning, and non-urgent conditions to dashboards or ticket queues.
- Prefer symptom and error-budget burn alerts over raw CPU, memory, pod, or log-volume thresholds.
- Run new alerts in shadow mode for at least seven days unless an emergency exception is approved.
- Review alerts with less than 50% actionability, more than three firings in seven days, or repeated no-action acknowledgements within two business days.
- Require compensating detection and an owner before suppressing or removing an alert.
- Set a responder load target of no more than two after-hours pages per person per week. A breach creates a mandatory alert-remediation plan.
- Review alert actionability, duplication, missed detection, and page load monthly by domain.
- Prioritize the small number of rules producing most of the current 85% noise.
12. Detect payment and ledger failures before customers (depends on: 3, 11)
Shift detection from infrastructure symptoms to customer journeys and financial outcomes. Set internal objectives stricter than the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, webhooks, and reporting freshness.
- Run external synthetic transactions from outside the production boundary and through both regions at least every minute for critical paths.
- Monitor payment failure rates, latency, queue age, delayed value, settlement-window risk, regional asymmetry, and third-party response quality.
- Add continuous ledger controls for reconciliation breaks, unexpected balances, duplicate identifiers, replication lag, backup health, and failover readiness.
- Add tenant or cohort anomaly detection for high-value customers and common payment methods.
- Convert high-priority Support, account-manager, bank, and processor reports into incident candidates within five minutes.
- Review every customer-first incident as a missed-detection defect and create a corrective action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Codify escalation and live incident execution (depends on: 6, 10)
Create one time-bound path from signal to ownership and mitigation. Delivery of a notification does not count as acknowledgement.
- Page the owning Tier 0 or Tier 1 primary immediately; page the secondary after five unacknowledged minutes; page the domain manager at 10 minutes; escalate to the engineering director at 15 minutes.
- For SEV0–SEV2, page the duty commander immediately. Page the backup after five minutes and require the Engineering Director to assume command if no certified commander owns the event by 10 minutes.
- Automatically involve command for integrity concerns, security concerns, regional events, cross-team impact, critical customer journeys, or unresolved ownership.
- If impact remains materially unknown after 15 minutes, increase response posture rather than waiting for certainty.
- Open one incident channel, bridge, and system record. State severity, known impact, assigned roles, current objective, and next update time.
- Freeze unrelated changes during SEV0 and SEV1 unless the commander records an exception.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require controlled backlog processing and reconciliation before resolving payment or ledger incidents.
- Require a formal command handoff for incidents extending beyond four hours or when fatigue impairs a role holder.
14. Standardize internal and customer communications (depends on: 5, 6, 10)
Communicate known impact early without waiting for root cause. The Communications Lead uses approved facts and always states the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish customer status within 15 minutes when customer-visible and legally safe, then update at least every 30 minutes.
- For SEV1, brief internal stakeholders and publish status within 15 minutes, then update every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send notice within 30 minutes, then update every 60 minutes.
- For SEV3, communicate directly to affected customers only when impact or contract terms require it.
- Give account managers an approved statement and affected-customer list within 30 minutes for SEV0 or SEV1 and within 60 minutes for SEV2.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. Describe affected capabilities, symptoms, workarounds, and the next update.
- Do not speculate about root cause, blame, data integrity, security scope, or recovery time.
- Post a Monitoring update within 30 minutes of mitigation. Post Resolved only after stability and required reconciliation.
- Provide a customer-facing incident summary within five business days for qualifying incidents.
- Record and approve any delay or restriction of public details during an active security threat.
15. Operationalize regulatory, partner, contract, and credit decisions (depends on: 4, 5, 14)
Give Legal and Compliance a timed decision process while keeping technical command with the incident commander. Record both reportable and non-reportable determinations.
- Build a jurisdiction and obligation matrix covering applicable NYDFS requirements, state breach laws, GLBA or FTC requirements, PCI obligations, money-transmitter rules, sponsor banks, payment networks, cyber insurance, and customer contracts.
- Validate applicability and deadlines with counsel rather than assuming every operational incident is reportable.
- Begin reportability assessment immediately for every SEV0, security event, suspected ledger-integrity event, and relevant SEV1.
- Target a documented initial legal assessment within one hour for SEV0 and within two hours for other potentially reportable events.
- Record the decision, evidence, approver, legal deadline, submission owner, and confirmation of delivery.
- Maintain tested 24×7 contacts for regulators, banks, networks, insurers, outside counsel, and critical vendors.
- Encode customer-specific notification clocks and channels in the customer record used by the Communications Lead.
- Have Finance calculate affected minutes, likely credits, and contractual exposure from the incident record within five business days.
- Review whether proactive credits or claims-based handling applies by customer segment and contract.
16. Make postmortems mandatory, consistent, and blameless (depends on: 5, 6, 10)
Use one review standard to learn from incidents and test whether controls worked. Keep learning reviews separate from performance or misconduct processes.
- Require a postmortem for every SEV0 and SEV1.
- Require one for customer-visible SEV2, customer-first detection, incidents lasting more than two hours, contractual breaches, repeat failures, control gaps, and ledger-integrity near misses.
- Produce the factual draft within three business days, conduct the review within five, and publish the approved version within 10.
- Use one template covering summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, lessons, and actions.
- Analyze why detection was not earlier and why mitigation took as long as it did.
- Examine technical, organizational, process, testing, dependency, and incentive factors rather than forcing a single root cause.
- Use a trained facilitator and written blameless rules. Describe decisions in the context and information available at the time.
- Publish broadly useful findings internally while maintaining restricted versions for security, privacy, personnel, or privileged material.
- Review postmortem quality and recurring factors in a weekly Incident Review Board.
17. Enforce action ownership and effectiveness tracking (depends on: 16)
Treat corrective actions as risk commitments rather than suggestions. Closing a ticket is insufficient without evidence that the control or system behavior improved.
- Give every action one named individual owner, manager, priority, due date, expected risk reduction, verification method, and linked work item.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with intermediate milestones.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Place recurrence-prevention actions for SEV0 and SEV1 ahead of discretionary feature work unless an executive accepts the residual risk.
- Escalate overdue high-risk actions to the manager after seven days, director after 14 days, and CTO after 30 days.
- Require written residual-risk acceptance, compensating controls, and a new review date when high-risk work is deferred.
- Verify completed actions through tests, telemetry, drills, or production evidence.
- Triage the 53 currently open historical actions within 30 days. Complete, re-plan, or formally accept the risk, prioritizing ledger, regional, security, and detection items.
18. Measure outcomes, controls, business impact, and human load (depends on: 10, 11, 14, 17)
Use a balanced scorecard so teams are not rewarded for suppressing alerts or avoiding incident declarations. Report medians and 90th percentiles, not averages alone.
- Measure time from first impact to internal detection, declaration, acknowledgement, command assignment, mitigation, and resolution.
- Track customer-first detection, missed escalations, status-page timeliness, update-cadence compliance, and role conflicts.
- Track incident count, recurrence, availability, error-budget burn, affected payment value, delayed transactions, reconciliation breaks, and SLA credits.
- Track page volume, actionability, duplicates, after-hours pages, missed acknowledgements, and load per responder.
- Track postmortem timeliness, action completion, action age, verified effectiveness, and repeated contributing factors.
- Track rotation size, duty frequency, recovery days, swaps, attrition signals, and quarterly responder sentiment.
- Hold a weekly Incident Review Board for incidents, actions, missed controls, and noisy alerts.
- Hold a monthly executive reliability review for trends, investment decisions, contractual exposure, and accepted risks.
- Hold a quarterly resilience and controls review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile dashboard data against customer cases and sampled incident records monthly to identify missing incidents or metric gaming.
19. Train participants and address pager resistance (depends on: 6, 7, 8, 13, 14, 16)
Introduce the operating model as a fair exchange, not an audit mandate. The message is that people are paid, paged only for accepted services, supported by trained command, and given capacity to fix defects.
- Train all employees to recognize impact, declare an incident, and find the incident channel and customer status page.
- Give all 260 engineers role-based training in severity, escalation, evidence preservation, handoff, and financial-integrity precautions.
- Certify incident commanders through instruction, simulation, two shadowed events or exercises, and observed performance.
- Train Communications Leads in status writing, account-manager briefing, contractual clocks, and legal escalation.
- Train scribes in timeline quality, decision capture, fact-versus-hypothesis labeling, and evidence handling.
- Require responders to demonstrate access, dashboards, rollback, runbooks, and escalation competence before primary duty.
- Require two shadow shifts before independent primary on-call.
- Appoint one adoption champion in each team and hold weekly office hours during rollout.
- Publish compensation, fatigue protections, service boundaries, and the accommodation process before assigning shifts.
- Use paid working time for training, exercises, shadowing, runbook work, and certification.
- Survey engineers at baseline, day 60, day 120, and quarterly thereafter.
20. Pilot on the payment critical path (depends on: 8, 9, 11, 12, 13, 14, 17, 19)
Run a four-to-six-week pilot across the highest-risk customer journey. Use real incidents and exercises to correct the process before wider rollout.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, authentication or API edge, settlement, and Support intake.
- Activate paid primary and secondary rotations, central command, communications, the incident record, status templates, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new system authoritative.
- Review every pilot page within one business day for routing, actionability, responder load, and missing context.
- Have the program team coach incidents without silently taking command from assigned role holders.
- Hold a weekly pilot retrospective and correct critical process or tool defects within 48 hours.
- Require 95% command assignment within five minutes, 95% timely communications, no unpaid pages, complete postmortems, and at least a 50% page-noise reduction before expansion.
21. Roll out by customer journey and risk (depends on: 18, 20)
Expand in fixed waves rather than waiting for every team to become perfect. Apply explicit readiness gates and time-limited exceptions.
- Days 0–7: operate the interim declaration, command, and communications process.
- By day 30: complete Tier 0 ownership, activate the first certified command roster, approve compensation, and begin the pilot.
- By day 60: provide compensated 24×7 coverage for every Tier 0 and Tier 1 customer journey and route all critical pages through the new platform.
- By day 90: complete critical detection upgrades, customer communications, regulatory playbooks, and the first major alert-noise reduction.
- By day 120: assign every production service a response tier, owner, tested escalation, and appropriate coverage model.
- Roll out remaining teams in two-to-three-week waves ordered by customer and dependency risk.
- Require each wave to pass ownership, alert, runbook, access, training, compensation, and tabletop gates.
- Disable legacy paging paths after verified cutover rather than leaving ambiguous parallel obligations.
- Publish a weekly adoption dashboard by team and escalate failed gates as business risks.
- Never start mandatory night coverage before compensation, staffing, training, and access are ready.
22. Exercise command, regional resilience, and ledger recovery (depends on: 9, 13, 14, 15, 19)
Validate the process under realistic conditions before depending on it during a crisis. Use the same action-tracking rules for exercises and real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops covering payment failure, customer-first detection, third-party failure, and ambiguous ownership.
- Exercise loss of one AWS region, Kubernetes degradation, PostgreSQL failure, suspected duplicate payments, queue backlog, and settlement risk.
- Exercise simultaneous operational and security events to test command and disclosure boundaries.
- Exercise loss of chat, status-page, identity, or paging providers using telephone and offline fallbacks.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors.
- Validate backups, restoration, RPO, RTO, failover prerequisites, financial controls, and post-recovery reconciliation.
- Avoid uncontrolled production-ledger experiments; use staging, replicas, simulations, or tightly governed production tests.
- Complete at least two cross-company exercises before the audit, including one overnight or unannounced paging test.
23. Test SOC 2 operating effectiveness before the auditor (depends on: 4, 18, 21, 22)
Demonstrate that the controls operate consistently, not merely that policies exist. Correct failures through tracked remediation rather than rewriting historical records.
- Preserve policy approvals, service ownership, schedules, compensation activation, access reviews, training, certifications, incidents, communications, postmortems, actions, and exercises.
- Sample evidence monthly from initial signal through verified action closure.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Use exceptions to document missed acknowledgements, late communications, incomplete records, and compensating controls.
- Conduct a formal mock audit in month 7 using the populations and interview roles expected by the external auditor.
- Trace at least one SEV0 or SEV1, one SEV2, a customer-reported event, and an exercise end to end.
- Remediate evidence and operating gaps before external fieldwork.
- Brief commanders, responders, Support, and Compliance for auditor interviews without scripting inaccurate answers.
24. Institutionalize continuous improvement (depends on: 18, 21, 23)
Keep incident management active after the audit by assigning permanent owners, budgets, and review cycles. Use incident trends to drive architectural investment.
- Assign permanent owners for the policy, service catalog, paging platform, status page, training program, metrics, and evidence repository.
- Review severity thresholds, communications timing, staffing, and compensation annually and after material process failures.
- Recertify commanders and communications leads annually through observed exercises.
- Review recurring failure families quarterly and require executive action when remediation repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fatigue, fairness, tool friction, and psychological safety.
- Report severe incidents, SLA exposure, overdue risks, resilience investments, and customer-first detection to the board or risk committee quarterly.
- Prioritize reduction of the shared-ledger concentration risk, stronger regional independence, deployment safety, graceful degradation, and automated mitigation.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
Previous Proposal 3 (ID: 2437bc94-f488-46b5-866e-01e1d4afbf3b, Agent: qwen3.8-max_refine_3, LLM: alibaba/qwen3.8-max):
Estimated Complexity: high
Success Metrics: - Median time to detect customer-impacting incidents falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median mitigation time falls from 3h10 to under 90 minutes by day 120 and under 60 minutes by month 9.
- A named incident commander is assigned within 5 minutes in at least 95% of SEV1 and SEV2 incidents; zero incidents remain unowned for more than 10 minutes.
- First status-page update occurs within 15 minutes of SEV1 declaration and 30 minutes of SEV2 declaration in at least 95% of qualifying incidents.
- Monthly paging volume falls from 3,400 to under 500 actionable pages within six months, with alert actionability above 80%.
- Out-of-hours pages average no more than 2 per responder per week; sustained breaches trigger mandatory alert remediation.
- Six legacy alerting tools are consolidated into one paging and incident platform, with legacy paging paths disabled by week 16.
- 100% of Tier 0 and Tier 1 services have a named owning team, escalation path, dashboard, and runbook by day 60.
- All 28 teams are onboarded with readiness gates by week 22; every 24x7 critical-path rotation has at least six trained responders.
- 30 or more certified incident commanders and 20 or more certified communications leads provide continuous primary and secondary coverage.
- Paid on-call is approved by HR, legal, and finance and active in payroll before any new mandatory night rotation begins.
- 100% of required SEV1 and SEV2 postmortems are drafted within 3 business days and published within 10 business days in the standard template.
- Postmortem action closure rises from 17% to at least 90% of high-priority actions completed by their due date within two quarters.
- All 53 open historical actions are triaged within 30 days; high-risk unaccepted items are completed or formally risk-accepted within 90 days.
- Repeat incidents from a known unaddressed contributing factor decline by at least 50% within six months.
- Annualized SLA credits fall from $1.3M to under $400k within 12 months.
- Monthly availability meets or exceeds 99.95% by month 6, with exceptions reviewed at the executive reliability meeting.
- At least two cross-company exercises, including regional and ledger scenarios, are completed before the SOC 2 audit, with critical findings tracked.
- The month-6 internal dry-run audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete evidence.
- SOC 2 Type II incident-response controls pass with zero exceptions at the month-8 audit.
- On-call sentiment improves quarter over quarter; fewer than 10% of engineers report unwillingness to participate in their owned-service rotation by month 6.
- No increase in attrition among engineers on active rotations compared with baseline.
Steps (27):
1. Executive mandate, program funding, and governance
Convert the CEO email into a **written mandate within 48 hours**. Name one accountable program owner and a small decision group. Time-box design to five weeks so rollout starts well before the audit.
- Appoint the CTO as executive sponsor and a Head of Reliability or Incident Management as program owner with full-time authority.
- Create an 8–10 person design working group: SRE/platform lead, payments and ledger engineering managers, support lead, compliance, legal, HR, finance, and two rotating engineering managers.
- Approve budget lines: tooling consolidation, on-call compensation, training, coaching, and 2–3 dedicated program staff.
- State the non-negotiables: paid on-call, named service ownership, one severity scale, one paging platform, mandatory postmortems, and protected engineering capacity for reliability actions.
- Fix the master timeline: interim controls week 1, design weeks 1–5, pilot weeks 6–11, full rollout weeks 12–22, internal audit rehearsal month 6, audit-ready month 7.
- Publish a one-page charter to the company that says incident response is a company operating process, not an optional team practice.
2. Evidence baseline from incidents, alerts, and coverage gaps (depends on: 1)
Before changing anything, create an auditable **before picture** from the last 12 months. This baseline drives the severity design, staffing model, and executive reporting.
- Reconstruct all 31 customer-impacting incidents: detection source, first owner, severity, mitigation time, credits paid, and whether command was clear.
- Write a specific case review of the two incidents with unclear ownership for more than one hour.
- Inventory the six alert tools, alert volume by team, noise rate, rules with no owner, and rules with no runbook.
- Map the 16 teams without on-call and the 12 teams with unpaid on-call.
- Review the 53 open postmortem actions and triage the highest-risk items first.
- Freeze baseline metrics: 22-minute detection, 40% customer-first detection, 3h10 mitigation, $1.3M credits, 3,400 alerts, 85% noise, and 17% action closure.
3. Stakeholder listening, resistance mapping, and co-design group (depends on: 1)
Treat engineer pushback as **design input, not an attitude problem**. The objection to carrying a pager for other teams' code must shape ownership, routing, compensation, and command staffing.
- Interview leads from all 28 teams plus support, customer success, sales, compliance, and HR within two weeks.
- Separate the real causes of resistance: unpaid night work, unfamiliar systems, poor runbooks, unfair load, fear of blame, or unclear authority.
- Recruit 10–15 respected engineers and managers as co-designers so the model is built with the teams.
- Survey baseline on-call sentiment, trust in alerts, and psychological safety; repeat at 90, 180, and 365 days.
- Document the central promise: responders are paged for services they own, trained incident commanders coordinate, and all on-call work is compensated.
4. Service ownership catalog, criticality tiers, and dependency map (depends on: 2)
No alert can page the right team until every service has a **named owner**. Build the catalog as the routing foundation for on-call, severity impact mapping, status-page components, and audit evidence.
- Assign one accountable team to each of the 180 services, with an engineering manager, Slack channel, escalation policy, dashboard, and runbook link.
- Tier services: Tier 0 for money movement, ledger integrity, authentication, settlement, and shared PostgreSQL; Tier 1 for customer-facing degradable services; Tier 2 for internal or batch services; Tier 3 for non-critical services.
- Map customer journeys to services, databases, regions, third parties, and contractual SLA components.
- Mark orphan and shared services; require ownership, reassignment, or decommissioning within 30 days.
- Document cross-region dependencies, failover constraints, and services that defeat nominal regional redundancy.
- Treat missing ownership for Tier 0 or Tier 1 services as an executive escalation and a release-blocking risk.
5. Severity scale, declaration rights, and automatic triggers (depends on: 2, 4)
Adopt one severity scale so declaration is a lookup, not a debate. **Anyone may declare; only the incident commander may downgrade.** When uncertain, start higher.
- SEV1: money movement stopped or incorrect, ledger integrity in doubt, confirmed security or data event, both regions impaired, or broad SLA-credit exposure.
- SEV2: major degradation, settlement window at risk, one or more critical customers fully down, or likely contractual breach.
- SEV3: limited impact with a workaround, no financial-integrity or security risk.
- SEV4: no customer impact; handled by ticket during business hours.
- Define fixed triggers for each level: pages, command staffing, bridge call, status-page timing, account-manager outreach, executive notification, and postmortem obligation.
- Add automatic escalation: SEV3 open more than 2 hours becomes SEV2; any incident touching the shared ledger or unclear ownership becomes SEV2 unless the commander documents otherwise.
- Publish a decision tree and 10–12 worked examples from the actual 31 incidents.
6. Incident roles, authority, and handover rules (depends on: 5)
Solve the **nobody-was-in-charge failure** by making command explicit, trained, and transferable. Separate coordination from technical remediation so responders are not asked to debug unfamiliar code.
- Incident Commander: owns severity, priorities, escalation, mitigation strategy, role assignment, handoffs, and closure. Does not write code during the incident. May freeze deploys, invoke pre-approved failover, and pull any on-call responder.
- Communications Lead: owns status page, internal updates, account-manager briefs, and coordination with legal or compliance.
- Scribe: maintains the timestamped record of decisions, actions, role changes, and customer communications.
- Subject-matter responders: engineers from owning teams who diagnose and remediate within their accepted domain.
- Executive liaison: required for SEV1; shields the commander from executive questions and owns board or regulator escalation.
- Require the commander to claim the role within 5 minutes of SEV1 or SEV2 declaration, state it in the incident channel, and record every handover.
- Allow role combination only for SEV3 or SEV4; require separate people for command, communications, and primary technical work at SEV1.
7. Paid on-call, fatigue safeguards, and HR compliance (depends on: 3)
Unpaid on-call is a retention, fairness, and legal risk in New York. **Pay must be approved before any new mandatory rotation starts.** This is the fastest way to reduce pager resistance.
- Create a weekly stipend for primary and secondary on-call, differentiated by tier and night responsibility.
- Add per-page or per-incident pay for out-of-hours activation, plus guaranteed recovery time after night work or major incidents.
- Create a separate stipend for duty incident commanders and communications leads because command is a heavier burden.
- Verify FLSA, New York wage-hour, overtime, holiday, and payroll treatment with HR, legal, and finance.
- Cap rotation frequency: no more than one primary week in four or one week in six depending on staffing; no simultaneous primary assignments.
- Require at least six trained responders for any 24x7 rotation; fund hiring or service reassignment where teams are too small.
- Publish the compensation package and payroll start date before asking engineers to join rotations.
8. On-call staffing model and rotation rules (depends on: 4, 6, 7)
Do not force 28 identical night rotations. Use a **layered model**: central command coverage, critical-path team coverage, and business-hours coverage for lower-tier services.
- Company incident-command and communications rotation: 24x7 pool of 25–35 trained volunteers and designated senior staff, giving primary plus secondary coverage at all times.
- Critical-path teams: 24x7 primary and secondary on-call for Tier 0 and Tier 1 domains, including payments, ledger, authentication, API edge, Kubernetes platform, and shared PostgreSQL.
- Other teams: business-hours on-call with a documented night escalation list owned by the engineering manager.
- Group services into 8–12 coherent domains so rotations are sustainable; no rotation with fewer than six people is approved without an executive exception.
- Require handoff overlap, shadow shifts before first primary duty, and no on-call during approved PTO.
- Define cross-team pull rules: the commander may page another team's on-call with a 10-minute acknowledgement obligation; this is coordinated by command, not pushed onto responders.
- Publish schedules, swap rules, and load limits in the paging tool.
9. Alert quality standard, page budget, and noise controls (depends on: 2, 4)
With 3,400 alerts and 85% noise, detection fails because people stop trusting pages. Make alert quality a **condition of paging anyone**.
- Every page must have a named owning team, customer or SLO impact, severity, dashboard, runbook, expected action, and escalation policy.
- Page on customer-visible symptoms: payment success rate, latency, settlement deadlines, error-budget burn, ledger integrity, and replication health.
- Demote cause-based CPU, memory, or infrastructure-only alerts to dashboards or tickets unless they map to a customer journey.
- Set a page budget: maximum 2 out-of-hours pages per person per week; breach triggers a mandatory alert-tuning sprint.
- Auto-flag alerts that fire repeatedly without action or have high no-action acknowledgement rates.
- Require shadow mode for new alerts before they page humans, except for documented emergencies.
- Review alert quality monthly by team and publish a noise leaderboard.
- Target fewer than 500 actionable pages per month and above 80% actionability within six months.
10. Detection uplift across payments, ledger, and customer signals (depends on: 4, 9)
Customers detected 40% of incidents first. Detection must shift to **payment outcomes, ledger integrity, and inbound customer signals**, not host metrics alone.
- Define SLIs and SLOs for payment initiation, authorization, settlement, reconciliation, refunds, API availability, and reporting freshness.
- Run synthetic end-to-end payment tests from outside the platform in both AWS regions every 60 seconds.
- Add continuous ledger assurance: double-entry reconciliation, replication lag, failover readiness, disk pressure, and settlement-window countdown alerts.
- Monitor per-customer anomalies for top accounts so a single-tenant outage is detected before the account manager calls.
- Convert high-priority support tickets and account-manager reports into incident candidates within 5 minutes.
- Track detection source for every incident; make customer-first detection a reviewed defect with a corrective action.
- Add partner, banking, and card-network notification intake into the same declaration path.
11. Single incident platform and alert-tool consolidation (depends on: 5, 8, 9)
Collapse six tools into **one paging and incident-management system** with one queue, one timeline, and one audit trail. Do not create a parallel audit process.
- Select an integrated stack: paging and on-call scheduling, incident workflow, Slack or chat integration, bridge calling, status-page API, and ticketing integration.
- Implement one-command declaration that creates the incident record, channel, bridge, severity label, role prompts, and clock.
- Migrate alert sources by wave; retire a legacy paging path only after routing, ownership, and acknowledgement tests pass.
- Automate evidence capture: timestamps, acknowledgements, role assignments, severity changes, communications, and postmortem links.
- Integrate with the service catalog, on-call schedules, Jira or equivalent action tracker, customer-success tooling, and the status page.
- Verify the platform works during a single-region failure, including SMS, phone, and offline fallback paths.
- Set a hard date after which pages outside the chosen platform are not valid on-call obligations.
12. Escalation paths and five-minute command rule (depends on: 6, 8, 11)
Create one path from **signal to named commander in under five minutes**, any hour of the day. Make escalation automatic and time-bound.
- Accept declarations from automated alerts, engineers, support, account managers, partners, and customers through the same command.
- For SEV1 and SEV2, page the duty incident commander and owning-team primary immediately.
- Acknowledgement ladder: primary 5 minutes, secondary 10 minutes, manager 15 minutes, director or executive 20 minutes.
- If no commander claims the incident within 5 minutes, the platform assigns and announces one; the assignee may hand over but cannot leave the incident unowned.
- For SEV3, require team acknowledgement within 30 minutes; otherwise create a tracked work item.
- Give the commander pre-approved authority to invoke regional failover, ledger read-only mode, feature kill switches, and partner notifications without waiting for executive sign-off.
- Record every missed acknowledgement and escalation failure for weekly review.
13. Internal communications protocol (depends on: 6, 11)
Separate the **working incident channel from the audience channel** so responders can work and executives, support, and sales stay informed without interrupting the commander.
- Create one incident channel and one bridge per SEV1–SEV3 incident; use a read-only broadcast channel for executives and support.
- Set update cadence: every 15 minutes for SEV1, 30 minutes for SEV2, and at state changes for SEV3.
- Use a fixed update template: impact, customer-visible symptoms, current action, ETA or next update, commander, and communications lead.
- Brief support and customer success with a live affected-customer list and approved holding statements within 15 minutes of SEV1 or SEV2.
- Require executives to route questions through the executive liaison; the commander is not interrupted.
- Define fatigue rules and formal handover for incidents lasting more than 4 hours.
14. Customer status page, account-manager outreach, and SLA credit workflow (depends on: 5, 13)
Replace ad-hoc status updates with a **timed, owned, template-driven customer communication process**. Link incidents to SLA credits so finance and customer success are not surprised.
- Publish status-page updates within 15 minutes of SEV1 declaration and 30 minutes for SEV2; update every 30 or 60 minutes until resolution.
- Assign the communications lead as the single author; use pre-approved templates reviewed by legal and communications.
- Map status-page components to customer capabilities: payments, settlement, reporting, API, onboarding, and regional availability.
- For top accounts, require direct account-manager outreach within 30 minutes of SEV1 with an approved briefing pack.
- Send proactive email or webhook notifications to subscribed customers for SEV1 and SEV2.
- Publish a resolution notice within 30 minutes of mitigation and a customer-facing incident summary within 5 business days for SEV1.
- Automate SLA credit calculation from incident duration and affected capability; review with finance and legal within 5 business days.
- Track credits by incident and root-cause family to guide reliability investment.
15. Regulatory, legal, and partner notification playbook (depends on: 5, 14)
In payments, some incidents are reportable and the clock starts at detection. Make regulatory assessment a **mandatory step in the incident process**, not an afterthought.
- Map obligations: NYDFS cybersecurity-event rules, state breach laws, GLBA safeguards, PCI DSS if in scope, money-transmitter duties, card-network and sponsor-bank contracts, and cyber-insurer notice.
- Require a legal or compliance reportability assessment within 2 hours for every SEV1 and every security-related SEV2, even when the answer is not reportable.
- Maintain a 24x7 contact matrix for regulators, sponsor banks, card networks, outside counsel, insurers, and law enforcement.
- Pre-draft notification templates and preserve legal review.
- Encode customer-contract notification deadlines into account tiers and the communications workflow.
- Record every reportability decision, approver, deadline, and submission confirmation in the incident record.
16. Postmortem standard and blameless review (depends on: 5, 6)
Replace inconsistent postmortems with a **mandatory, blameless, single-format process**. The discipline comes from deadlines, facilitation, and action tracking.
- Require postmortems for all SEV1 and SEV2 incidents, customer-detected incidents, incidents over 2 hours, repeat failures, and ledger near-misses.
- Draft within 3 business days, peer review within 5, publish within 10 for SEV1 and SEV2.
- Use one template: summary, impact, timeline, detection analysis, response analysis, contributing factors, what worked, what failed, and action items.
- Make blamelessness explicit: focus on systems and decisions, not individual fault; never use postmortems in performance discipline.
- Hold a weekly incident review board to review postmortems, ratify severity, and challenge weak actions.
- Maintain a searchable postmortem library and quarterly recurring-cause analysis.
- Require a trained facilitator for major reviews; the incident commander attends but does not facilitate.
17. Action-item tracking, ownership, and delivery gates (depends on: 16, 11)
Only 11 of 64 actions were closed. Give postmortem actions the same status as **customer commitments**, with named owners and visible escalation.
- Create every action as a ticket with one named individual owner, priority, due date, and verification method.
- Use delivery classes: containment within 7 days, corrective work within 30 days, strategic work within 90 days.
- Reserve 15–20% of team sprint capacity for reliability and incident actions.
- Escalate overdue items: manager at 7 days, director at 14 days, CTO dashboard at 30 days.
- Block related feature releases when overdue P0 actions prevent recurrence of a severe incident.
- Require director approval and documented residual risk acceptance for overdue high-risk items.
- Verify effectiveness after completion; closing a ticket without evidence does not close the action.
- Target 90% of high-priority actions completed on time within two quarters.
18. Runbooks, critical-incident playbooks, and readiness bar (depends on: 4, 8)
Poor runbooks are a real cause of pager resistance and slow mitigation. Define a **minimum readiness bar** before a service is allowed to page anyone at night.
- Require for every Tier 0 and Tier 1 service: architecture summary, dependencies, dashboards, alert-to-runbook map, rollback procedure, feature flags, escalation contacts, and customer-impact statement.
- Write major playbooks for shared PostgreSQL ledger failure, regional failover, Kubernetes control-plane loss, payment-processor outage, settlement-window breach, duplicate-payment suspicion, and security compromise.
- Define ledger recovery rules: failover procedure, read-only degraded mode, reconciliation, RPO/RTO, and data-loss tolerance approved by executives.
- Test runbooks in drills at least twice a year; mark untested runbooks stale.
- Prevent paging alerts for services without readiness sign-off unless the engineering manager accepts the gap in writing.
- Keep runbooks linked from every alert and incident template.
19. Training, certification, and role readiness (depends on: 6, 12, 13, 16)
Command and communications are skills. Build a **tiered certification path** so rotations are staffed by people who have practiced, not by whoever is around.
- All employees: 1-hour module on declaring incidents, finding the incident channel, and reading the status page.
- All responders: half-day training on severity, acknowledgement, escalation, runbooks, and evidence hygiene.
- Incident commanders: 2-day course plus two shadowed incidents and one simulation before certification.
- Communications leads: training on status-page writing, customer language, account-manager briefs, and regulatory triggers.
- Scribes: training on timeline discipline and audit evidence.
- Certify for 12 months; renew through a simulation.
- Require shadow shifts before independent primary duty; no new hire holds primary within 90 days.
- Publish the certification register as an audit artifact.
20. Simulation program and game days (depends on: 19, 11, 18)
Rehearse the process before it meets a real SEV1. Simulations build commander confidence, expose runbook gaps, and produce audit evidence.
- Run monthly 60-minute tabletops using real incidents from the 31-incident baseline.
- Run quarterly game days covering regional failover, ledger replica promotion, dependency failure, partner outage, and security event.
- Run twice-yearly unannounced paging drills to measure night acknowledgement times.
- Include support, account managers, legal, compliance, and executives in at least one exercise per quarter.
- Produce tracked action items from every exercise using the same board as real incidents.
- Measure time to commander, time to first status update, and time to mitigation decision.
21. Critical-path pilot and gate review (depends on: 7, 10, 11, 14, 18, 19)
Prove the model on the highest-risk services with willing teams before full rollout. Run a **six-week pilot with daily feedback and public exit criteria**.
- Pilot with payments, ledger/database, platform/Kubernetes, API edge, authentication, and support intake.
- Activate severity scale, duty commanders, paid rotations, single paging platform, alert budget, status-page policy, and postmortem process.
- Hold a weekly pilot retrospective and fix process defects quickly.
- Validate night acknowledgement, cross-team pull response, severity clarity, and compensation payroll.
- Exit gate: commander assigned within 5 minutes in 95% of incidents, status page on time, page noise down at least 50%, postmortems on time, and positive on-call sentiment.
- Publish pilot results to the whole company as the main adoption argument.
22. Wave rollout across all 28 teams (depends on: 21)
Roll out by criticality and dependency, not by calendar alone. Use **readiness gates** so teams are not forced live without coverage.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1 teams. Wave 3: Tier 2 teams. Wave 4: Tier 3 and internal platforms.
- Gate per team: catalog entry complete, alerts migrated, runbooks ready, six-person rotation staffed, one commander candidate nominated, compensation in payroll, and one tabletop passed.
- Assign a program coach to each wave for three weeks.
- Freeze legacy paging paths for each team after successful onboarding.
- Publish a live adoption scoreboard by team.
- Complete all teams by week 22, leaving months of operating evidence before the audit.
23. Metrics, dashboards, and review cadence (depends on: 11, 16, 21)
Instrument the process itself. Leadership must see whether the program is working, and auditors must see **operating evidence, not retrospective paperwork**.
- Track response metrics: MTTD, time to declare, time to commander, acknowledgement time, mitigation time, resolution time, and customer-first detection rate.
- Track quality metrics: page volume, alert actionability, missed pages, postmortem timeliness, action closure rate, and status-page compliance.
- Track business metrics: availability against 99.95%, SLA credits, repeat incidents, and error-budget consumption.
- Track people metrics: on-call load, night pages per person, recovery time usage, sentiment, and attrition signals.
- Hold weekly incident review, monthly reliability review, quarterly executive review, and annual policy review.
- Publish dashboards internally and give every metric a target and owner.
24. Change management, incentives, and culture (depends on: 3, 7, 21)
The process will be judged on fairness. Communicate the deal repeatedly and make participation **recognized, compensated, and safe**.
- Core message: you are paid, paged only for what you own, supported by a trained commander, and given real capacity for actions.
- Run CTO all-hands, team roadshows, office hours, FAQ, and an internal incident-management hub.
- Add incident response and reliability work to promotion criteria and manager objectives.
- Recognize good postmortems, alert-noise reduction, and calm incident leadership.
- Provide a written path for engineers who cannot do nights; cover those shifts with paid volunteers or adjusted staffing.
- Prohibit retaliation for good-faith declaration or escalation.
- Publish sentiment survey results, including bad news, to maintain credibility.
25. SOC 2 evidence design and internal dry-run audit (depends on: 16, 17, 22, 23)
Design the process so audit evidence is a **by-product of normal operations**. Test it internally before the external auditor does.
- Map the process to SOC 2 criteria: incident identification, response, recovery, monitoring, communication, control activities, availability, and corrective action.
- Approve versioned policies: incident response policy, severity standard, on-call policy, communications policy, postmortem standard, and exception process.
- Retain incident records, paging logs, status-page history, postmortems, action tickets, training records, drill records, and access reviews for the audit period.
- Record regulatory reportability decisions even when no notification is required.
- Run an internal dry-run audit at month 6: sample at least 15 incidents and trace evidence end to end.
- Fix gaps with at least four weeks before the audit.
- Maintain an exception and remediation log instead of claiming perfection.
26. Program risk register and contingency planning (depends on: 1)
Name the likely failure modes now and pre-commit responses. Review the register monthly with the executive sponsor.
- Commander volunteer shortfall: roster command duty for senior engineers and managers until the pool reaches 25 trained people.
- Compensation delay: use immediate time-off-in-lieu plus phased stipend, but do not start mandatory night rotations without pay.
- Tool migration slip: preserve single paging and routing scope; defer advanced automation if needed.
- Alert pruning causing missed incidents: move noisy alerts to ticket first, observe 30 days, then delete.
- Burnout in experienced teams: monitor load weekly and cap pages per person.
- Major SEV1 during rollout: program owner shifts to incident support, wave schedule slips one wave, sponsor is informed same day.
- Shared ledger concentration: track blast-radius reduction and failover improvements as top strategic actions.
27. Continuous improvement, maturity roadmap, and post-audit sustainability (depends on: 23, 25)
Prevent the classic failure where the process decays after the audit. Build the second-year operating plan before the first year ends.
- Hold quarterly process retrospectives with commanders, responders, support, and communications leads.
- Re-baseline targets every six months; tighten goals once current targets are met.
- Move from lagging metrics to leading indicators: error-budget burn, near-miss rate, drill performance, and action ageing.
- Plan year-two improvements: follow-the-sun coverage, automated mitigation, ledger blast-radius reduction, error-budget release gates, and per-customer real-time impact reporting.
- Keep annual policy review, certification renewal, drills, and board reporting on a permanent calendar independent of the audit cycle.
- Report quarterly to the board or risk committee on availability, credits, severe incidents, overdue actions, and resilience investment.
Previous Proposal 4 (ID: 185429cf-5cd1-451d-ab7e-34d64ca403df, Agent: grok4.6_refine_4, LLM: xai/grok-4.6):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 10 minutes by month 3 and under 5 minutes by month 9.
- Customer-first detection falls from 40% to under 20% by month 3 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by month 4 and under 60 minutes by month 12.
- A named Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes after week 4.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases by month 3, with required update cadence met in 95% of cases.
- Monthly pages fall from 3,400 to under 500 within 6 months, with actionability above 75%; out-of-hours pages at or under 2 per person per week.
- All six legacy alerting tools route through one paging platform by week 16; legacy paging paths disabled per wave.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by due date within two quarters; all 53 currently open historical actions triaged within 30 days of S2.
- SLA credits fall from $1.3M to under $400k in the first 12 months; customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known unaddressed cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar; every Layer B rotation has 6+ certified responders or a time-limited executive exception.
- Paid on-call policy approved by HR, Legal, and Finance and in payroll before any mandatory night rotation starts.
- All 28 teams onboarded by week 24 with 24x7 command cover from 25+ certified ICs and 15+ certified Comms Leads.
- Internal dry-run at month 6 passes a 15-incident evidence walkthrough; month-7 mock audit finds no unowned high-risk control gap; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment improves quarter over quarter; no increase in attrition among engineers on rotation; pulse at 60 and 120 days shows at least 70% agree rotations are fair and limited to services they own.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked actions; two cross-company exercises completed before the audit.
Steps (24):
1. Program charter, mandate, and funding
Convert the CEO email into a named program with one owner, a budget, and a deadline earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, 28 teams will opt out.
- Appoint a Program Lead (Head of Reliability) with direct CTO and CEO sponsorship.
- Form a small steering group: CTO, VP Eng, Head of CS/Support, CISO, Legal, Finance, HR. Not a 28-team committee.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Timeline: operating floor in 7 days; design weeks 1–6; pilot weeks 7–12; rollout weeks 13–24; mock audit week 28; SOC 2 at month 8.
- Fund tooling ($150–250k/yr), on-call pay (~$600k–1M/yr), 2–3 program FTEs, and reserved engineering capacity. Anchor the ask against $1.3M in credits plus unmeasured incident cost.
- Freeze baselines now: 31 incidents, MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, $1.3M credits, 3,400 alerts/month at 85% noise, 11 of 64 actions closed.
2. Immediate 7-day operating floor (depends on: 1)
Do not wait for tooling, compensation, or the audit. Put a minimum process in place this week so the next outage has a named commander.
Start retaining artifacts on day one. This week becomes the first audit evidence.
- Publish a one-page interim severity guide and a single declaration path through Slack, phone, and pager.
- Staff interim primary and backup Incident Commander 24x7 from existing on-call veterans and engineering managers. Compensate this duty retroactively.
- Require a named IC within 10 minutes for every suspected major incident. If nobody claims it, the duty manager is IC.
- Use one channel, one bridge, one timeline doc, and one naming convention for every major incident.
- Direct Support to escalate credible customer reports immediately. They do not wait for engineering confirmation.
- Triage all 53 open historical actions. Close, re-plan, or formally risk-accept. Ledger integrity, payment duplication, regional resilience, security, and detection come first.
- Hold a daily 15-minute ops review until the permanent process is live.
3. Forensic baseline of incidents and alerts (depends on: 1)
Rebuild the facts before locking design. This is the before picture for the CEO and the auditor.
Every later design choice should trace to this evidence.
- Re-code each of the 31 incidents: trigger, service, detection source, timestamps, who led, credits paid, root-cause family.
- Quantify the 40% customer-first detections and name the missing signal in each case.
- Reconstruct the two nobody-in-charge incidents minute by minute. Use them as the burning-platform story.
- Audit the six alerting tools: volume per tool and team, top 50 noisy rules, rules with no owner or runbook.
- Freeze the baseline numbers. Do not let them drift during design.
4. Listening tour and the fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the main delivery risk. Treat it as a design constraint, not an attitude problem.
The answer that will actually stick is **the deal**: you are paid; you are paged only for services you own; a trained commander runs the room; your postmortem actions get real sprint capacity.
- Interview all 28 teams plus Support, CS, and Sales in two weeks.
- Separate the objections: unpaid work, nights, unfamiliar code, bad runbooks, fear of blame. Each needs a different fix.
- Collect informal practices from the 12 teams already on-call. They are the pilot candidates and the veteran pool.
- Recruit 10–15 credible engineers as a design working group so the process is co-authored.
- Baseline sentiment on alert trust, on-call willingness, and burnout. Re-measure at 6 and 12 months.
5. Service ownership catalog and criticality tiers (depends on: 3)
You cannot page the right person across 180 services until each one has a named owner. This is the foundation of fairness, routing, and audit evidence.
Build a machine-readable catalog as the single source of truth.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, the shared PostgreSQL cluster, queues, partner banks, and customer-facing endpoints.
- Assign one owning team, named engineering manager, Slack channel, escalation policy, and dependency list per service.
- Tier 0: money movement, ledger, auth, shared Postgres, regional control plane. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map Tier 0/1 services to customer capabilities: initiation, authorization, settlement, reporting, onboarding.
- Assign coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
6. Severity scale and declaration rules (depends on: 3, 5)
Replace in-the-moment debate with a lookup. Four payments-specific levels, with worked examples drawn from the 31 real incidents.
Anyone may declare. Only the Incident Commander may downgrade, with recorded rationale. When unsure, start high.
- **SEV1**: money movement stopped or incorrect; ledger integrity in doubt; material security or data exposure; both regions impaired; or more than 10% of customers impacted. Pages IC, comms, scribe, SMEs, and exec liaison. Bridge in 5 minutes. Status page in 15 minutes. Mandatory postmortem and regulator assessment.
- SEV2: severe degradation; settlement window at risk; a strategic customer fully down; SLA breach likely. IC and SMEs paged. Status page in 30 minutes. Mandatory postmortem.
- SEV3: partial impact with a workaround; no credit exposure. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, or a repeat.
- SEV4: internal or minor. Ticket only. No page.
- Auto-escalate: any SEV3 open more than 2 hours, or any incident touching the shared ledger, becomes SEV2. Unknown impact after 15 minutes is raised, not sat on.
- Nobody is punished for over-declaring. Publish that rule in writing and repeat it.
7. Roles, authority, and ledger dual-control (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, transferable, and logged.
The Incident Commander owns the incident, not the fix, and **never types** in a production terminal.
- Incident Commander: declares severity, pulls anyone, freezes deploys, invokes failover, authorises spend. One IC at a time. Assumed within 5 minutes and announced in channel.
- Communications Lead: single voice to customers, status page, account managers, and the exec summary.
- Scribe: timestamped timeline, decisions, and open questions. Feeds the postmortem and the audit trail. Required for SEV1 and SEV2.
- Subject-matter responders: diagnose and mitigate only services they own, with access and runbooks.
- Executive Liaison (SEV1): shields the IC from exec questions; owns regulator and board escalation.
- The IC may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access rules.
- Handover is verbal and written, with time and new owner recorded. Roles may combine below SEV2; never at SEV1. The IC stays in charge if a VP joins.
8. Three-layer 24x7 coverage model (depends on: 5, 7)
Do not put 28 teams on night rotation. That is what engineers are rejecting.
Staff command centrally. Page engineers only for code their team owns.
- Layer A, company command: Duty IC plus secondary, Duty Comms, and a scribe pool. Target 25–35 certified ICs and 15–20 comms leads. Roughly one week per person every 6–8 months. Five-minute ack SLA, then secondary, then on-call director.
- Layer B, Tier 0/1 domains: group 28 teams into 8–12 product and platform domains (ledger app, Postgres platform, payments orchestration, auth, API edge, Kubernetes/platform, settlement). Primary plus secondary 24x7. Minimum six trained people. Target one week in six, never worse than one in four.
- Layer C, Tier 2/3: business-hours on-call. After hours the IC pages the EM, who holds a written escalation list.
- No engineer joins another team's responder pool without training, access, runbooks, shadow shifts, and both teams' acceptance.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner.
- Evaluate follow-the-sun coverage as a 12-month option, not a year-1 dependency.
9. Paid on-call, NY labor compliance, and fatigue rules (depends on: 4, 8)
Unpaid on-call is a retention risk and a New York legal exposure. Pay must be live in payroll before any mandatory night rotation starts.
Treat interrupted nights and recovery time as compensable work.
- Weekly stipend by layer, higher for 24x7 Tier 0/1 and Duty IC, lower for business-hours, with holiday and weekend premiums.
- After-hours call-out pay or TOIL. A paid recovery day after prolonged overnight work, a SEV1, or a qualifying SEV2. Managers cover the next day.
- HR, Legal, and Finance publish dollar amounts, FLSA exempt/non-exempt treatment, NY wage-hour rules, tax treatment, and payroll timing within 14 days of charter.
- Load rules: no primary on two rotations; no consecutive primary weeks; no on-call the week after a SEV1 you commanded.
- A person may declare temporarily unfit after overnight work with no performance penalty.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- If a rotation cannot staff six people, merge domains or hire. Do not run two-person 24x7.
- Model annual cost against $1.3M in credits and get it as a CFO/board line item.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts a month at 85% noise is why detection takes 22 minutes. Make quality a condition of paging a human.
A page is a product with a quality bar, not a dump of host metrics.
- Every paging alert must have a named owner, customer or SLO impact, runbook, tested threshold, severity mapping, dashboard, expected action, and dedup key. Fail any of these and it becomes a ticket or is deleted.
- Page on customer symptoms: SLO burn, error budget, settlement-queue age. Cause-based CPU and memory alerts become dashboards or tickets.
- Set a **page budget** of at most two out-of-hours pages per person per week. A breach triggers a mandatory tuning sprint and blocks new paging alerts for that team.
- Auto-quarantine alerts that fire more than five times a month with no action, or that have more than 70% no-action acks. Never silently disable without compensating detection and a recorded decision.
- Run new alerts in shadow for seven days unless an emergency exception is approved.
- Monthly per-team kill, tune, or keep review. Target under 500 pages a month and actionability above 75% within six months.
11. Customer-journey detection and ledger assurance (depends on: 5, 10)
Stop customers being the monitoring system. Detect on money-path outcomes, not host metrics.
Every postmortem will ask why a customer saw it first.
- Define SLOs per Tier 0/1 capability: initiation success, auth latency, settlement timeliness, API availability, reporting freshness. Set an internal target stricter than 99.95%.
- Run synthetic full-lifecycle payments from outside the platform, in both regions, every 60 seconds. Include a small real-value canary where legally feasible.
- Add ledger assurance: continuous double-entry reconciliation, replication lag, failover-readiness, unexpected balances, duplicate identifiers, settlement-window countdown.
- Add per-customer anomaly detection for the top 100 accounts.
- Auto-create a triage incident within five minutes when support tickets or AM reports match impact keywords.
12. Single incident and paging platform (depends on: 6, 8, 10)
Collapse six alerting tools into one queue, one timeline, and one audit record.
The platform must still work if one AWS region is down. Verify SMS and phone fallback, plus an offline runbook.
- Select one paging product, one incident layer, and one hosted status page.
- One Slack command declares the incident, creates the channel and bridge, pages the Duty IC, sets severity, and starts the clock.
- Ingest the six existing tools first. Deduplicate and route. Retire a legacy path only after named owners and two weeks of verified operation.
- Auto-capture timestamps, acknowledgements, role assignments, severity changes, comms, mitigated time, and resolved time. Export for SOC 2.
- Integrate the service catalog, Jira, Salesforce or CS tooling for affected-customer lists, and Zoom or Slack Huddle.
- Test paging, escalation, status publication, and conference access every week.
- After a team's wave, old paging paths are disabled, not left as a fallback.
13. Detection-to-command escalation path (depends on: 7, 8, 12)
Write the single path from something looks wrong to someone is in charge. Target: a commander in under five minutes, any hour.
If nobody claims IC in five minutes, the platform assigns the Duty IC. The assignee can hand over, not decline.
- Converge every entry point on the same declare command: alert, engineer, support, account manager, partner bank, SEV hotline.
- SEV1: page the owning primary immediately; secondary at 5 minutes unacked; domain manager and Duty IC at 10; exec liaison at 15.
- SEV2: primary ack in 10 minutes; IC assigned in 15.
- Human acknowledgement is required. Delivery to a device does not count.
- The IC can page any team's on-call, with a 10-minute ack obligation. This reciprocity makes single-team ownership viable.
- Unowned alerts go to Layer A command, then the missing owner record is a control defect.
- Pre-authorise regional failover, ledger read-only mode, and partner-bank notice so the IC does not wait for an executive. Dual-control still applies to ledger writes.
14. Live execution and major-incident playbooks (depends on: 7, 13)
Limit customer and financial harm before proving root cause. One procedure from the first minute to handback.
A service cannot page at night until it meets the readiness bar.
- Open channel, bridge, record, and timeline immediately for SEV1 and SEV2. The IC states severity, known impact, hypothesis, objective, roles, and next update time.
- Freeze unrelated production changes during SEV1. Record exceptions the IC approves.
- Prefer reversible mitigation: rollback, feature flag, traffic isolation, rate limit, partner reroute.
- Write playbooks first for Postgres ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank outage, settlement-window breach, suspected security compromise, and suspected duplicate payments.
- Guard against split-brain, replay, and duplication during regional or database recovery. Reconcile and process backlog before calling the incident resolved.
- Mitigation means customer impact ended. Resolution means stable, backlog processed, and ledger reconciled.
- Formal IC handover after 4 hours. Reopen if impact recurs during the stability window.
- Readiness bar per Tier 0/1 service: diagram, dependencies, dashboard, runbook, rollback, kill switch, escalation contacts, RPO/RTO. Untested runbooks are marked stale.
15. Internal, customer, and regulatory communications (depends on: 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved process. The Communications Lead is the single author.
State impact and the next update time. Never speculate on cause.
- Internal: one working channel, one bridge, one read-only broadcast for execs, Support, and Sales. SEV1 update every 30 minutes even if unchanged; SEV2 every 60 minutes.
- Executives ask questions only of the Executive Liaison. Publish this as a signed exec behaviour rule.
- Status page: SEV1 within 15 minutes, SEV2 within 30 minutes; then 30/60-minute updates; resolve notice within 30 minutes of mitigation. Templates pre-approved with Legal.
- Top 100 accounts: named AM contact within 30 minutes of SEV1, with a briefing pack from Comms. Long tail gets the status page plus email or webhook.
- Customer-facing summary within 5 business days for SEV1.
- Mandatory regulatory checkpoint on every SEV1 and every security SEV2, within 2 hours, recorded even when not reportable. Cover NYDFS Part 500 72-hour clock, state breach laws, GLBA/FTC, PCI if in scope, sponsor-bank and card-network windows, FinCEN/OFAC if relevant.
- Legal owns outbound regulatory letters. The IC owns facts. Security or Legal may limit public detail during an active threat, with the reason recorded.
- Encode bespoke customer-contract notice SLAs into account tiering.
16. SLA credit and financial-impact workflow (depends on: 6, 15)
Link incidents to money so severity, credits, and investment stay consistent. Finance should not learn about outages from invoices.
Make credit calculation an output of the incident record, not a negotiation.
- Agree availability measurement per contract and component with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and capability telemetry. Produce a proposed credit schedule within 5 business days of resolution.
- Posture: proactive credits for top-tier accounts; claims-based for the rest. Document the approval chain.
- Track credits per incident and root-cause family. Quarterly report which reliability investments would have prevented which credits.
- Target: cut credits from $1.3M to under $400k in 12 months. Use that delta as the ongoing business case.
17. Blameless postmortems and action tracking (depends on: 6, 12)
Eleven of 64 actions closed is the clearest process failure. Make learning mandatory and actions as binding as customer commitments.
Closing a ticket without evidence of effectiveness does not close the action.
- Mandatory for every SEV1 and SEV2, every customer-first detection, every incident over 2 hours, every repeat of a known cause, and every ledger near-miss.
- Draft within 3 business days, review within 5, published internally within 10. The IC owns the draft. The owning EM is accountable.
- One template: timeline, customer and financial impact, why detection was late, why mitigation took that long, contributing factors, what went well, actions.
- Blameless in writing. No individual named as a cause. HR and management commit that postmortems are never used in performance reviews.
- Every action gets a named person, priority, due date, Jira ticket, and verification method. P0 (prevents SEV1 recurrence) due in 30 days and committed into the next sprint before roadmap work. P1 in 60 days. P2 in 90 days.
- Teams reserve 15–20% of sprint capacity for reliability and incident actions.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Weekly Incident Review Board with engineering directors. Target 90% of P0/P1 actions closed on time within two quarters.
18. Training, certification, and commander academy (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Nobody holds independent duty untrained. Training happens in paid working time.
The certification register is an audit artefact.
- All employees, 1 hour: recognise impact, declare, find the channel and status page.
- Responder, half day: severity, escalation, runbooks, comms hygiene. Mandatory before joining a rotation.
- Scribe, 2 hours: timeline discipline. The entry point.
- Incident Commander, two days plus shadowing: command presence, decisions under uncertainty, severity, handover, exec management. Two shadowed incidents and one simulated SEV1 before certification.
- Communications Lead, one day: status writing, customer tiering, legal boundaries, regulator triggers.
- Certification valid 12 months, renewed via simulation.
- New joiners shadow two shifts and do not hold primary in their first 90 days.
- Appoint a champion in each of the 28 teams.
19. Simulations and game days (depends on: 14, 15, 18)
Rehearse before the next real SEV1. Use a standing calendar, not a one-off exercise.
Do not inject uncontrolled changes into the production ledger.
- Monthly 60-minute tabletop per engineering group, using a real incident from the 31.
- Quarterly full-scale game day: regional failover, ledger replica promotion, dependency failure. Whole role structure, timed.
- Twice-yearly unannounced paging drill, including nights, to measure real acknowledgement times.
- One security incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Also exercise status-page failure and loss of the primary chat or pager.
- Use replicas, staging, or tightly governed tests for ledger scenarios.
- Every exercise produces actions in the same tracker as real incidents.
- Complete two cross-company exercises before the SOC 2 audit.
20. Critical-path pilot (depends on: 9, 11, 12, 18, 19)
Prove the process on the highest-risk surface with willing teams before asking 28 teams to adopt it.
Publish a one-page result company-wide. That result is the adoption argument.
- Six-week Wave 0: ledger app, Postgres platform, payments orchestration, Kubernetes/platform, API gateway, Support intake, plus two of the 12 teams already on-call.
- Activate the full stack: severity scale, Duty IC, one pager, page budget, status page, mandatory postmortems, paid on-call.
- Parallel-run old paths for one week, then cut over. The program lead coaches every SEV2+ and does not secretly take command.
- Weekly retro. Expect 20–40 process defects and fix them in the standard before rollout.
- Exit gate: MTTD under 10 minutes on pilot services; IC assigned within 5 minutes in 95% of incidents; page volume down 50%; all postmortems on time; no unpaid pages; sentiment not worse.
21. Metrics, reviews, and error budgets (depends on: 12, 17, 20)
Instrument the process itself. Do not reward hiding incidents or suppressing pages.
Every metric has a target and a named owner. Dashboards are public inside the company.
- Response: MTTD, time to declare, time to IC, MTTA, MTTM, MTTR, incidents by severity, percent customer-first. Report median and p90, by tier, journey, and region.
- Quality: pages per person per week, actionability, budget breaches, postmortem on-time rate, action closure and age, missed acknowledgements.
- Business: credits, 99.95% per capability, error-budget burn, failed-payment volume, reconciliation breaks, repeat-incident rate.
- People: rotation size, frequency, after-hours pages, recovery days, sentiment, attrition among on-call staff.
- Cadence: weekly Incident Review Board; monthly Reliability Review; quarterly exec and board review; annual policy review.
- Error budgets on Tier 0/1 SLOs. Burn too fast and the team pauses features to pay down reliability.
- Reconcile dashboards monthly against a sample of incident records and customer cases so missing incidents cannot hide.
22. Wave rollout with readiness gates (depends on: 20, 21)
Roll out in four waves by criticality, every three weeks. Gates keep the standard credible. A missed gate is rescheduled, not waived.
Finish all 28 teams by week 24 so roughly three months of operating evidence remain before audit fieldwork.
- Wave 1: remaining Tier 0. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team kit: catalog complete, alerts migrated and inside budget, runbooks at the readiness bar, rotation of 6+ certified responders, one IC nominee, one drill passed, paid rotation in payroll.
- Named coach for three weeks. The director signs the gate.
- Freeze legacy paging per wave.
- Publish a live adoption scoreboard.
- Any production service that cannot provide sustainable ownership gets an executive-reviewed deadline, compensating control, and expiry date. No indefinite verbal exceptions.
23. Change management, incentives, and the deal (depends on: 1, 4, 9)
Run this from day one in parallel with design. Engineers will judge fairness. Executives will judge visible results.
Repeat the deal until it is muscle memory: paid on-call, paged only for what you own, trained commander, real sprint capacity for actions.
- Launch via CTO all-hands, per-team roadshows, a one-page laptop card, an internal wiki, and a Slack help channel with a 4-hour answer SLA.
- Put incident-response contribution into promotion criteria. Award best postmortem and biggest noise cut quarterly. Thank people publicly after every SEV1.
- Put adoption, alert hygiene, action closure, and on-call load fairness into every engineering manager's quarterly objectives.
- Write an exception path for engineers who cannot do nights because of caring responsibilities or health, covered by stipended volunteers.
- Prohibit retaliation for good-faith declaration or escalation.
- Pulse-survey at 60 and 120 days. If fairness or load is red, pause expansion until fixed.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
24. SOC 2 evidence, mock audit, and sustainability (depends on: 17, 21, 22)
Evidence is a by-product of doing the work, not a reconstruction the week before the auditor. Protect the process after SOC 2 is signed.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to CC7.2–CC7.4, CC2.2/CC2.3, CC5, and availability A1.2. Confirm interpretation early.
- Maintain versioned, signed policies for incident response, severity, on-call, communications, and postmortems, reviewed annually.
- Automate evidence: incident records, paging and ack logs, status history, postmortem library, action closure, training register, drill records, and reportability decisions including not reportable.
- Internal dry-run at month 6 on a 15-incident evidence walkthrough. Mock audit at month 7. Fix with weeks to spare.
- Assign permanent owners for policy, pager, status page, catalog, training, and metrics.
- Year-two roadmap: follow-the-sun cell, self-healing for the top recurring causes, error-budget release gates, and blast-radius reduction on the shared ledger.
- Quarterly board summary of severe incidents, credits, overdue P0s, and resilience investment so attention does not die after the audit.
Previous Proposal 5 (ID: 0ab2b1c8-46af-4001-ab14-0c074d30a026, Agent: deepseek-v4-pro_refine_5, LLM: deepseek/deepseek-v4-pro):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 6 months of full rollout.
- Customer-first detection falls from 40% to below 10% within 6 months and below 5% within 12.
- Median time to mitigate SEV0/SEV1 falls from 3h10 to under 60 minutes within 12 months; SEV2 under 2 hours.
- Incident Commander assigned and announced within 5 minutes for 95% of SEV0/SEV1 incidents; zero incidents with unclear command beyond 10 minutes.
- Status page updated within policy time for 95% of SEV0/SEV1/SEV2 (10/15/30 minutes).
- Monthly pages fall from 3,400 to under 500 with actionability above 75%.
- All six legacy alerting tools decommissioned by week 16.
- 100% of SEV0/SEV1 incidents have a blameless postmortem published within 10 business days.
- Postmortem action closure rises from 17% to 90% for P0/P1 actions on time.
- SLA credits fall from $1.3M to under $400K in first 12 months.
- Customer-impacting incidents decline to under 15 per year; repeat root causes under 10%.
- 100% of 180 services have a named owning team and criticality tier.
- All 28 teams onboarded by week 24; every Tier 0/1 team has 24x7 primary+secondary coverage with 6+ certified responders.
- At least 30 certified Incident Commanders and 20 certified Communications Leads active.
- Paid on-call policy is approved and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among on-call engineers.
- Quarterly game days and twice-yearly unannounced paging drills executed on schedule, each with tracked action items.
Steps (26):
1. Executive mandate, program governance, and interim incident command
Convert the CEO's outage complaint into a company-level improvement program with a named owner, budget, and deadlines. In the first week, establish an interim command process so no incident remains unowned while the permanent process is designed.
- Appoint Head of Reliability as program owner and CTO/CEO as executive sponsor.
- Form steering group with Engineering, SRE, Support, CS, Legal, Compliance, HR, Finance, Security.
- Approve non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Fund tooling, on-call compensation, training, and 3 dedicated program FTEs; anchor to $1.3M credits.
- Set timeline: design weeks 1-4, tooling and pilot weeks 5-10, rollout weeks 11-20, audit evidence collection from week 8, dry-run month 6.
- Stand up interim 24x7 duty officer and single declaration path within 48 hours, compensated retroactively.
2. Baseline data and alert estate analysis (depends on: 1)
Re-open the last 31 incidents and profile the current alert estate so every later design decision is evidence-based.
- Re-code each incident: detection source, timestamps, owner, severity, credits, root-cause family.
- Quantify customer-first detections and missing signals.
- Analyze two 'nobody in charge' incidents minute-by-minute.
- Inventory six alerting tools: volume, noise, owner, runbook coverage, top 50 noisy rules.
- Freeze baseline metrics: MTTD 22m, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed.
3. Stakeholder listening and resistance mapping (depends on: 1, 2)
Treat engineer pushback as design input. Interview all 28 teams plus Support, CS, and Sales to understand real objections and find early adopters.
- Test objection: unpaid work, night load, unfamiliar code, poor runbooks, or blame.
- Document current informal practices from 12 on-call teams.
- Recruit 10-15 credible engineers as co-design working group.
- Survey baseline sentiment on on-call and alerts.
- Publish promise: paid on-call, page only for owned services, trained commander, reserved reliability capacity.
4. Service ownership catalog and criticality tiering (depends on: 2, 3)
Build a machine-readable service catalog that assigns each of the 180 services a single owning team, escalation policy, and criticality tier.
- Define fields: owner team, manager, Slack channel, escalation policy, dependencies, dashboards, runbooks.
- Tier 0: ledger, money movement, auth, shared PostgreSQL; Tier 1 customer-facing; Tier 2 internal; Tier 3 non-critical.
- Map customer-visible capabilities and dependencies across regions.
- Identify orphan and shared services; force ownership decision within 30 days or schedule decommissioning.
- Publish coverage gaps by tier; Tier 0 gaps become executive escalations.
5. Define severity levels and declaration triggers (depends on: 4)
Adopt a five-level severity model with objective payment-specific triggers, so declaration is a lookup rather than debate. Anyone may declare; only the Incident Commander may downgrade.
- SEV0: unauthorized, lost, duplicated, or corrupted money movement; ledger integrity loss; confirmed data breach; both regions failed. Page all roles, exec, legal, and consider payment pause.
- SEV1: widespread payment failure, no workaround, severe SLA breach, or single-region total loss. Full role activation and public status page.
- SEV2: significant degradation, multiple customers, or workaround available but SLA risk. IC and SMEs paged; status page if customer-visible.
- SEV3: limited impact with workaround; team-led, business-hours response.
- SEV4: internal/no customer impact; ticket only.
- Auto-escalation: unresolved SEV3 >2h becomes SEV2; unresolved SEV2 >1h becomes SEV1; ledger or security always at least SEV2.
- Provide decision tree and 12 worked examples from actual incidents.
6. Define incident roles, decision authority, and handover rules (depends on: 5)
Codify five roles with written responsibilities and explicit authority, so there is never confusion about who is in charge.
- Incident Commander: owns severity, priorities, cross-team pulls, mitigation decisions; does not code.
- Communications Lead: owns status page, internal updates, account manager briefs, regulator coordination.
- Scribe: maintains timeline, decisions, and evidence for postmortem/audit.
- Subject-Matter Responders: diagnose and remediate only their owned services.
- Executive Liaison (SEV0/SEV1): handles exec communications and external stakeholders.
- Rule: IC identified within 5 minutes and announced in channel; handover announced and logged; roles may combine below SEV2, never at SEV0/SEV1.
7. Design 24x7 incident command and comms staffing model (depends on: 6, 4)
Create a central, trained incident command rotation instead of making each of 28 teams field its own commander.
- Recruit 24-30 certified ICs and 12-16 Comms Leads from a cross-team volunteer pool with manager approval.
- Weekly rotations, primary and secondary; 5-minute acknowledgement SLA with auto-escalation.
- Scribe pool as entry-level rotation.
- Night coverage: paid command rotation in New York timezone initially; evaluate follow-the-sun coverage later.
- Eligibility: certification required; commanders can leave with 30 days notice.
- Ensure distinct persons for IC, CL, and primary SME on SEV0/SEV1.
8. Design team on-call rotations and cross-team escalation policies (depends on: 4, 6)
Define tiered team on-call obligations so engineers are only paged for services they own, and cross-team pages go through the IC.
- Tier 0/1 services: 24x7 primary+secondary, at least 6 trained responders per rotation, one-week shifts.
- Tier 2/3: business-hours on-call; after-hours escalation list to engineering manager.
- Platform, infrastructure, database: 24x7 due shared ledger and Kubernetes.
- Routing: every page resolves via service catalog to owning team's escalation policy; cross-team pages only by IC.
- Guardrails: max one primary week in four, no consecutive weeks, no on-call in week after leading SEV0/1, protected recovery time after night work.
9. Design paid on-call compensation and fatigue safeguards (depends on: 8, 3)
Make on-call paid and legally compliant before any new rotation starts, and convert unpaid pager culture into a fair employment condition.
- Weekly stipends: primary 24x7 $800-1,200, secondary 30-50%, business-hours $300-500; holidays premium.
- Out-of-hours incident pay: $150 per night page plus hourly beyond one hour; time-off-in-lieu after overnight work.
- Additional command rotation stipend and SEV0/1 bonus for active responders.
- Verify FLSA/NY wage rules with HR, Legal, Finance; document exemption treatment.
- Base annual budget against current $1.3M credits.
- Track load and trigger staffing review if >2 after-hours pages per person per week sustained.
10. Enforce alert quality standards and noise budget (depends on: 2, 4, 8)
Replace 3,400 monthly alerts at 85% noise with a contractual paging standard that makes on-call sustainable.
- Every paging alert must have owning team, customer impact statement, runbook link, severity mapping, and tested threshold.
- Page only on customer-visible symptoms or SLO burn; cause-based alerts become tickets or dashboards.
- Set page budget: max 2 out-of-hours pages per person per week; breach triggers mandatory alert-tuning sprint.
- Auto-quarantine alerts with >5 firings/month without action or >70% no-action acknowledgements.
- Target <500 actionable pages/month and >75% actionability within 6 months.
- Weekly per-team alert review, monthly cross-team review.
11. Build detection uplift: synthetics, SLOs, and support intake (depends on: 4, 10)
Shift detection from host metrics to customer outcomes so the company stops hearing about outages from clients first.
- Define SLOs per Tier 0/1 capability: payment initiation, auth, settlement timeliness, API availability, ledger consistency.
- Deploy external synthetic transactions from both regions every 60 seconds, covering full payment flow and ledger write.
- Add ledger assurance checks: replication lag, double-entry balance, settlement window countdown.
- Top-100 customer anomaly detection to catch single-tenant outages.
- Auto-create triage incident from support tickets or account manager keywords within 5 minutes.
- Track customer-detected-first as a defect and require a postmortem action.
12. Consolidate alerting and incident tooling (depends on: 5, 8, 10, 11)
Collapse six alerting tools into one integrated paging and incident management platform to create a single system of record for people and audit.
- Select paging/on-call platform and incident management layer (e.g., PagerDuty + incident.io/FireHydrant).
- Implement one-command Slack declaration that auto-creates channel, bridge, pages roles, sets severity, starts timeline.
- Migrate all monitoring sources into the one tool; decommission legacy paging only after two weeks verified.
- Integrate service catalog, status page, Jira action tracking, Salesforce/CS customer lists, and conference bridge.
- Ensure out-of-band paging and offline fallback if a region or chat tool is down.
- Automate evidence capture for SOC2: timestamps, role assignments, severity changes, comms sent.
13. Define acknowledgement and escalation paths (depends on: 12, 6, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts at first signal.
- For SEV0/SEV1: page primary on-call; after 5 minutes unacked page secondary; at 10 page manager and duty IC; at 15 page executive officer.
- For SEV2: primary ack within 10 minutes, IC assigned within 15; escalate on miss.
- For SEV3: ack within 30 minutes or create work item.
- Automatic duty IC page for ledger/security, cross-team, or unresolved ownership.
- If impact unknown after 15 minutes, raise severity.
- Cross-team responders summoned by IC have 10-minute acknowledgement obligation.
- Route alerts with no owner to command rotation, then treat missing ownership as control defect.
- Human acknowledgement required; delivery confirmation is not sufficient.
14. Standardize internal communications (depends on: 6, 5)
Separate the working war room from executive and stakeholder updates, with fixed cadence and pre-approved templates.
- Auto-create incident channel and one read-only broadcast channel for execs, support, sales.
- SEV0: internal update every 15 minutes; SEV1 every 30; SEV2 every 60; SEV3 on state change.
- Template: impact, what we know, what we are doing, ETA/next update, current IC and CL.
- Executives questions through Executive Liaison only; IC is not interrupted.
- Support/CS receives affected-customer list and holding statement within 15 minutes for SEV0/SEV1 and 30 for SEV2.
- Handover protocol for incidents lasting >4 hours: formal IC handover and fatigue check.
15. Standardize customer and status page communications (depends on: 5, 14)
Replace ad-hoc status page updates with a timed, role-owned, template-driven process, including account manager outreach.
- Status page timing: SEV0 initial post within 10 minutes, SEV1 within 15, SEV2 within 30; updates every 15/30/60 minutes until resolved.
- Resolution notice within 30 minutes of mitigation; customer-facing summary in 5 business days for SEV0/SEV1.
- Pre-approve 12-15 templates with Legal and Comms.
- Tiered outreach: top 100 accounts get direct call/email from AM within 30 minutes of SEV0/SEV1; long tail via subscription.
- Use factual language: state impact and next update; never speculate cause or blame vendor.
- Comms Lead is sole author for customer language.
16. Regulatory, legal, and account manager notification playbook (depends on: 5, 15)
Build a notification decision tree and contact matrix so legal/regulatory obligations are assessed early and never forgotten.
- Map obligations: NYDFS Part 500 72-hour cybersecurity event notification, state breach laws, GLBA/FTC, PCI, sponsor bank/card network contractual windows, FinCEN/OFAC if relevant.
- Add regulatory assessment checkpoint for every SEV0 and security SEV1 within 2 hours, even if not reportable.
- Maintain 24x7 contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel with named backups.
- Pre-draft notification templates and test quarterly.
- Encode customer-specific notification SLAs from enterprise contracts into customer tiering.
- Account managers receive legal-approved script and affected-customer list.
17. SLA credit workflow and financial impact measurement (depends on: 5, 15)
Link every incident to money automatically so severity, credits, and prioritization stay consistent and Finance is never surprised.
- Define availability measurement per contract and capability with Legal/Finance.
- Auto-compute affected minutes per customer from incident record and telemetry; generate credit proposal within 5 business days.
- Decide proactive credits for top tier vs claims-based for others; document approval chain.
- Track credits by root-cause family and incident; feed quarterly reliability investment decisions.
- Target reducing annual credits from $1.3M to under $400K in first year.
18. Standardize blameless postmortems (depends on: 5, 6)
Make postmortems mandatory with fixed deadlines, a single format, and a blameless review forum, replacing 'various formats'.
- Mandatory: all SEV0/SEV1; SEV2 with customer impact, credits, >2h, repeat cause, or detected by customer; near-miss involving ledger.
- Draft within 3 business days, peer review within 5, publish within 10.
- One template: timeline, customer/financial impact, detection gap, response gap, contributing factors, what went well.
- Blameless rules: context-based, no individual blame, never used in performance reviews.
- Weekly Incident Review Board reviews all postmortems and challenges quality.
- Publish searchable postmortem library and quarterly recurring causes report.
19. Track postmortem actions with owner and due date (depends on: 18, 12)
Fix the 11/64 completion rate by giving every action the same status as customer commitments, with capacity and escalation.
- Every action gets named owner, priority, due date, and Jira ticket auto-created from postmortem.
- P0 actions prevent SEV0 recurrence, due 30 days; P1 60; P2 90.
- Reserve 15-20% team sprint capacity for incident actions.
- Escalation ladder: manager at 7 days overdue, director at 14, CTO dashboard at 30; P0 overdue blocks features.
- Monthly reporting of closure rate in engineering leadership.
- Target 90% closure of P0/P1 on time in two quarters.
20. Create runbooks and service readiness bar (depends on: 4, 8, 11)
Ensure every service is prepared for 3 a.m. response before it is allowed to page anyone.
- Readiness checklist for Tier 0/1: architecture diagram, dependencies, dashboards, rollback, feature flags, escalation contacts, data-loss statement.
- Write major-incident playbooks: shared PostgreSQL failure, cross-region failover, Kubernetes control plane loss, partner outage, settlement breach, security compromise.
- Prioritize ledger playbooks: documented failover, read-only mode, reconciliation, signed RPO/RTO.
- Runbooks must be tested twice yearly; stale runbooks marked in catalog.
- No paging alerts without readiness sign-off; gap reported to director.
21. Train, certify, and simulate incident response (depends on: 6, 13, 14, 15, 18, 20)
Build a certification path so 24x7 roles are staffed by people who have practiced, and validate the process through drills.
- Scribe course (2 hrs); Responder (half-day); Communications Lead (1 day); Incident Commander (2 days + shadowing/tabletop).
- Certify ICs and CLs; recertify annually.
- New responders shadow two shifts before primary; no primary in first 90 days.
- Monthly tabletop using real past incidents; quarterly game day including regional failover or ledger scenario.
- Twice-yearly unannounced paging drill; one regulator/legal exercise annually.
- Track drill metrics and action items.
22. Pilot on critical services and iterate (depends on: 12, 13, 15, 19, 21)
Prove the process on the highest-risk surface before full rollout. Run a six-week pilot with tight measurement and public results.
- Select 5-6 teams: payments core, ledger/database, platform/Kubernetes, API gateway, plus two existing on-call teams.
- Activate full stack: severity, roles, command rotation, paid on-call, single tooling, alert budget, status page, postmortems.
- Weekly pilot retro; fix defects within 48 hours.
- Exit criteria: MTTD <10 min, IC assigned within 5 min in 95% incidents, pages down 50%, postmortems on time, positive sentiment.
- Publish one-page pilot result as adoption argument.
23. Wave rollout to all teams and decommission legacy paths (depends on: 22, 20)
Roll out to all 28 teams in four waves by criticality, with explicit readiness gates and legacy tool shutdown.
- Wave 1 remaining Tier 0; Wave 2 Tier 1; Wave 3 Tier 2; Wave 4 Tier 3/internal.
- Per-team onboarding: catalog complete, alerts pruned, runbooks ready, rotation staffed with 6+ certified responders, one IC candidate, one drill passed.
- Coach per wave 3 weeks.
- Gate signed by director; failures rescheduled, not waived.
- After onboarding, freeze legacy alerting paths; no fallback to old tools.
- Publish adoption scoreboard.
24. Establish metrics, dashboards, and review cadence (depends on: 12, 19, 22)
Make process performance visible with a small set of metrics and fixed review meetings, so the program is owned by data.
- Response: MTTD, time to IC, MTTA, MTTM, MTTR, % customer-detected-first.
- Quality: page volume per person, alert actionability, postmortem on-time, action closure rate.
- Business: SLA credits, availability vs 99.95, repeat incidents.
- People: on-call load, page per engineer, sentiment, attrition.
- Cadence: weekly Incident Review Board, monthly Reliability Review, quarterly Executive/Board review.
- Dashboards self-serve with targets and named owners.
25. SOC 2 readiness and internal dry-run audit (depends on: 19, 23, 24)
Design evidence as a by-product and test it with an internal walkthrough before the external auditor arrives.
- Map process to SOC2 CC7.3/CC7.4, CC7.2, CC2.2/2.3, CC5, availability criteria.
- Publish versioned policy documents: Incident Response, Severity, On-Call, Communication, Postmortem.
- Automate evidence: incident records with timestamps, role assignments, paging logs, status history, postmortems, action board, training register, drill records.
- Ensure process operates at least 3 months before fieldwork.
- Run internal dry-run month 6; sample 15 incidents; fix gaps with 8 weeks to spare.
- Keep remediation log for process deviations.
26. Continuous improvement, culture, and sustainability (depends on: 23, 25)
Prevent the process from decaying after the audit by embedding review, feedback, and roadmap ownership.
- Quarterly process retrospective with IC pool and responders.
- Re-baseline metrics every six months; raise targets.
- Year-two roadmap: follow-the-sun coverage, self-healing top 3 causes, error budgets gating releases, blast-radius reduction for shared ledger.
- Annual policy review and certification renewal.
- Quarterly on-call sentiment survey with published actions.
- Board quarterly report on availability, credits, and incident trends.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Charter the program: one owner, one mandate, funded, dated before the audit", "description": "Turn the CEO email into a chartered company program with a single accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.\n\n- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.\n- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.\n- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.\n- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.\n- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.\n- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.", "dependencies": []}, {"step_id": "S2", "title": "Seven-day operating floor so the next outage already has an owner", "description": "Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.\n\n- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.\n- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.\n- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.\n- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.\n- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.\n- Run a 15-minute daily operations review until the permanent process is live.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Forensic baseline of incidents, alerts and money lost", "description": "Rebuild the facts before designing anything. This is both the design input and the frozen \"before\" picture for the executive team and the auditor.\n\n- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.\n- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.\n- Reconstruct the two \"nobody in charge\" incidents minute by minute. They are the **burning-platform narrative** and the test case for every design decision.\n- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.\n- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.\n- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Listening tour, resistance map and the written on-call deal", "description": "Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.\n\n- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.\n- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.\n- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.\n- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.\n- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.", "dependencies": ["S1"]}, {"step_id": "S5", "title": "Service catalog, ownership, tiering and customer-journey map", "description": "You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.\n\n- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.\n- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).\n- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.\n- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.\n- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.\n- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.", "dependencies": ["S3"]}, {"step_id": "S6", "title": "Severity standard, declaration rights and incident lifecycle", "description": "Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.\n\n- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.\n- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.\n- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.\n- **SEV4:** no customer impact. Ticket only; never pages a human.\n- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.\n- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.\n- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.\n- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.\n- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.", "dependencies": ["S3", "S5"]}, {"step_id": "S7", "title": "Roles, decision authority and handover discipline", "description": "Solve \"nobody in charge for an hour\" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.\n\n- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.\n- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.\n- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.\n- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.\n- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.\n- Rules: command claimed within 5 minutes and stated in channel (\"I am IC\"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.\n- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.", "dependencies": ["S6"]}, {"step_id": "S8", "title": "Three-layer 24x7 coverage: command corps, domain rotations, triage desk", "description": "Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.\n\n- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.\n- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.\n- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.\n- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to \"I will not carry a pager for other teams' code.\"\n- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.\n- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.\n- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.", "dependencies": ["S5", "S7"]}, {"step_id": "S9", "title": "Paid on-call, New York labour compliance and fatigue safeguards", "description": "Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.\n\n- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.\n- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.\n- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.\n- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.\n- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.\n- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.\n- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.", "dependencies": ["S8", "S4"]}, {"step_id": "S10", "title": "Alert quality contract and page budget", "description": "3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.\n\n- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.\n- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.\n- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.\n- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.\n- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.\n- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.", "dependencies": ["S3", "S5"]}, {"step_id": "S11", "title": "Detection uplift on the money path, validated by incident replay", "description": "The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.\n\n- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.\n- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.\n- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.\n- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.\n- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish \"replay coverage\" as a leading metric and close the gaps it exposes.\n- Record detection source on every incident. \"Customer detected first\" becomes a named defect class with a mandatory tracked action.", "dependencies": ["S10", "S5"]}, {"step_id": "S12", "title": "One pager, one incident record, one status page", "description": "Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.\n\n- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.\n- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.\n- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.\n- Route every page through the service catalog: service label → owning domain schedule → escalation policy.\n- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.\n- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.\n- Set a hard date after which a page outside this platform creates no on-call obligation.", "dependencies": ["S6", "S8", "S10"]}, {"step_id": "S13", "title": "Escalation ladder and the five-minute command rule", "description": "Write one unskippable path from \"something looks wrong\" to \"someone is in charge\", and make the default action never be waiting.\n\n- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.\n- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.\n- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.\n- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.\n- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.\n- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.\n- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.", "dependencies": ["S7", "S8", "S12"]}, {"step_id": "S14", "title": "Live execution doctrine and payments safety rules", "description": "Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.\n\n- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.\n- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.\n- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.\n- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.\n- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.\n- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.\n- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.", "dependencies": ["S7", "S12", "S13"]}, {"step_id": "S15", "title": "Internal communications protocol", "description": "Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.\n\n- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.\n- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.\n- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.\n- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.\n- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.\n- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.", "dependencies": ["S7", "S12"]}, {"step_id": "S16", "title": "Customer communications and status page policy", "description": "Replace \"whoever is around\" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.\n\n- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.\n- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.\n- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.\n- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.\n- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.\n- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.", "dependencies": ["S6", "S15"]}, {"step_id": "S17", "title": "Regulatory, partner and legal notification playbook", "description": "In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.\n\n- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.\n- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is \"not reportable\"**, with evidence, decision-maker and timestamp.\n- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.\n- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.\n- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.\n- Rehearse the playbook quarterly as part of the exercise programme.", "dependencies": ["S6", "S16"]}, {"step_id": "S18", "title": "SLA credit and financial impact workflow", "description": "Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.\n\n- Agree with Legal and Finance the availability measurement method per contract and per component.\n- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.\n- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.\n- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.\n- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.\n- Use the credit delta as the standing business case for on-call pay and reserved reliability capacity.", "dependencies": ["S6", "S16"]}, {"step_id": "S19", "title": "Blameless postmortem standard and Incident Review Board", "description": "Replace \"some incidents, various formats\" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.\n\n- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.\n- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.\n- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.\n- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**\n- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.\n- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.\n- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly \"top five recurring causes\" analysis, with restricted versions for security or privileged content.", "dependencies": ["S6", "S7"]}, {"step_id": "S20", "title": "Action ownership, reserved capacity and enforcement", "description": "Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.\n\n- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.\n- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.\n- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.\n- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.\n- Verify effectiveness before closing. A closed ticket without evidence does not close the action.\n- Report closure rate by team monthly and include it in engineering manager objectives.", "dependencies": ["S19", "S12"]}, {"step_id": "S21", "title": "Runbooks, readiness bar and ledger blast-radius reduction", "description": "Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.\n\n- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.\n- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.\n- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.\n- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.\n- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.", "dependencies": ["S5", "S8"]}, {"step_id": "S22", "title": "Training and certification academy", "description": "Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.\n\n- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.\n- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.\n- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.\n- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.\n- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.\n- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.\n- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.", "dependencies": ["S7", "S13", "S15", "S19"]}, {"step_id": "S23", "title": "Exercise programme: tabletops, game days and unannounced drills", "description": "The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.\n\n- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.\n- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.\n- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.\n- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.\n- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.\n- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.\n- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.", "dependencies": ["S22", "S12", "S21"]}, {"step_id": "S24", "title": "Publish Incident Management Policy v1 and the exception register", "description": "Collapse the design into a document people will actually open mid-outage, and make it official.\n\n- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.\n- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.\n- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.\n- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.\n- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.", "dependencies": ["S6", "S7", "S8", "S9", "S10", "S13", "S15", "S16", "S17", "S19", "S20"]}, {"step_id": "S25", "title": "Pilot on the payments critical path", "description": "Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.\n\n- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.\n- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.\n- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.\n- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.\n- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.\n- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.", "dependencies": ["S24", "S11", "S12", "S21", "S22", "S9"]}, {"step_id": "S26", "title": "Metrics, dashboards, review cadence and anti-gaming", "description": "Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.\n\n- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.\n- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.\n- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.\n- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.\n- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.\n- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.\n- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.\n- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.", "dependencies": ["S12", "S19", "S25"]}, {"step_id": "S27", "title": "Alert noise burn-down campaign", "description": "Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.\n\n- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.\n- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.\n- Publish a weekly noise leaderboard that names systems, never people.\n- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.\n- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.\n- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.", "dependencies": ["S10", "S12", "S25"]}, {"step_id": "S28", "title": "Wave rollout to all 28 teams with readiness gates", "description": "Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.\n\n- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.\n- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.\n- Each wave gets a named coach from the program office for three weeks.\n- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.\n- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.\n- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.", "dependencies": ["S25", "S26"]}, {"step_id": "S29", "title": "Change management, fairness and pager culture", "description": "Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.\n\n- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.\n- Publicly close the two historical \"nobody in charge\" incidents with a written account of what would be different now.\n- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.\n- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.\n- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.\n- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.", "dependencies": ["S4", "S9", "S25"]}, {"step_id": "S30", "title": "SOC 2 evidence by design, internal testing and mock audit", "description": "Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.\n\n- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.\n- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including \"not reportable\", postmortems, action closure with verification, training register, drill records, access reviews.\n- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.\n- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.\n- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.\n- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.", "dependencies": ["S24", "S26", "S28"]}, {"step_id": "S31", "title": "Risk register and pre-committed contingencies", "description": "Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.\n\n- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.\n- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.\n- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.\n- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.\n- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.\n- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.\n- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.", "dependencies": ["S1"]}, {"step_id": "S32", "title": "Ninety-day inspect-and-adapt, then year-two sustainability", "description": "Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.\n\n- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.\n- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.\n- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.\n- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.\n- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.\n- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.", "dependencies": ["S26", "S28", "S30"]}], "estimated_complexity": "high", "success_metrics": "- A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.\n- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.\n- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.\n- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.\n- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.\n- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.\n- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.\n- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.\n- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.\n- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.\n- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.\n- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.\n- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.\n- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.\n- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.\n- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.\n- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.\n- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.\n- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.\n- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of \"not reportable\".\n- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.\n- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.\n- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.\n- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly."}Renamed the scale to SEV1–SEV4, ending the SEV0/SEV1 ambiguity, and split seven previously buried topics (fairness contract, alert contract, customer comms, actions, policy, metrics, sustainability) into their own steps. Compensation moved from "fixed weekly stipends" to actual dollar bands and an annual budget.
- Step 6 adopts SEV1–SEV4 and keeps the guardrail caveat, so the scale is now comparable with every other plan and no longer has two crisis tiers.
- New step 4 promotes the listening tour and written fairness contract to an early standalone workstream with baseline, day-60, day-120 and quarterly surveys.
- Step 9 gives approvable numbers: $800–1,200 primary, $300–500 secondary, $900–1,300 duty commander, ~$0.9M–1.2M/yr, plus FLSA/NY treatment and a 15% sprint reduction.
- Standalone alert-quality contract (11), customer/status-page comms (15), corrective actions (18), policy-and-certification (20) and metrics (24) make each control separately testable.
- Best audit sequencing in the round: control mapping, retention, legal hold and observation-window confirmation in step 3; monthly sampling and tests in months 4 and 6 (step 25).
- Fastest credible schedule — Tier 0/1 coverage by week 10, all 28 teams by week 18 — giving the longest operating-evidence window before fieldwork.
- Unique operational realism: surge roster for simultaneous incidents and a 24x7 Executive Duty Officer (step 8); six-responder rule scoped to direct 24x7 rotations only (step 22).
- Still the only plan without a risk register or pre-committed contingencies for compensation delay, commander shortfall, tool-migration slip or a SEV1 mid-rollout.
- No quota-driven noise burn-down campaign; step 11 relies on review cadence plus "burn down the top 50 first", which is weaker than P1/P3's leaderboard and per-sprint quota.
- No decision tree with worked examples from the 31 real incidents, so severity calibration rests on prose alone.
- Dropped its round-1 numeric declaration thresholds without substituting P1-style payment anchors, leaving SEV1/SEV2 boundaries more subjective than before.
- Proposal 1 : Four-level SEV1–SEV4 scale with SEV4 as a never-paging ticket class.
- Proposal 1 : Listening tour, resistance map and a written on-call deal as a first-class early step.
- Proposal 1 : Published dollar figures for on-call stipends and a named annual budget.
- Proposal 1 : Separate customer-communication step with timed status-page posts and tiered account-manager outreach.
- Proposal 1 : Versioned policy suite plus one-page severity/role/escalation cards inside the incident tool.
- Proposal 1 : A standalone risk register with pre-committed contingencies reviewed monthly with the sponsor.
- Proposal 1 : Quota-driven noise burn-down campaign with a weekly leaderboard and an automatic 8-week downgrade rule.
- Proposal 1 : Consolidating the 28 teams into 10–12 domains as the default coverage design.
+ Install a seven-day incident-response floor+ Create the factual, legal, and audit baseline+ Co-design the fairness contract with engineers+ Enforce an alert-quality contract and page budget+ Standardize status-page and customer communications+ Give corrective actions enforceable ownership+ Publish policy and certify every response role+ Measure performance and operate fixed review forums+ Institutionalize improvement and reduce structural riskInstall an interim process in seven daysDesign compliance and evidence controls from day oneEnforce alert quality and burn down noise safelyEnforce action ownership and effectiveness trackingMeasure outcomes, controls, business impact, and human loadTrain participants and address pager resistanceInstitutionalize continuous improvement
The plan produced
1. Establish the mandate, owner, funding, and schedule
Launch incident management as a company operating program within 48 hours. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Form a decision group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Approve the non-negotiables: one severity scale, one human-paging platform, one incident record, paid on-call, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, and incident actions.
- Fund tooling, compensation, training, exercises, and reliability work. Use the $1.3M in credits as the minimum financial comparison.
- Set milestones: interim process by day 7, policy and critical-path pilot by week 6, Tier 0/1 coverage by week 10, company rollout by week 18, internal audit in month 6, and mock audit in month 7.
2. Install a seven-day incident-response floor (after 1) new
Do not wait for policy design or tooling migration. Put a minimum process into operation immediately and begin retaining evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communication schedule.
- Create one monitored declaration path through chat, telephone, and the current paging environment.
- Use one incident channel, bridge, timeline document, and naming convention for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced on-call engineers and engineering leaders.
- Require a named commander within 10 minutes. The duty engineering director assumes command if nobody else does.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Compensate interim duty retroactively under the permanent policy.
- Hold a 15-minute daily operational review until the permanent process is live.
3. Create the factual, legal, and audit baseline (after 1) new
Build one defensible baseline for process design, executive decisions, and SOC 2 testing. Confirm the auditor's expected Type II observation period immediately.
- Reconstruct all 31 incidents from first customer impact through resolution, communications, credits, postmortem, and corrective action.
- Analyze the two incidents with unclear command minute by minute.
- Identify the missing detection signal for every customer-first incident.
- Inventory the six alert sources, 3,400 monthly alert events, noisy rules, duplicates, missing owners, and missing runbooks.
- Record the current rotations, unpaid work, after-hours load, and teams without coverage.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M credits, 85% noise, and 11 of 64 actions closed.
- Map incident-response controls to applicable SOC 2 criteria with Compliance and the auditor.
- Establish retention, confidentiality, legal-hold, and access requirements for incident evidence.
4. Co-design the fairness contract with engineers (after 1) new
Treat pager resistance as a valid design constraint. Make the employment and ownership bargain explicit before expanding on-call.
- Interview representatives from all 28 teams and the Support, Customer Success, Security, and Operations groups.
- Distinguish objections involving unpaid work, unfamiliar code, bad alerts, weak runbooks, sleep disruption, or blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for services they own or have formally accepted, supported by trained command, and given capacity to remove recurring defects.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Measure baseline trust, fairness, fatigue, and psychological safety. Repeat the survey at days 60 and 120, then quarterly.
5. Build the service catalog and criticality model (after 3)
Make a machine-readable catalog the source of truth for routing, escalation, customer impact, and audit evidence. Every production service must have one accountable owner.
- Record the owning team, manager, product capability, escalation policy, communication channel, dashboard, runbook, dependencies, regions, and data stores for all 180 services.
- Map payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, and reporting to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing services whose failure can create material contractual impact.
- Classify Tier 2 as internal or deferrable services, and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional mode, recovery method, and contractual obligations for Tier 0 and Tier 1.
- Give orphan services an owner or approved decommission date within 30 days.
- Maintain separate but coordinated ownership for the ledger application and PostgreSQL platform.
6. Adopt one severity scale and incident lifecycle (after 3, 5)
Use four impact-based levels. Classify on actual or credible customer, financial, security, regulatory, and contractual harm rather than organizational seniority.
- SEV1 — crisis: incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; both regions impaired; or a core payment journey broadly unavailable. Page every role immediately, open a bridge, notify executives, publish customer status within 15 minutes when applicable, begin legal assessment, and require a postmortem.
- SEV2 — major: material payment degradation; settlement deadline at risk; a critical customer or material cohort unavailable; regional impairment with reduced resilience; or likely SLA breach. Page command and technical roles, publish customer status within 30 minutes when customer-visible, and require a postmortem.
- SEV3 — limited: narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads; page only when immediate action is necessary.
- SEV4 — operational event: no current customer impact and no urgent risk. Create a ticket and handle during normal operations.
- Treat percentages as supporting guardrails, never reasons to under-classify integrity, settlement, security, or contractual risk.
- Anyone may declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Automatically use at least SEV2 posture for suspected ledger-integrity events, cross-team incidents with unknown ownership, and materially unknown impact lasting 15 minutes.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed. Resolution requires stability, backlog recovery, and necessary reconciliation.
7. Define roles, authority, and handoffs (after 6)
Separate command, communications, recordkeeping, and technical repair. One named person must hold command throughout every SEV1 and SEV2.
- Incident Commander: owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not serve as the primary technical operator.
- Communications Lead: owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language.
- Scribe: maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- Subject-Matter Responders: diagnose and mitigate only services for which they have ownership, access, training, or formally accepted support responsibility.
- Executive Duty Officer: removes organizational barriers and makes exceptional business decisions without displacing the commander.
- Security, Legal, Compliance, Vendor Management, and Finance join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Require distinct commander, communications, scribe, and primary technical lead for SEV1. Communications and scribe may combine temporarily for bounded SEV2 incidents.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, and reconciliation requirements for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, with the exact transfer time.
8. Create sustainable 24x7 coverage across 28 teams (after 4, 5, 7)
Use central command coverage and risk-based technical coverage rather than creating 28 fragile night rotations. All services receive a response path, but only critical domains maintain direct overnight technical rotations.
- Create a 24x7 Incident Command corps of 24–30 certified people, with primary and secondary coverage at all times.
- Create 18–24 trained Communications Leads and a similarly sized scribe pool using Support, Customer Operations, Engineering Operations, and qualified managers.
- Maintain a surge roster for simultaneous incidents and a 24x7 Executive Duty Officer schedule.
- Group Tier 0 and Tier 1 ownership into approximately 8–12 coherent domains only where responders have training, access, runbooks, and explicit acceptance.
- Staff each critical domain with primary and secondary responders and at least six qualified people. Target no more than one primary week in six.
- Give Tier 2 and Tier 3 services business-hours coverage plus a maintained manager and director escalation path.
- When lower-tier impact becomes SEV1 or SEV2, central command activates the manager escalation path and obtains the necessary owner.
- Route unknown-owner pages to Duty Command and platform triage temporarily. Log each one as a catalog control failure.
- Do not merge small teams merely for schedule convenience. Provide training, reassign services, add staffing, or decommission unsupported services.
9. Approve compensation and fatigue protections (after 4, 8)
End unpaid on-call before expanding mandatory night coverage. HR, Finance, Payroll, and employment counsel should approve the policy within 14 days.
- Use market-validated weekly bands, initially budgeting approximately $800–$1,200 for Tier 0/1 primary duty and $300–$500 for secondary duty.
- Budget approximately $900–$1,300 for Duty Incident Commander weeks and $400–$800 for Communications Lead or scribe duty, adjusted for actual burden.
- Pay holiday premiums and compensate active after-hours work according to exempt or non-exempt status and applicable federal and New York rules.
- Provide a protected paid recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal sprint commitments by approximately 15% during a primary on-call week.
- Prohibit simultaneous primary assignments, consecutive primary weeks, on-call during leave, and invisible schedule swaps.
- Allow responders to declare themselves temporarily unfit after disruptive night work without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Budget roughly $0.9M–$1.2M annually, then refine using actual rotation count, employment classification, and activation data.
10. Establish one paging and incident system of record (after 3, 5, 8)
Monitoring tools may remain specialized, but all human pages must enter one controlled platform. This removes conflicting schedules and creates one evidence trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page.
- Ingest events from the six existing monitoring tools before disabling their direct paging paths.
- Route pages through service-catalog ownership and deduplicate related events.
- Provide a single declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, severity changes, decisions, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status-page publishing, and conference facilities.
- Require role-based access, MFA, access reviews, and immutable or tamper-evident history.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, or an AWS region.
- Retire each legacy human-paging path only after ownership review, end-to-end tests, and two weeks of verified operation.
11. Enforce an alert-quality contract and page budget (after 3, 5, 10) from P4 step 10
Treat every page as a production interface with an owner and an expected action. Noise reduction must not create detection gaps.
- Require every paging rule to identify the service, owning team, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Page only when prompt human judgment or intervention can materially reduce customer, financial, security, or contractual risk.
- Route informational, capacity, and non-urgent infrastructure conditions to dashboards or ticket queues.
- Prefer payment-outcome, settlement-risk, queue-age, and error-budget-burn alerts over raw CPU, memory, pod, or log thresholds.
- Run new paging rules in shadow mode for at least seven days unless an emergency exception is approved.
- Review rules with repeated no-action acknowledgements, low actionability, or excessive firing within two business days.
- Set a load budget of no more than two after-hours pages per responder per week, measured over four weeks.
- Require an owner, compensating detection, and recorded approval before suppressing or deleting a rule.
- Burn down the top 50 noisy rules first. Review missed detections and noise together so teams cannot improve metrics by becoming blind.
12. Detect payment and ledger failures before customers (after 5, 11)
Move detection from host health to customer journeys and financial outcomes. Use internal SLOs with enough headroom to protect the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, API access, webhooks, and reporting freshness.
- Run external synthetic transactions through critical payment journeys at least every minute and from paths independent of the production platform.
- Validate each AWS region and expose dependencies that undermine nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed value relative to settlement deadlines, not only queue depth.
- Add tenant and cohort anomaly detection for high-value customers and major payment methods.
- Convert high-priority Support, account-manager, processor, sponsor-bank, and network reports into incident candidates within five minutes.
- Record the detection source for every incident. Treat customer-first detection as a mandatory missed-detection review.
13. Codify detection, escalation, and live execution (after 7, 10)
Create one time-bound path from the first credible signal to named command and mitigation. Notification delivery does not count as human acknowledgement.
- Page the owning critical-domain primary and Duty Incident Commander immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at five minutes, manager at 10, and director at 15.
- Escalate an unclaimed command page to backup commander at five minutes. The Executive Duty Officer assumes command at 10 minutes until a certified handoff occurs.
- Require Support and account managers to use the same declaration path as automated monitoring and engineers.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Begin with a standard statement of severity, known impact, assigned roles, current objective, workstreams, and next update time.
- Freeze unrelated changes during SEV1 and normally during SEV2. Record any exception.
- Prefer reversible mitigation such as rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require formal command handoff for long incidents, shift changes, or fatigue. Do not leave an incident unowned during transfer.
- Reconcile the ledger and safely drain backlogs before resolving payment or ledger incidents.
14. Standardize internal incident communications (after 7, 10, 13)
Give responders one working room and stakeholders one controlled information source. Executives must not interrupt the technical command path.
- Create one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue an initial internal brief within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within 15 minutes for SEV1 and 30 minutes for SEV2.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record material decisions and outbound messages in the incident timeline.
15. Standardize status-page and customer communications (after 6, 10, 14)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate impact before root cause is known.
- Publish customer status within 15 minutes of declaring a customer-visible SEV1 and within 30 minutes for customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice promptly after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map status-page components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, data-integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, known workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Give affected strategic accounts direct account-manager outreach within 30 minutes for SEV1 and 60 minutes for SEV2.
- Require account managers to use the approved briefing and prohibit independent technical explanations.
- Provide a customer-facing incident summary within five business days for SEV1 and qualifying SEV2 events.
- Record any legally necessary delay or restriction of public detail, its approver, and the alternative communication plan.
16. Operationalize legal, regulatory, contractual, and credit decisions (after 3, 6, 15)
Some payments incidents start external notification clocks. Make assessment mandatory without assuming that every operational incident is reportable.
- Build a counsel-validated matrix covering applicable NYDFS rules, state breach laws, GLBA or FTC obligations, PCI requirements, money-transmitter obligations, sponsor-bank and network contracts, cyber insurance, and customer contracts.
- Engage Legal and Compliance immediately for every SEV1, security event, and suspected ledger-integrity event.
- Record an initial reportability assessment within one hour for SEV1 and within two hours for other potentially reportable events.
- Record non-reportable decisions as evidence, including facts considered, approver, timestamp, and reassessment trigger.
- Maintain tested 24x7 contacts for counsel, regulators, sponsor banks, networks, insurers, and critical vendors.
- Encode customer-specific notice deadlines and channels in the customer record used by the Communications Lead.
- Have Legal own regulatory text and submission. Keep technical command with the Incident Commander.
- Have Finance calculate affected minutes, delayed value, likely credits, and contractual exposure within five business days.
- Track credits by incident and recurring cause to support reliability investment decisions.
17. Make postmortems mandatory, consistent, and blameless (after 6, 7, 10)
Use one learning standard with fixed deadlines. Keep postmortems separate from performance, misconduct, and disciplinary processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact over two hours, SLA credits, contractual breach, repeated contributing factors, control failures, and ledger-integrity near misses.
- Produce the factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the Incident Commander responsible for the timeline and the owning engineering director accountable for completion.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit analysis of why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the incident commander.
- Describe decisions in the context and information available at the time. Do not name an individual as the root cause.
- Publish broadly useful findings internally while restricting security, privacy, personnel, or privileged material appropriately.
18. Give corrective actions enforceable ownership (after 10, 17) new
Treat incident actions as risk commitments, not suggestions. A ticket is not complete until the expected risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Escalate overdue high-risk actions to the manager after seven days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit overdue actions to block related releases where the unrepaired condition could reproduce severe impact.
- Verify effectiveness through tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days. Complete, re-plan, or formally accept each risk.
19. Set the service-readiness bar and critical playbooks (after 5, 8, 11, 12)
A critical service must be supportable at 3 a.m. before its team is placed on direct overnight coverage. Existing critical detection must remain active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and data-integrity constraints.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional failure, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document controlled failover, split-brain prevention, replay protection, backlog recovery, and post-recovery reconciliation.
- Exercise critical runbooks at least twice per year and after material changes.
- Block new Tier 0/1 releases and new paging rules when readiness requirements are missing.
- Handle existing gaps through a named owner, compensating control, executive-approved expiry date, and remediation plan rather than disabling detection.
20. Publish policy and certify every response role (after 6, 7, 8, 9, 11, 13, 14, 15, 16, 17, 18) new
Convert the operating design into concise, signed documents and practical training. Training and exercises occur during paid working time.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, and exception standards.
- Provide one-page severity, role, authority, escalation, and communication cards inside the incident tool.
- Train all employees to recognize and declare incidents.
- Train all engineers in severity, acknowledgement, evidence preservation, handoff, and financial-integrity precautions.
- Certify responders only after they demonstrate access, dashboards, runbooks, rollback, and escalation competence.
- Certify Incident Commanders through formal instruction, simulation, and at least two shadowed incidents or exercises.
- Train Communications Leads in status writing, account segmentation, contractual clocks, and legal escalation.
- Train scribes in timeline quality and fact-versus-hypothesis labeling.
- Require two shadow shifts before independent primary duty and annual recertification.
- Maintain the training and certification register as operational and audit evidence.
21. Pilot on the payment critical path (after 9, 10, 12, 19, 20)
Run a four-to-six-week pilot across the highest-risk journey before expanding. Use real incidents and exercises to correct the model quickly.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, API edge, authentication, settlement, and Support intake.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently assuming command.
- Correct critical process or tool defects within 48 hours and update the standard visibly.
- Exit only after 95% timely command assignment, 95% communication compliance, no unpaid pages, complete required postmortems, and at least 50% lower noise.
22. Roll out by customer journey and risk (after 21)
Expand in controlled waves while the interim process remains active company-wide. Complete rollout early enough to accumulate operating evidence before the audit.
- Roll out remaining Tier 0 domains first, then Tier 1, Tier 2, and Tier 3.
- Use two-to-three-week waves with a named program coach and director sign-off.
- Gate each team on catalog ownership, appropriate coverage, compensation, trained responders, tested escalation, alert quality, runbooks, access, and a passed tabletop.
- Require six or more responders only for direct 24x7 technical rotations. Apply business-hours coverage and manager escalation to lower tiers.
- Reschedule failed gates or approve a time-limited executive exception with a compensating control.
- Disable legacy human-paging paths after verified cutover for each wave.
- Publish an internal adoption dashboard by team, service tier, and control gap.
- Finish critical coverage by week 10 and all 28 teams by week 18.
23. Exercise command, communications, regional recovery, and fallbacks (after 19, 20)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous security and operational events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors where appropriate.
- Conduct at least one unannounced after-hours paging test before the audit.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure time to acknowledgement, command, customer notice, mitigation decision, handoff, and recovery.
- Create tracked actions for every material exercise finding.
24. Measure performance and operate fixed review forums (after 3, 10, 17, 18) new
Use a balanced scorecard that exposes weak controls without rewarding hidden incidents or suppressed alerts. Report medians and 90th percentiles rather than averages alone.
- Measure time from first impact to detection, declaration, acknowledgement, command assignment, mitigation, recovery, and resolution.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track pages, actionability, duplicates, after-hours load, missed detections, and page-budget breaches.
- Track postmortem timeliness, action age, closure by due date, verified effectiveness, and recurring contributing factors.
- Track availability by customer journey, error-budget burn, failed or delayed payment value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight activations, recovery days, schedule exceptions, sentiment, and attrition.
- Hold a weekly Incident Review Board for incidents, postmortems, control failures, noisy alerts, and overdue actions.
- Hold a monthly executive reliability review for trends, funding, contractual exposure, and accepted risks.
- Hold a quarterly control and resilience review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile incident records monthly against support cases, status history, credits, and customer complaints to detect under-reporting.
25. Prove SOC 2 operating effectiveness before fieldwork (after 20, 22, 23, 24)
Generate audit evidence through normal operation rather than reconstructing it later. Test both control design and consistent execution.
- Maintain approved and versioned policies, exceptions, catalog records, schedules, compensation activation, access reviews, training, incidents, communications, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, non-reportable legal decisions, missed timelines, and exercises in the population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the populations, evidence requests, and interviews expected from the external auditor.
- Record deviations honestly with owners, remediation dates, and compensating controls. Never rewrite historical records.
- Confirm that evidence retention covers the full auditor-defined observation period.
- Brief commanders, responders, Support, and Compliance on the real process without scripting inaccurate answers.
26. Institutionalize improvement and reduce structural risk (after 22, 24, 25) new
Prevent the program from decaying after the audit. Use incident evidence to drive permanent ownership and architectural investment.
- Assign permanent owners for policy, service catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review severity thresholds, staffing, compensation, and communication timings annually and after material process failures.
- Recertify command and communications personnel annually.
- Review recurring failure families quarterly and require executive decisions where corrective work repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fairness, fatigue, and tool friction.
- Report severe incidents, credits, overdue high-risk actions, and resilience investment to the board or risk committee quarterly.
- Treat the shared ledger cluster as a strategic concentration risk. Fund failover assurance, blast-radius reduction, isolation of non-critical readers, and stronger regional independence.
- Evaluate follow-the-sun command or technical coverage using one year of page-load and staffing data.
- Build year-two plans for automated mitigation, deployment safety, graceful degradation, and error-budget release controls.
- By day 7, every suspected major incident uses one incident record, one coordination channel, and a named commander.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes; at least 95% receive command within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and technical coverage.
- By week 18, all 180 services have an owner, tier, tested escalation path, and coverage appropriate to their risk.
- No mandatory night rotation starts before compensation, training, access, runbooks, and staffing controls are active.
- Every direct 24x7 technical rotation has at least six qualified responders or a documented, expiring executive exception.
- No responder is routinely assigned primary duty more often than one week in six or simultaneously assigned to two primary rotations.
- At least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of qualifying SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet their required internal and customer update cadence by month 3.
- Monthly human pages fall from 3,400 alert events to no more than 1,500 by day 90 and no more than 500 actionable pages by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, with no decline in Tier 0/1 detection coverage.
- Average after-hours load remains at or below two pages per responder per week; every sustained breach produces a remediation plan.
- All six monitoring sources route human pages through the single paging platform by week 12; direct legacy paging paths are disabled by week 18.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Monthly contracted availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- Quarterly on-call surveys reach at least 75% favorable responses on fairness, ownership boundaries, compensation, and sustainability by month 6.
- The month-7 mock audit finds no unowned high-risk control gap, and at least 95% of sampled incidents contain complete operating evidence.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 3efb9ff1-28a5-40aa-8f73-9853e91aa095, Agent: opus5_refine_1, LLM: anthropic/claude-opus-5):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (32):
1. Executive mandate, single owner, funding and non-negotiables
Convert the CEO email into a chartered program with one accountable owner and authority over all 28 teams. Incident response becomes a **company operating process**, not a per-team choice.
- Name an executive sponsor (CTO) and one accountable owner (Head of Reliability / Incident Management) with a small permanent office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group with Engineering, SRE/Platform, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours. It is not a 28-person committee.
- Fix the non-negotiables now: one severity scale, one paging tool, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.
- Set the clock deliberately earlier than the audit: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready week 28.
- Approve budget against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises.
2. Seven-day interim command bridge (depends on: 1)
Do not let design work leave the company unprotected for six weeks. Put a crude but real process in place within seven days and improve it later.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster from engineering managers and senior SREs, primary plus backup, 24x7. Pay it retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Tell Support to escalate credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items: complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This becomes both the design input and the "before" picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start / detect / declare / commander-assigned / mitigate / resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. This list drives the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome. Identify the top 50 rules producing most of the noise and every rule with no owner or runbook.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS and Sales in two weeks. Test what the objection actually is: unpaid work, lost sleep, unfamiliar code, missing runbooks, or fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey (fairness, trust in alerts, willingness, burnout) to re-measure at 60, 120 and 365 days.
5. Service catalog, ownership and money-path tiering (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that is the single source of truth for paging, impact and audit.
- One accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not by technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions and third parties, including sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Every orphan service gets an owner in 30 days or a decommission date approved by the sponsor. Tier 0 without an owner is an executive escalation.
6. Severity scale, declaration rules and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- **SEV3 (major, contained):** narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- **SEV4:** no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring; only the commander may downgrade, with the evidence recorded.
- Lifecycle: Detected → Declared → Triaged → **Mitigated** (customer impact ends) → Monitoring → **Resolved** (backlog processed and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Incident roles, authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, roles, cadence and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit spend. Retains command when a VP joins.
- **Communications Lead:** single voice for status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates for SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team; they mitigate, they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance and Vendor Management join on defined triggers.
- Rules: command claimed within 5 minutes, stated in channel ("I am IC"), distinct people for command, comms and technical lead at SEV1/SEV2, and every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation or privileged-access rules.
8. 24x7 coverage model: central command corps, local expertise (depends on: 5, 7)
Do not create 28 night rotations. Centralise coordination in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- **Layer A — Incident Command corps:** ~30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of ~18 from Support, CS and engineering management; a scribe pool used as the training entry point.
- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- **Layer C — everyone else:** business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing, with cost and timeline: a US-only paid night rotation now, a Lisbon or APAC follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
9. On-call compensation, labour compliance and fatigue safeguards (depends on: 8, 4)
Unpaid on-call in New York is both a retention problem and a legal exposure. Pay for it before asking anyone to sign up, and publish the numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 for business-hours rotations, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover rather than expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, tax treatment and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a ~15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
10. Alert quality standard, page budget and noise burn-down (depends on: 3, 5)
3,400 alerts at 85% noise is the reason detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human, and burn the backlog down deliberately rather than by mass silencing.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a **page budget** of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements; return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
- Target 3,400 → under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
11. Detection uplift on the money path (depends on: 10, 5)
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls; five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrating without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack or the identity provider is unavailable. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live incident execution doctrine (depends on: 7, 12, 13)
Give responders one short operating procedure for the first minutes through closure. Priority is limiting customer and financial harm, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; controlled backlog drain; **reconciliation completed before any payment or ledger incident is declared resolved**.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and Customer Success, and reopening if impact recurs.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook and RSS; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack; the long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity or blame; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded **even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review; Legal owns outbound regulatory text, the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
18. SLA credit and financial impact workflow (depends on: 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one, and use the delta as the standing business case for on-call pay and reliability capacity.
19. Blameless postmortem standard and Incident Review Board (depends on: 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first?** and **why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, capacity reservation and enforcement (depends on: 19, 12)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
21. Runbooks, readiness bar and ledger blast-radius reduction (depends on: 5, 8)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the single largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel workstream on blast-radius reduction — tenant or function partitioning, read replicas, and isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (depends on: 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days + shadowing):** command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence for their own domain before independent primary duty; two shadow shifts minimum, never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (depends on: 22, 12, 21)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 11, 12, 21, 22, 9)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect (from first impact), time to declare, time to commander, acknowledgement, mitigate, resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; team scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Wave rollout to all 28 teams with readiness gates (depends on: 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director; failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
28. Alert noise burn-down campaign (depends on: 10, 12, 25)
Run the noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
29. Change management, fairness and pager culture (depends on: 4, 9, 25)
Run this from day one in parallel. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (depends on: 24, 26, 27)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and contingencies (depends on: 1)
Name the ways this programme fails and pre-commit the response. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 26, 27, 30)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 2 (ID: 2e38c07e-bf01-4b70-ba95-0ca1ef39e2d8, Agent: gpt5.6-sol_refine_2, LLM: openai/gpt-5.6-sol):
Estimated Complexity: high
Success Metrics: - By day 7, every suspected SEV0–SEV2 uses one incident record, one coordination channel, and one named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named commander for more than 10 minutes; at least 95% are assigned within 5 minutes.
- By day 30, 100% of Tier 0 services have a named owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, every Tier 0 and Tier 1 customer journey has compensated 24×7 command and subject-matter coverage.
- By day 120, all 180 production services have an owner, response tier, tested escalation path, and appropriate coverage model.
- No mandatory night rotation starts before its compensation, training, access, and staffing controls are active.
- Every critical primary rotation has at least six qualified responders or a documented, expiring executive exception by day 120.
- No responder is routinely scheduled for primary duty more often than one week in six by day 120.
- At least 95% of critical pages are acknowledged within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by day 120 and under 60 minutes by month 8.
- At least 95% of SEV0 and SEV1 customer notices are issued within 15 minutes of declaration by month 3.
- At least 95% of customer-visible SEV2 notices are issued within 30 minutes by month 3.
- At least 95% of incidents meet their required update cadence by month 3.
- Monthly paging volume falls from 3,400 to no more than 1,500 by day 90 and no more than 700 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6 without loss of critical detection coverage.
- All new paging alerts satisfy the owner, impact, action, dashboard, runbook, deduplication, and escalation standard by day 60.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 currently open historical actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date by month 6, with effectiveness evidence.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Two end-to-end cross-company exercises, including regional failure and ledger recovery, are completed before the audit.
- Monthly availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- SLA credits decline by at least 50% on an annualized trailing basis within 12 months.
- Quarterly on-call surveys show improving fairness and sustainability, with at least 75% favorable responses by month 6.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (24):
1. Create the mandate, ownership, and funding
Launch incident management as a company operating program within 48 hours. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable program owner.
- Form a small steering group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Fund two to three implementation staff, paging and incident tooling, observability work, training, exercises, and on-call compensation.
- Reserve 10%–15% of engineering capacity for alert remediation, runbooks, and incident actions.
- Authorize incident commanders to freeze deployments, order rollback, disable features, shift traffic, and invoke continuity plans.
- Preserve financial controls. Incident commanders may coordinate ledger recovery but may not bypass dual approval, privileged-access controls, or reconciliation.
- Set the delivery target: critical controls operational within 60 days, enterprise rollout within 120 days, and a mock audit in month 7.
2. Install an interim process in seven days (depends on: 1)
Do not wait for new tools or the final policy. Put a minimum viable incident process into operation immediately and start collecting evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communications schedule.
- Establish one monitored declaration path through chat, telephone, and the existing paging tools.
- Create a standard incident channel, bridge, incident document, and naming convention.
- Staff temporary primary and backup incident commanders 24×7 from the existing on-call teams and engineering leadership.
- Compensate interim duties retroactively under the final compensation policy.
- Require a named incident commander within 10 minutes for every suspected major incident.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Hold a daily 15-minute control review until the permanent process is live.
3. Build the baseline, ownership catalog, and risk map (depends on: 1)
Establish the facts behind the current failures and assign every production component an owner. Use the resulting catalog as the source for routing, escalation, and audit evidence.
- Reconstruct all 31 customer-impacting incidents, including first impact, detection source, declaration, command assignment, mitigation, resolution, customer communications, and credits.
- Analyze the two incidents with no clear leader and every case detected first by customers.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, databases, queues, regional dependencies, payment processors, and banking partners.
- Assign one accountable team, engineering manager, product owner, primary escalation, secondary escalation, dashboard, and runbook to each service.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 based on financial integrity, customer impact, contractual exposure, and dependency centrality.
- Map critical customer journeys to their application, PostgreSQL, Kubernetes, regional, and third-party dependencies.
- Inventory the six alert sources and all 3,400 monthly alerts by owner, volume, actionability, and duplication.
- Interview representatives from all 28 teams and baseline on-call sentiment, fatigue, and objections.
- Give orphan services an owner or decommissioning decision within 30 days.
4. Design compliance and evidence controls from day one (depends on: 1)
Map the operating process to audit, legal, contractual, and record-retention requirements before finalizing it. Confirm the expected SOC 2 observation period with the auditor immediately.
- Map controls to applicable SOC 2 criteria for monitoring, incident identification, response, recovery, communications, corrective action, access, and availability.
- Define evidence required for declarations, pages, acknowledgements, role assignments, decisions, status updates, postmortems, actions, training, drills, and exceptions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve the Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, and Postmortem Standard.
- Record control exceptions with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing evidence before the audit.
5. Adopt one severity and lifecycle standard (depends on: 2, 3, 4)
Use four impact-based severity levels across operational, security, data, and third-party incidents. Start at the highest credible severity when facts are uncertain, then downgrade with recorded evidence.
- **SEV0 — financial or security crisis:** suspected ledger corruption, unauthorized or duplicated funds movement, material data compromise, both-region loss, or a decision to suspend payment processing. Page all roles immediately; engage executives, Security, Legal, Compliance, and Risk; assess regulatory duties within one hour; use distinct role holders; require a postmortem.
- **SEV1 — critical availability event:** a core payment journey is unavailable, payment failures exceed an initial 10% guardrail for five minutes, regional loss has impaired failover, a settlement deadline is at imminent risk, or rapid error-budget burn makes a material SLA breach likely. Assign all roles; notify internal stakeholders within 10 minutes; publish customer status within 15 minutes; update every 30 minutes; require a postmortem.
- **SEV2 — major bounded event:** approximately 1%–10% of payment attempts fail, a material customer subset or critical customer is down, degradation has a workaround, or contractual impact is likely. Assign an incident commander and responders; add communications and scribe roles for customer impact; publish status within 30 minutes; update every 60 minutes; require a postmortem for customer-visible events.
- **SEV3 — limited event:** localized impact, a safe workaround, and no financial-integrity, security, regulatory, or material contractual risk. The owning team leads; page only if immediate action is necessary; use a ticket otherwise.
- Treat the percentage thresholds as declaration guardrails, not reasons to under-classify integrity, settlement, security, or strategic-customer risk.
- Permit any employee to declare an incident. Only the incident commander may lower severity, with the rationale logged.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of active customer harm. Declare resolution only after stability, backlog recovery, and required ledger reconciliation.
6. Define roles, authority, and handoffs (depends on: 5)
Separate command, communication, recordkeeping, and technical repair. One named person must hold command at every moment of a major incident.
- **Incident Commander:** owns severity, priorities, role assignment, escalation, decision cadence, mitigation coordination, and closure. The commander does not act as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language. Legal or Compliance retains ownership of regulatory submissions.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, and status changes.
- **Subject-Matter Responders:** diagnose and mitigate services for which they have accepted ownership, access, training, and runbooks.
- **Executive Duty Officer:** removes organizational obstacles and approves exceptional business decisions without displacing the incident commander.
- Require separate commander, communications, scribe, and primary technical lead for SEV0 and SEV1.
- For customer-visible SEV2, keep the commander separate from the primary technical responder; communications and scribe may be combined if workload permits.
- Announce every role assignment and transfer in the incident channel. Require verbal and written handoff with current impact, decisions, risks, and next actions.
- Keep executives and account managers out of the technical command path; questions flow through the Executive Duty Officer or Communications Lead.
7. Create sustainable 24×7 coverage across the service estate (depends on: 3, 6)
Use central command coverage and service-domain responder coverage rather than creating 28 fragile night rotations. Engineers remain responsible only for services they own or have formally trained to support.
- Build a company incident-command pool of approximately 18–24 certified senior engineers and managers, with a primary and backup scheduled at all times.
- Build 12–16-person communications and scribe pools using Support, Customer Operations, Engineering Operations, and qualified engineering managers.
- Schedule an Engineering Director or equivalent as the 24×7 Executive Duty Officer.
- Group related service owners into approximately 8–12 coherent responder domains only where members have training, access, and explicit acceptance.
- Require Tier 0 and Tier 1 domains to provide 24×7 primary and secondary responders, normally with at least six trained people in each sustainable rotation.
- Give Tier 2 services business-hours coverage plus a maintained manager escalation path. Treat Tier 3 conditions as tickets unless their impact changes.
- Maintain distinct but coordinated coverage for the ledger application and the PostgreSQL platform.
- Route unknown-owner events to the duty commander and platform triage temporarily. Treat every such event as an ownership control defect.
- If a small team cannot staff a fair rotation, merge coverage only after training or provide headcount, service reassignment, or decommissioning.
8. Approve compensation and fatigue protections (depends on: 7)
End unpaid on-call before expanding mandatory coverage. Publish the policy through HR, Legal, Finance, and Payroll within 14 days.
- Pay fixed weekly stipends for primary and secondary service rotations.
- Pay separate stipends for duty commander, communications, and scribe assignments.
- Provide additional call-out compensation or equivalent paid recovery time for material after-hours work.
- Apply overtime and reporting rules correctly for non-exempt employees under federal and New York requirements.
- Pay higher rates for company holidays and provide a protected recovery day after qualifying overnight work, SEV0 events, or prolonged SEV1 response.
- Target no more than one primary week in six and prohibit simultaneous primary assignments.
- Avoid consecutive primary weeks and make all swaps visible in the paging system.
- Reduce sprint commitments for people carrying primary duty rather than expecting normal delivery capacity.
- Provide a documented accommodation path for health, disability, or caregiving constraints without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
9. Set the service readiness and runbook standard (depends on: 3, 5, 7)
A team cannot respond effectively at night without ownership, access, telemetry, and rehearsed recovery procedures. Apply a formal readiness gate to every Tier 0 and Tier 1 service.
- Require a current architecture diagram, dependency map, dashboards, SLOs, runbooks, rollback method, feature-control method, contacts, and tested escalation path.
- Record RTO, RPO, data-integrity requirements, regional mode, and customer-facing capabilities in the service catalog.
- Link every paging alert to the exact runbook step expected from the responder.
- Prevent new paging alerts for services that fail readiness review. Preserve existing critical detection through a documented exception until a safe replacement exists.
- Write priority playbooks for PostgreSQL failure, ledger-integrity investigation, regional failover, Kubernetes control-plane degradation, payment-processor failure, queue backlog, credential compromise, and payment suspension.
- For the ledger, document read-only or stop-processing modes, failover controls, replay and duplicate protections, backlog handling, and post-recovery reconciliation.
- Require runbook review after material incidents and at least twice per year through exercises.
10. Establish one paging and incident system of record (depends on: 4, 5, 6, 7)
Consolidate paging and incident coordination without requiring an unsafe big-bang replacement of every monitoring system. Monitoring sources may remain specialized, but all human pages must enter one controlled platform.
- Select one enterprise paging platform and one integrated incident record, with chat, telephone, SMS, conference, status-page, ticketing, and service-catalog integrations.
- Initially ingest events from all six current tools, then deduplicate, correlate, and route by service ownership.
- Automate incident channel and bridge creation, role paging, timeline capture, severity changes, communications reminders, and postmortem creation.
- Preserve immutable records of declarations, acknowledgements, role assignments, decisions, and messages.
- Use role-based access, multifactor authentication, break-glass controls, and periodic access reviews.
- Provide telephone and offline fallback procedures for loss of chat, identity, the paging vendor, or an AWS region.
- Test paging and fallback paths weekly.
- Retire a legacy paging route only after its signals have owners, quality review, successful end-to-end tests, and at least two weeks of verified operation in the new path.
11. Enforce alert quality and burn down noise safely (depends on: 3, 10)
Treat paging alerts as production products with owners and quality requirements. Do not reduce noise by silently disabling detection.
- Require every page to identify the service, owner, customer or SLO risk, urgency, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Define an actionable page as one requiring prompt human judgment or intervention that materially reduces customer, financial, security, or contractual risk.
- Route informational, capacity-planning, and non-urgent conditions to dashboards or ticket queues.
- Prefer symptom and error-budget burn alerts over raw CPU, memory, pod, or log-volume thresholds.
- Run new alerts in shadow mode for at least seven days unless an emergency exception is approved.
- Review alerts with less than 50% actionability, more than three firings in seven days, or repeated no-action acknowledgements within two business days.
- Require compensating detection and an owner before suppressing or removing an alert.
- Set a responder load target of no more than two after-hours pages per person per week. A breach creates a mandatory alert-remediation plan.
- Review alert actionability, duplication, missed detection, and page load monthly by domain.
- Prioritize the small number of rules producing most of the current 85% noise.
12. Detect payment and ledger failures before customers (depends on: 3, 11)
Shift detection from infrastructure symptoms to customer journeys and financial outcomes. Set internal objectives stricter than the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, webhooks, and reporting freshness.
- Run external synthetic transactions from outside the production boundary and through both regions at least every minute for critical paths.
- Monitor payment failure rates, latency, queue age, delayed value, settlement-window risk, regional asymmetry, and third-party response quality.
- Add continuous ledger controls for reconciliation breaks, unexpected balances, duplicate identifiers, replication lag, backup health, and failover readiness.
- Add tenant or cohort anomaly detection for high-value customers and common payment methods.
- Convert high-priority Support, account-manager, bank, and processor reports into incident candidates within five minutes.
- Review every customer-first incident as a missed-detection defect and create a corrective action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Codify escalation and live incident execution (depends on: 6, 10)
Create one time-bound path from signal to ownership and mitigation. Delivery of a notification does not count as acknowledgement.
- Page the owning Tier 0 or Tier 1 primary immediately; page the secondary after five unacknowledged minutes; page the domain manager at 10 minutes; escalate to the engineering director at 15 minutes.
- For SEV0–SEV2, page the duty commander immediately. Page the backup after five minutes and require the Engineering Director to assume command if no certified commander owns the event by 10 minutes.
- Automatically involve command for integrity concerns, security concerns, regional events, cross-team impact, critical customer journeys, or unresolved ownership.
- If impact remains materially unknown after 15 minutes, increase response posture rather than waiting for certainty.
- Open one incident channel, bridge, and system record. State severity, known impact, assigned roles, current objective, and next update time.
- Freeze unrelated changes during SEV0 and SEV1 unless the commander records an exception.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require controlled backlog processing and reconciliation before resolving payment or ledger incidents.
- Require a formal command handoff for incidents extending beyond four hours or when fatigue impairs a role holder.
14. Standardize internal and customer communications (depends on: 5, 6, 10)
Communicate known impact early without waiting for root cause. The Communications Lead uses approved facts and always states the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish customer status within 15 minutes when customer-visible and legally safe, then update at least every 30 minutes.
- For SEV1, brief internal stakeholders and publish status within 15 minutes, then update every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send notice within 30 minutes, then update every 60 minutes.
- For SEV3, communicate directly to affected customers only when impact or contract terms require it.
- Give account managers an approved statement and affected-customer list within 30 minutes for SEV0 or SEV1 and within 60 minutes for SEV2.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. Describe affected capabilities, symptoms, workarounds, and the next update.
- Do not speculate about root cause, blame, data integrity, security scope, or recovery time.
- Post a Monitoring update within 30 minutes of mitigation. Post Resolved only after stability and required reconciliation.
- Provide a customer-facing incident summary within five business days for qualifying incidents.
- Record and approve any delay or restriction of public details during an active security threat.
15. Operationalize regulatory, partner, contract, and credit decisions (depends on: 4, 5, 14)
Give Legal and Compliance a timed decision process while keeping technical command with the incident commander. Record both reportable and non-reportable determinations.
- Build a jurisdiction and obligation matrix covering applicable NYDFS requirements, state breach laws, GLBA or FTC requirements, PCI obligations, money-transmitter rules, sponsor banks, payment networks, cyber insurance, and customer contracts.
- Validate applicability and deadlines with counsel rather than assuming every operational incident is reportable.
- Begin reportability assessment immediately for every SEV0, security event, suspected ledger-integrity event, and relevant SEV1.
- Target a documented initial legal assessment within one hour for SEV0 and within two hours for other potentially reportable events.
- Record the decision, evidence, approver, legal deadline, submission owner, and confirmation of delivery.
- Maintain tested 24×7 contacts for regulators, banks, networks, insurers, outside counsel, and critical vendors.
- Encode customer-specific notification clocks and channels in the customer record used by the Communications Lead.
- Have Finance calculate affected minutes, likely credits, and contractual exposure from the incident record within five business days.
- Review whether proactive credits or claims-based handling applies by customer segment and contract.
16. Make postmortems mandatory, consistent, and blameless (depends on: 5, 6, 10)
Use one review standard to learn from incidents and test whether controls worked. Keep learning reviews separate from performance or misconduct processes.
- Require a postmortem for every SEV0 and SEV1.
- Require one for customer-visible SEV2, customer-first detection, incidents lasting more than two hours, contractual breaches, repeat failures, control gaps, and ledger-integrity near misses.
- Produce the factual draft within three business days, conduct the review within five, and publish the approved version within 10.
- Use one template covering summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, lessons, and actions.
- Analyze why detection was not earlier and why mitigation took as long as it did.
- Examine technical, organizational, process, testing, dependency, and incentive factors rather than forcing a single root cause.
- Use a trained facilitator and written blameless rules. Describe decisions in the context and information available at the time.
- Publish broadly useful findings internally while maintaining restricted versions for security, privacy, personnel, or privileged material.
- Review postmortem quality and recurring factors in a weekly Incident Review Board.
17. Enforce action ownership and effectiveness tracking (depends on: 16)
Treat corrective actions as risk commitments rather than suggestions. Closing a ticket is insufficient without evidence that the control or system behavior improved.
- Give every action one named individual owner, manager, priority, due date, expected risk reduction, verification method, and linked work item.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with intermediate milestones.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Place recurrence-prevention actions for SEV0 and SEV1 ahead of discretionary feature work unless an executive accepts the residual risk.
- Escalate overdue high-risk actions to the manager after seven days, director after 14 days, and CTO after 30 days.
- Require written residual-risk acceptance, compensating controls, and a new review date when high-risk work is deferred.
- Verify completed actions through tests, telemetry, drills, or production evidence.
- Triage the 53 currently open historical actions within 30 days. Complete, re-plan, or formally accept the risk, prioritizing ledger, regional, security, and detection items.
18. Measure outcomes, controls, business impact, and human load (depends on: 10, 11, 14, 17)
Use a balanced scorecard so teams are not rewarded for suppressing alerts or avoiding incident declarations. Report medians and 90th percentiles, not averages alone.
- Measure time from first impact to internal detection, declaration, acknowledgement, command assignment, mitigation, and resolution.
- Track customer-first detection, missed escalations, status-page timeliness, update-cadence compliance, and role conflicts.
- Track incident count, recurrence, availability, error-budget burn, affected payment value, delayed transactions, reconciliation breaks, and SLA credits.
- Track page volume, actionability, duplicates, after-hours pages, missed acknowledgements, and load per responder.
- Track postmortem timeliness, action completion, action age, verified effectiveness, and repeated contributing factors.
- Track rotation size, duty frequency, recovery days, swaps, attrition signals, and quarterly responder sentiment.
- Hold a weekly Incident Review Board for incidents, actions, missed controls, and noisy alerts.
- Hold a monthly executive reliability review for trends, investment decisions, contractual exposure, and accepted risks.
- Hold a quarterly resilience and controls review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile dashboard data against customer cases and sampled incident records monthly to identify missing incidents or metric gaming.
19. Train participants and address pager resistance (depends on: 6, 7, 8, 13, 14, 16)
Introduce the operating model as a fair exchange, not an audit mandate. The message is that people are paid, paged only for accepted services, supported by trained command, and given capacity to fix defects.
- Train all employees to recognize impact, declare an incident, and find the incident channel and customer status page.
- Give all 260 engineers role-based training in severity, escalation, evidence preservation, handoff, and financial-integrity precautions.
- Certify incident commanders through instruction, simulation, two shadowed events or exercises, and observed performance.
- Train Communications Leads in status writing, account-manager briefing, contractual clocks, and legal escalation.
- Train scribes in timeline quality, decision capture, fact-versus-hypothesis labeling, and evidence handling.
- Require responders to demonstrate access, dashboards, rollback, runbooks, and escalation competence before primary duty.
- Require two shadow shifts before independent primary on-call.
- Appoint one adoption champion in each team and hold weekly office hours during rollout.
- Publish compensation, fatigue protections, service boundaries, and the accommodation process before assigning shifts.
- Use paid working time for training, exercises, shadowing, runbook work, and certification.
- Survey engineers at baseline, day 60, day 120, and quarterly thereafter.
20. Pilot on the payment critical path (depends on: 8, 9, 11, 12, 13, 14, 17, 19)
Run a four-to-six-week pilot across the highest-risk customer journey. Use real incidents and exercises to correct the process before wider rollout.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, authentication or API edge, settlement, and Support intake.
- Activate paid primary and secondary rotations, central command, communications, the incident record, status templates, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new system authoritative.
- Review every pilot page within one business day for routing, actionability, responder load, and missing context.
- Have the program team coach incidents without silently taking command from assigned role holders.
- Hold a weekly pilot retrospective and correct critical process or tool defects within 48 hours.
- Require 95% command assignment within five minutes, 95% timely communications, no unpaid pages, complete postmortems, and at least a 50% page-noise reduction before expansion.
21. Roll out by customer journey and risk (depends on: 18, 20)
Expand in fixed waves rather than waiting for every team to become perfect. Apply explicit readiness gates and time-limited exceptions.
- Days 0–7: operate the interim declaration, command, and communications process.
- By day 30: complete Tier 0 ownership, activate the first certified command roster, approve compensation, and begin the pilot.
- By day 60: provide compensated 24×7 coverage for every Tier 0 and Tier 1 customer journey and route all critical pages through the new platform.
- By day 90: complete critical detection upgrades, customer communications, regulatory playbooks, and the first major alert-noise reduction.
- By day 120: assign every production service a response tier, owner, tested escalation, and appropriate coverage model.
- Roll out remaining teams in two-to-three-week waves ordered by customer and dependency risk.
- Require each wave to pass ownership, alert, runbook, access, training, compensation, and tabletop gates.
- Disable legacy paging paths after verified cutover rather than leaving ambiguous parallel obligations.
- Publish a weekly adoption dashboard by team and escalate failed gates as business risks.
- Never start mandatory night coverage before compensation, staffing, training, and access are ready.
22. Exercise command, regional resilience, and ledger recovery (depends on: 9, 13, 14, 15, 19)
Validate the process under realistic conditions before depending on it during a crisis. Use the same action-tracking rules for exercises and real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops covering payment failure, customer-first detection, third-party failure, and ambiguous ownership.
- Exercise loss of one AWS region, Kubernetes degradation, PostgreSQL failure, suspected duplicate payments, queue backlog, and settlement risk.
- Exercise simultaneous operational and security events to test command and disclosure boundaries.
- Exercise loss of chat, status-page, identity, or paging providers using telephone and offline fallbacks.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors.
- Validate backups, restoration, RPO, RTO, failover prerequisites, financial controls, and post-recovery reconciliation.
- Avoid uncontrolled production-ledger experiments; use staging, replicas, simulations, or tightly governed production tests.
- Complete at least two cross-company exercises before the audit, including one overnight or unannounced paging test.
23. Test SOC 2 operating effectiveness before the auditor (depends on: 4, 18, 21, 22)
Demonstrate that the controls operate consistently, not merely that policies exist. Correct failures through tracked remediation rather than rewriting historical records.
- Preserve policy approvals, service ownership, schedules, compensation activation, access reviews, training, certifications, incidents, communications, postmortems, actions, and exercises.
- Sample evidence monthly from initial signal through verified action closure.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Use exceptions to document missed acknowledgements, late communications, incomplete records, and compensating controls.
- Conduct a formal mock audit in month 7 using the populations and interview roles expected by the external auditor.
- Trace at least one SEV0 or SEV1, one SEV2, a customer-reported event, and an exercise end to end.
- Remediate evidence and operating gaps before external fieldwork.
- Brief commanders, responders, Support, and Compliance for auditor interviews without scripting inaccurate answers.
24. Institutionalize continuous improvement (depends on: 18, 21, 23)
Keep incident management active after the audit by assigning permanent owners, budgets, and review cycles. Use incident trends to drive architectural investment.
- Assign permanent owners for the policy, service catalog, paging platform, status page, training program, metrics, and evidence repository.
- Review severity thresholds, communications timing, staffing, and compensation annually and after material process failures.
- Recertify commanders and communications leads annually through observed exercises.
- Review recurring failure families quarterly and require executive action when remediation repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fatigue, fairness, tool friction, and psychological safety.
- Report severe incidents, SLA exposure, overdue risks, resilience investments, and customer-first detection to the board or risk committee quarterly.
- Prioritize reduction of the shared-ledger concentration risk, stronger regional independence, deployment safety, graceful degradation, and automated mitigation.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
Previous Proposal 3 (ID: 2437bc94-f488-46b5-866e-01e1d4afbf3b, Agent: qwen3.8-max_refine_3, LLM: alibaba/qwen3.8-max):
Estimated Complexity: high
Success Metrics: - Median time to detect customer-impacting incidents falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median mitigation time falls from 3h10 to under 90 minutes by day 120 and under 60 minutes by month 9.
- A named incident commander is assigned within 5 minutes in at least 95% of SEV1 and SEV2 incidents; zero incidents remain unowned for more than 10 minutes.
- First status-page update occurs within 15 minutes of SEV1 declaration and 30 minutes of SEV2 declaration in at least 95% of qualifying incidents.
- Monthly paging volume falls from 3,400 to under 500 actionable pages within six months, with alert actionability above 80%.
- Out-of-hours pages average no more than 2 per responder per week; sustained breaches trigger mandatory alert remediation.
- Six legacy alerting tools are consolidated into one paging and incident platform, with legacy paging paths disabled by week 16.
- 100% of Tier 0 and Tier 1 services have a named owning team, escalation path, dashboard, and runbook by day 60.
- All 28 teams are onboarded with readiness gates by week 22; every 24x7 critical-path rotation has at least six trained responders.
- 30 or more certified incident commanders and 20 or more certified communications leads provide continuous primary and secondary coverage.
- Paid on-call is approved by HR, legal, and finance and active in payroll before any new mandatory night rotation begins.
- 100% of required SEV1 and SEV2 postmortems are drafted within 3 business days and published within 10 business days in the standard template.
- Postmortem action closure rises from 17% to at least 90% of high-priority actions completed by their due date within two quarters.
- All 53 open historical actions are triaged within 30 days; high-risk unaccepted items are completed or formally risk-accepted within 90 days.
- Repeat incidents from a known unaddressed contributing factor decline by at least 50% within six months.
- Annualized SLA credits fall from $1.3M to under $400k within 12 months.
- Monthly availability meets or exceeds 99.95% by month 6, with exceptions reviewed at the executive reliability meeting.
- At least two cross-company exercises, including regional and ledger scenarios, are completed before the SOC 2 audit, with critical findings tracked.
- The month-6 internal dry-run audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete evidence.
- SOC 2 Type II incident-response controls pass with zero exceptions at the month-8 audit.
- On-call sentiment improves quarter over quarter; fewer than 10% of engineers report unwillingness to participate in their owned-service rotation by month 6.
- No increase in attrition among engineers on active rotations compared with baseline.
Steps (27):
1. Executive mandate, program funding, and governance
Convert the CEO email into a **written mandate within 48 hours**. Name one accountable program owner and a small decision group. Time-box design to five weeks so rollout starts well before the audit.
- Appoint the CTO as executive sponsor and a Head of Reliability or Incident Management as program owner with full-time authority.
- Create an 8–10 person design working group: SRE/platform lead, payments and ledger engineering managers, support lead, compliance, legal, HR, finance, and two rotating engineering managers.
- Approve budget lines: tooling consolidation, on-call compensation, training, coaching, and 2–3 dedicated program staff.
- State the non-negotiables: paid on-call, named service ownership, one severity scale, one paging platform, mandatory postmortems, and protected engineering capacity for reliability actions.
- Fix the master timeline: interim controls week 1, design weeks 1–5, pilot weeks 6–11, full rollout weeks 12–22, internal audit rehearsal month 6, audit-ready month 7.
- Publish a one-page charter to the company that says incident response is a company operating process, not an optional team practice.
2. Evidence baseline from incidents, alerts, and coverage gaps (depends on: 1)
Before changing anything, create an auditable **before picture** from the last 12 months. This baseline drives the severity design, staffing model, and executive reporting.
- Reconstruct all 31 customer-impacting incidents: detection source, first owner, severity, mitigation time, credits paid, and whether command was clear.
- Write a specific case review of the two incidents with unclear ownership for more than one hour.
- Inventory the six alert tools, alert volume by team, noise rate, rules with no owner, and rules with no runbook.
- Map the 16 teams without on-call and the 12 teams with unpaid on-call.
- Review the 53 open postmortem actions and triage the highest-risk items first.
- Freeze baseline metrics: 22-minute detection, 40% customer-first detection, 3h10 mitigation, $1.3M credits, 3,400 alerts, 85% noise, and 17% action closure.
3. Stakeholder listening, resistance mapping, and co-design group (depends on: 1)
Treat engineer pushback as **design input, not an attitude problem**. The objection to carrying a pager for other teams' code must shape ownership, routing, compensation, and command staffing.
- Interview leads from all 28 teams plus support, customer success, sales, compliance, and HR within two weeks.
- Separate the real causes of resistance: unpaid night work, unfamiliar systems, poor runbooks, unfair load, fear of blame, or unclear authority.
- Recruit 10–15 respected engineers and managers as co-designers so the model is built with the teams.
- Survey baseline on-call sentiment, trust in alerts, and psychological safety; repeat at 90, 180, and 365 days.
- Document the central promise: responders are paged for services they own, trained incident commanders coordinate, and all on-call work is compensated.
4. Service ownership catalog, criticality tiers, and dependency map (depends on: 2)
No alert can page the right team until every service has a **named owner**. Build the catalog as the routing foundation for on-call, severity impact mapping, status-page components, and audit evidence.
- Assign one accountable team to each of the 180 services, with an engineering manager, Slack channel, escalation policy, dashboard, and runbook link.
- Tier services: Tier 0 for money movement, ledger integrity, authentication, settlement, and shared PostgreSQL; Tier 1 for customer-facing degradable services; Tier 2 for internal or batch services; Tier 3 for non-critical services.
- Map customer journeys to services, databases, regions, third parties, and contractual SLA components.
- Mark orphan and shared services; require ownership, reassignment, or decommissioning within 30 days.
- Document cross-region dependencies, failover constraints, and services that defeat nominal regional redundancy.
- Treat missing ownership for Tier 0 or Tier 1 services as an executive escalation and a release-blocking risk.
5. Severity scale, declaration rights, and automatic triggers (depends on: 2, 4)
Adopt one severity scale so declaration is a lookup, not a debate. **Anyone may declare; only the incident commander may downgrade.** When uncertain, start higher.
- SEV1: money movement stopped or incorrect, ledger integrity in doubt, confirmed security or data event, both regions impaired, or broad SLA-credit exposure.
- SEV2: major degradation, settlement window at risk, one or more critical customers fully down, or likely contractual breach.
- SEV3: limited impact with a workaround, no financial-integrity or security risk.
- SEV4: no customer impact; handled by ticket during business hours.
- Define fixed triggers for each level: pages, command staffing, bridge call, status-page timing, account-manager outreach, executive notification, and postmortem obligation.
- Add automatic escalation: SEV3 open more than 2 hours becomes SEV2; any incident touching the shared ledger or unclear ownership becomes SEV2 unless the commander documents otherwise.
- Publish a decision tree and 10–12 worked examples from the actual 31 incidents.
6. Incident roles, authority, and handover rules (depends on: 5)
Solve the **nobody-was-in-charge failure** by making command explicit, trained, and transferable. Separate coordination from technical remediation so responders are not asked to debug unfamiliar code.
- Incident Commander: owns severity, priorities, escalation, mitigation strategy, role assignment, handoffs, and closure. Does not write code during the incident. May freeze deploys, invoke pre-approved failover, and pull any on-call responder.
- Communications Lead: owns status page, internal updates, account-manager briefs, and coordination with legal or compliance.
- Scribe: maintains the timestamped record of decisions, actions, role changes, and customer communications.
- Subject-matter responders: engineers from owning teams who diagnose and remediate within their accepted domain.
- Executive liaison: required for SEV1; shields the commander from executive questions and owns board or regulator escalation.
- Require the commander to claim the role within 5 minutes of SEV1 or SEV2 declaration, state it in the incident channel, and record every handover.
- Allow role combination only for SEV3 or SEV4; require separate people for command, communications, and primary technical work at SEV1.
7. Paid on-call, fatigue safeguards, and HR compliance (depends on: 3)
Unpaid on-call is a retention, fairness, and legal risk in New York. **Pay must be approved before any new mandatory rotation starts.** This is the fastest way to reduce pager resistance.
- Create a weekly stipend for primary and secondary on-call, differentiated by tier and night responsibility.
- Add per-page or per-incident pay for out-of-hours activation, plus guaranteed recovery time after night work or major incidents.
- Create a separate stipend for duty incident commanders and communications leads because command is a heavier burden.
- Verify FLSA, New York wage-hour, overtime, holiday, and payroll treatment with HR, legal, and finance.
- Cap rotation frequency: no more than one primary week in four or one week in six depending on staffing; no simultaneous primary assignments.
- Require at least six trained responders for any 24x7 rotation; fund hiring or service reassignment where teams are too small.
- Publish the compensation package and payroll start date before asking engineers to join rotations.
8. On-call staffing model and rotation rules (depends on: 4, 6, 7)
Do not force 28 identical night rotations. Use a **layered model**: central command coverage, critical-path team coverage, and business-hours coverage for lower-tier services.
- Company incident-command and communications rotation: 24x7 pool of 25–35 trained volunteers and designated senior staff, giving primary plus secondary coverage at all times.
- Critical-path teams: 24x7 primary and secondary on-call for Tier 0 and Tier 1 domains, including payments, ledger, authentication, API edge, Kubernetes platform, and shared PostgreSQL.
- Other teams: business-hours on-call with a documented night escalation list owned by the engineering manager.
- Group services into 8–12 coherent domains so rotations are sustainable; no rotation with fewer than six people is approved without an executive exception.
- Require handoff overlap, shadow shifts before first primary duty, and no on-call during approved PTO.
- Define cross-team pull rules: the commander may page another team's on-call with a 10-minute acknowledgement obligation; this is coordinated by command, not pushed onto responders.
- Publish schedules, swap rules, and load limits in the paging tool.
9. Alert quality standard, page budget, and noise controls (depends on: 2, 4)
With 3,400 alerts and 85% noise, detection fails because people stop trusting pages. Make alert quality a **condition of paging anyone**.
- Every page must have a named owning team, customer or SLO impact, severity, dashboard, runbook, expected action, and escalation policy.
- Page on customer-visible symptoms: payment success rate, latency, settlement deadlines, error-budget burn, ledger integrity, and replication health.
- Demote cause-based CPU, memory, or infrastructure-only alerts to dashboards or tickets unless they map to a customer journey.
- Set a page budget: maximum 2 out-of-hours pages per person per week; breach triggers a mandatory alert-tuning sprint.
- Auto-flag alerts that fire repeatedly without action or have high no-action acknowledgement rates.
- Require shadow mode for new alerts before they page humans, except for documented emergencies.
- Review alert quality monthly by team and publish a noise leaderboard.
- Target fewer than 500 actionable pages per month and above 80% actionability within six months.
10. Detection uplift across payments, ledger, and customer signals (depends on: 4, 9)
Customers detected 40% of incidents first. Detection must shift to **payment outcomes, ledger integrity, and inbound customer signals**, not host metrics alone.
- Define SLIs and SLOs for payment initiation, authorization, settlement, reconciliation, refunds, API availability, and reporting freshness.
- Run synthetic end-to-end payment tests from outside the platform in both AWS regions every 60 seconds.
- Add continuous ledger assurance: double-entry reconciliation, replication lag, failover readiness, disk pressure, and settlement-window countdown alerts.
- Monitor per-customer anomalies for top accounts so a single-tenant outage is detected before the account manager calls.
- Convert high-priority support tickets and account-manager reports into incident candidates within 5 minutes.
- Track detection source for every incident; make customer-first detection a reviewed defect with a corrective action.
- Add partner, banking, and card-network notification intake into the same declaration path.
11. Single incident platform and alert-tool consolidation (depends on: 5, 8, 9)
Collapse six tools into **one paging and incident-management system** with one queue, one timeline, and one audit trail. Do not create a parallel audit process.
- Select an integrated stack: paging and on-call scheduling, incident workflow, Slack or chat integration, bridge calling, status-page API, and ticketing integration.
- Implement one-command declaration that creates the incident record, channel, bridge, severity label, role prompts, and clock.
- Migrate alert sources by wave; retire a legacy paging path only after routing, ownership, and acknowledgement tests pass.
- Automate evidence capture: timestamps, acknowledgements, role assignments, severity changes, communications, and postmortem links.
- Integrate with the service catalog, on-call schedules, Jira or equivalent action tracker, customer-success tooling, and the status page.
- Verify the platform works during a single-region failure, including SMS, phone, and offline fallback paths.
- Set a hard date after which pages outside the chosen platform are not valid on-call obligations.
12. Escalation paths and five-minute command rule (depends on: 6, 8, 11)
Create one path from **signal to named commander in under five minutes**, any hour of the day. Make escalation automatic and time-bound.
- Accept declarations from automated alerts, engineers, support, account managers, partners, and customers through the same command.
- For SEV1 and SEV2, page the duty incident commander and owning-team primary immediately.
- Acknowledgement ladder: primary 5 minutes, secondary 10 minutes, manager 15 minutes, director or executive 20 minutes.
- If no commander claims the incident within 5 minutes, the platform assigns and announces one; the assignee may hand over but cannot leave the incident unowned.
- For SEV3, require team acknowledgement within 30 minutes; otherwise create a tracked work item.
- Give the commander pre-approved authority to invoke regional failover, ledger read-only mode, feature kill switches, and partner notifications without waiting for executive sign-off.
- Record every missed acknowledgement and escalation failure for weekly review.
13. Internal communications protocol (depends on: 6, 11)
Separate the **working incident channel from the audience channel** so responders can work and executives, support, and sales stay informed without interrupting the commander.
- Create one incident channel and one bridge per SEV1–SEV3 incident; use a read-only broadcast channel for executives and support.
- Set update cadence: every 15 minutes for SEV1, 30 minutes for SEV2, and at state changes for SEV3.
- Use a fixed update template: impact, customer-visible symptoms, current action, ETA or next update, commander, and communications lead.
- Brief support and customer success with a live affected-customer list and approved holding statements within 15 minutes of SEV1 or SEV2.
- Require executives to route questions through the executive liaison; the commander is not interrupted.
- Define fatigue rules and formal handover for incidents lasting more than 4 hours.
14. Customer status page, account-manager outreach, and SLA credit workflow (depends on: 5, 13)
Replace ad-hoc status updates with a **timed, owned, template-driven customer communication process**. Link incidents to SLA credits so finance and customer success are not surprised.
- Publish status-page updates within 15 minutes of SEV1 declaration and 30 minutes for SEV2; update every 30 or 60 minutes until resolution.
- Assign the communications lead as the single author; use pre-approved templates reviewed by legal and communications.
- Map status-page components to customer capabilities: payments, settlement, reporting, API, onboarding, and regional availability.
- For top accounts, require direct account-manager outreach within 30 minutes of SEV1 with an approved briefing pack.
- Send proactive email or webhook notifications to subscribed customers for SEV1 and SEV2.
- Publish a resolution notice within 30 minutes of mitigation and a customer-facing incident summary within 5 business days for SEV1.
- Automate SLA credit calculation from incident duration and affected capability; review with finance and legal within 5 business days.
- Track credits by incident and root-cause family to guide reliability investment.
15. Regulatory, legal, and partner notification playbook (depends on: 5, 14)
In payments, some incidents are reportable and the clock starts at detection. Make regulatory assessment a **mandatory step in the incident process**, not an afterthought.
- Map obligations: NYDFS cybersecurity-event rules, state breach laws, GLBA safeguards, PCI DSS if in scope, money-transmitter duties, card-network and sponsor-bank contracts, and cyber-insurer notice.
- Require a legal or compliance reportability assessment within 2 hours for every SEV1 and every security-related SEV2, even when the answer is not reportable.
- Maintain a 24x7 contact matrix for regulators, sponsor banks, card networks, outside counsel, insurers, and law enforcement.
- Pre-draft notification templates and preserve legal review.
- Encode customer-contract notification deadlines into account tiers and the communications workflow.
- Record every reportability decision, approver, deadline, and submission confirmation in the incident record.
16. Postmortem standard and blameless review (depends on: 5, 6)
Replace inconsistent postmortems with a **mandatory, blameless, single-format process**. The discipline comes from deadlines, facilitation, and action tracking.
- Require postmortems for all SEV1 and SEV2 incidents, customer-detected incidents, incidents over 2 hours, repeat failures, and ledger near-misses.
- Draft within 3 business days, peer review within 5, publish within 10 for SEV1 and SEV2.
- Use one template: summary, impact, timeline, detection analysis, response analysis, contributing factors, what worked, what failed, and action items.
- Make blamelessness explicit: focus on systems and decisions, not individual fault; never use postmortems in performance discipline.
- Hold a weekly incident review board to review postmortems, ratify severity, and challenge weak actions.
- Maintain a searchable postmortem library and quarterly recurring-cause analysis.
- Require a trained facilitator for major reviews; the incident commander attends but does not facilitate.
17. Action-item tracking, ownership, and delivery gates (depends on: 16, 11)
Only 11 of 64 actions were closed. Give postmortem actions the same status as **customer commitments**, with named owners and visible escalation.
- Create every action as a ticket with one named individual owner, priority, due date, and verification method.
- Use delivery classes: containment within 7 days, corrective work within 30 days, strategic work within 90 days.
- Reserve 15–20% of team sprint capacity for reliability and incident actions.
- Escalate overdue items: manager at 7 days, director at 14 days, CTO dashboard at 30 days.
- Block related feature releases when overdue P0 actions prevent recurrence of a severe incident.
- Require director approval and documented residual risk acceptance for overdue high-risk items.
- Verify effectiveness after completion; closing a ticket without evidence does not close the action.
- Target 90% of high-priority actions completed on time within two quarters.
18. Runbooks, critical-incident playbooks, and readiness bar (depends on: 4, 8)
Poor runbooks are a real cause of pager resistance and slow mitigation. Define a **minimum readiness bar** before a service is allowed to page anyone at night.
- Require for every Tier 0 and Tier 1 service: architecture summary, dependencies, dashboards, alert-to-runbook map, rollback procedure, feature flags, escalation contacts, and customer-impact statement.
- Write major playbooks for shared PostgreSQL ledger failure, regional failover, Kubernetes control-plane loss, payment-processor outage, settlement-window breach, duplicate-payment suspicion, and security compromise.
- Define ledger recovery rules: failover procedure, read-only degraded mode, reconciliation, RPO/RTO, and data-loss tolerance approved by executives.
- Test runbooks in drills at least twice a year; mark untested runbooks stale.
- Prevent paging alerts for services without readiness sign-off unless the engineering manager accepts the gap in writing.
- Keep runbooks linked from every alert and incident template.
19. Training, certification, and role readiness (depends on: 6, 12, 13, 16)
Command and communications are skills. Build a **tiered certification path** so rotations are staffed by people who have practiced, not by whoever is around.
- All employees: 1-hour module on declaring incidents, finding the incident channel, and reading the status page.
- All responders: half-day training on severity, acknowledgement, escalation, runbooks, and evidence hygiene.
- Incident commanders: 2-day course plus two shadowed incidents and one simulation before certification.
- Communications leads: training on status-page writing, customer language, account-manager briefs, and regulatory triggers.
- Scribes: training on timeline discipline and audit evidence.
- Certify for 12 months; renew through a simulation.
- Require shadow shifts before independent primary duty; no new hire holds primary within 90 days.
- Publish the certification register as an audit artifact.
20. Simulation program and game days (depends on: 19, 11, 18)
Rehearse the process before it meets a real SEV1. Simulations build commander confidence, expose runbook gaps, and produce audit evidence.
- Run monthly 60-minute tabletops using real incidents from the 31-incident baseline.
- Run quarterly game days covering regional failover, ledger replica promotion, dependency failure, partner outage, and security event.
- Run twice-yearly unannounced paging drills to measure night acknowledgement times.
- Include support, account managers, legal, compliance, and executives in at least one exercise per quarter.
- Produce tracked action items from every exercise using the same board as real incidents.
- Measure time to commander, time to first status update, and time to mitigation decision.
21. Critical-path pilot and gate review (depends on: 7, 10, 11, 14, 18, 19)
Prove the model on the highest-risk services with willing teams before full rollout. Run a **six-week pilot with daily feedback and public exit criteria**.
- Pilot with payments, ledger/database, platform/Kubernetes, API edge, authentication, and support intake.
- Activate severity scale, duty commanders, paid rotations, single paging platform, alert budget, status-page policy, and postmortem process.
- Hold a weekly pilot retrospective and fix process defects quickly.
- Validate night acknowledgement, cross-team pull response, severity clarity, and compensation payroll.
- Exit gate: commander assigned within 5 minutes in 95% of incidents, status page on time, page noise down at least 50%, postmortems on time, and positive on-call sentiment.
- Publish pilot results to the whole company as the main adoption argument.
22. Wave rollout across all 28 teams (depends on: 21)
Roll out by criticality and dependency, not by calendar alone. Use **readiness gates** so teams are not forced live without coverage.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1 teams. Wave 3: Tier 2 teams. Wave 4: Tier 3 and internal platforms.
- Gate per team: catalog entry complete, alerts migrated, runbooks ready, six-person rotation staffed, one commander candidate nominated, compensation in payroll, and one tabletop passed.
- Assign a program coach to each wave for three weeks.
- Freeze legacy paging paths for each team after successful onboarding.
- Publish a live adoption scoreboard by team.
- Complete all teams by week 22, leaving months of operating evidence before the audit.
23. Metrics, dashboards, and review cadence (depends on: 11, 16, 21)
Instrument the process itself. Leadership must see whether the program is working, and auditors must see **operating evidence, not retrospective paperwork**.
- Track response metrics: MTTD, time to declare, time to commander, acknowledgement time, mitigation time, resolution time, and customer-first detection rate.
- Track quality metrics: page volume, alert actionability, missed pages, postmortem timeliness, action closure rate, and status-page compliance.
- Track business metrics: availability against 99.95%, SLA credits, repeat incidents, and error-budget consumption.
- Track people metrics: on-call load, night pages per person, recovery time usage, sentiment, and attrition signals.
- Hold weekly incident review, monthly reliability review, quarterly executive review, and annual policy review.
- Publish dashboards internally and give every metric a target and owner.
24. Change management, incentives, and culture (depends on: 3, 7, 21)
The process will be judged on fairness. Communicate the deal repeatedly and make participation **recognized, compensated, and safe**.
- Core message: you are paid, paged only for what you own, supported by a trained commander, and given real capacity for actions.
- Run CTO all-hands, team roadshows, office hours, FAQ, and an internal incident-management hub.
- Add incident response and reliability work to promotion criteria and manager objectives.
- Recognize good postmortems, alert-noise reduction, and calm incident leadership.
- Provide a written path for engineers who cannot do nights; cover those shifts with paid volunteers or adjusted staffing.
- Prohibit retaliation for good-faith declaration or escalation.
- Publish sentiment survey results, including bad news, to maintain credibility.
25. SOC 2 evidence design and internal dry-run audit (depends on: 16, 17, 22, 23)
Design the process so audit evidence is a **by-product of normal operations**. Test it internally before the external auditor does.
- Map the process to SOC 2 criteria: incident identification, response, recovery, monitoring, communication, control activities, availability, and corrective action.
- Approve versioned policies: incident response policy, severity standard, on-call policy, communications policy, postmortem standard, and exception process.
- Retain incident records, paging logs, status-page history, postmortems, action tickets, training records, drill records, and access reviews for the audit period.
- Record regulatory reportability decisions even when no notification is required.
- Run an internal dry-run audit at month 6: sample at least 15 incidents and trace evidence end to end.
- Fix gaps with at least four weeks before the audit.
- Maintain an exception and remediation log instead of claiming perfection.
26. Program risk register and contingency planning (depends on: 1)
Name the likely failure modes now and pre-commit responses. Review the register monthly with the executive sponsor.
- Commander volunteer shortfall: roster command duty for senior engineers and managers until the pool reaches 25 trained people.
- Compensation delay: use immediate time-off-in-lieu plus phased stipend, but do not start mandatory night rotations without pay.
- Tool migration slip: preserve single paging and routing scope; defer advanced automation if needed.
- Alert pruning causing missed incidents: move noisy alerts to ticket first, observe 30 days, then delete.
- Burnout in experienced teams: monitor load weekly and cap pages per person.
- Major SEV1 during rollout: program owner shifts to incident support, wave schedule slips one wave, sponsor is informed same day.
- Shared ledger concentration: track blast-radius reduction and failover improvements as top strategic actions.
27. Continuous improvement, maturity roadmap, and post-audit sustainability (depends on: 23, 25)
Prevent the classic failure where the process decays after the audit. Build the second-year operating plan before the first year ends.
- Hold quarterly process retrospectives with commanders, responders, support, and communications leads.
- Re-baseline targets every six months; tighten goals once current targets are met.
- Move from lagging metrics to leading indicators: error-budget burn, near-miss rate, drill performance, and action ageing.
- Plan year-two improvements: follow-the-sun coverage, automated mitigation, ledger blast-radius reduction, error-budget release gates, and per-customer real-time impact reporting.
- Keep annual policy review, certification renewal, drills, and board reporting on a permanent calendar independent of the audit cycle.
- Report quarterly to the board or risk committee on availability, credits, severe incidents, overdue actions, and resilience investment.
Previous Proposal 4 (ID: 185429cf-5cd1-451d-ab7e-34d64ca403df, Agent: grok4.6_refine_4, LLM: xai/grok-4.6):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 10 minutes by month 3 and under 5 minutes by month 9.
- Customer-first detection falls from 40% to under 20% by month 3 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by month 4 and under 60 minutes by month 12.
- A named Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes after week 4.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases by month 3, with required update cadence met in 95% of cases.
- Monthly pages fall from 3,400 to under 500 within 6 months, with actionability above 75%; out-of-hours pages at or under 2 per person per week.
- All six legacy alerting tools route through one paging platform by week 16; legacy paging paths disabled per wave.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by due date within two quarters; all 53 currently open historical actions triaged within 30 days of S2.
- SLA credits fall from $1.3M to under $400k in the first 12 months; customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known unaddressed cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar; every Layer B rotation has 6+ certified responders or a time-limited executive exception.
- Paid on-call policy approved by HR, Legal, and Finance and in payroll before any mandatory night rotation starts.
- All 28 teams onboarded by week 24 with 24x7 command cover from 25+ certified ICs and 15+ certified Comms Leads.
- Internal dry-run at month 6 passes a 15-incident evidence walkthrough; month-7 mock audit finds no unowned high-risk control gap; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment improves quarter over quarter; no increase in attrition among engineers on rotation; pulse at 60 and 120 days shows at least 70% agree rotations are fair and limited to services they own.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked actions; two cross-company exercises completed before the audit.
Steps (24):
1. Program charter, mandate, and funding
Convert the CEO email into a named program with one owner, a budget, and a deadline earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, 28 teams will opt out.
- Appoint a Program Lead (Head of Reliability) with direct CTO and CEO sponsorship.
- Form a small steering group: CTO, VP Eng, Head of CS/Support, CISO, Legal, Finance, HR. Not a 28-team committee.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Timeline: operating floor in 7 days; design weeks 1–6; pilot weeks 7–12; rollout weeks 13–24; mock audit week 28; SOC 2 at month 8.
- Fund tooling ($150–250k/yr), on-call pay (~$600k–1M/yr), 2–3 program FTEs, and reserved engineering capacity. Anchor the ask against $1.3M in credits plus unmeasured incident cost.
- Freeze baselines now: 31 incidents, MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, $1.3M credits, 3,400 alerts/month at 85% noise, 11 of 64 actions closed.
2. Immediate 7-day operating floor (depends on: 1)
Do not wait for tooling, compensation, or the audit. Put a minimum process in place this week so the next outage has a named commander.
Start retaining artifacts on day one. This week becomes the first audit evidence.
- Publish a one-page interim severity guide and a single declaration path through Slack, phone, and pager.
- Staff interim primary and backup Incident Commander 24x7 from existing on-call veterans and engineering managers. Compensate this duty retroactively.
- Require a named IC within 10 minutes for every suspected major incident. If nobody claims it, the duty manager is IC.
- Use one channel, one bridge, one timeline doc, and one naming convention for every major incident.
- Direct Support to escalate credible customer reports immediately. They do not wait for engineering confirmation.
- Triage all 53 open historical actions. Close, re-plan, or formally risk-accept. Ledger integrity, payment duplication, regional resilience, security, and detection come first.
- Hold a daily 15-minute ops review until the permanent process is live.
3. Forensic baseline of incidents and alerts (depends on: 1)
Rebuild the facts before locking design. This is the before picture for the CEO and the auditor.
Every later design choice should trace to this evidence.
- Re-code each of the 31 incidents: trigger, service, detection source, timestamps, who led, credits paid, root-cause family.
- Quantify the 40% customer-first detections and name the missing signal in each case.
- Reconstruct the two nobody-in-charge incidents minute by minute. Use them as the burning-platform story.
- Audit the six alerting tools: volume per tool and team, top 50 noisy rules, rules with no owner or runbook.
- Freeze the baseline numbers. Do not let them drift during design.
4. Listening tour and the fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the main delivery risk. Treat it as a design constraint, not an attitude problem.
The answer that will actually stick is **the deal**: you are paid; you are paged only for services you own; a trained commander runs the room; your postmortem actions get real sprint capacity.
- Interview all 28 teams plus Support, CS, and Sales in two weeks.
- Separate the objections: unpaid work, nights, unfamiliar code, bad runbooks, fear of blame. Each needs a different fix.
- Collect informal practices from the 12 teams already on-call. They are the pilot candidates and the veteran pool.
- Recruit 10–15 credible engineers as a design working group so the process is co-authored.
- Baseline sentiment on alert trust, on-call willingness, and burnout. Re-measure at 6 and 12 months.
5. Service ownership catalog and criticality tiers (depends on: 3)
You cannot page the right person across 180 services until each one has a named owner. This is the foundation of fairness, routing, and audit evidence.
Build a machine-readable catalog as the single source of truth.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, the shared PostgreSQL cluster, queues, partner banks, and customer-facing endpoints.
- Assign one owning team, named engineering manager, Slack channel, escalation policy, and dependency list per service.
- Tier 0: money movement, ledger, auth, shared Postgres, regional control plane. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map Tier 0/1 services to customer capabilities: initiation, authorization, settlement, reporting, onboarding.
- Assign coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
6. Severity scale and declaration rules (depends on: 3, 5)
Replace in-the-moment debate with a lookup. Four payments-specific levels, with worked examples drawn from the 31 real incidents.
Anyone may declare. Only the Incident Commander may downgrade, with recorded rationale. When unsure, start high.
- **SEV1**: money movement stopped or incorrect; ledger integrity in doubt; material security or data exposure; both regions impaired; or more than 10% of customers impacted. Pages IC, comms, scribe, SMEs, and exec liaison. Bridge in 5 minutes. Status page in 15 minutes. Mandatory postmortem and regulator assessment.
- SEV2: severe degradation; settlement window at risk; a strategic customer fully down; SLA breach likely. IC and SMEs paged. Status page in 30 minutes. Mandatory postmortem.
- SEV3: partial impact with a workaround; no credit exposure. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, or a repeat.
- SEV4: internal or minor. Ticket only. No page.
- Auto-escalate: any SEV3 open more than 2 hours, or any incident touching the shared ledger, becomes SEV2. Unknown impact after 15 minutes is raised, not sat on.
- Nobody is punished for over-declaring. Publish that rule in writing and repeat it.
7. Roles, authority, and ledger dual-control (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, transferable, and logged.
The Incident Commander owns the incident, not the fix, and **never types** in a production terminal.
- Incident Commander: declares severity, pulls anyone, freezes deploys, invokes failover, authorises spend. One IC at a time. Assumed within 5 minutes and announced in channel.
- Communications Lead: single voice to customers, status page, account managers, and the exec summary.
- Scribe: timestamped timeline, decisions, and open questions. Feeds the postmortem and the audit trail. Required for SEV1 and SEV2.
- Subject-matter responders: diagnose and mitigate only services they own, with access and runbooks.
- Executive Liaison (SEV1): shields the IC from exec questions; owns regulator and board escalation.
- The IC may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access rules.
- Handover is verbal and written, with time and new owner recorded. Roles may combine below SEV2; never at SEV1. The IC stays in charge if a VP joins.
8. Three-layer 24x7 coverage model (depends on: 5, 7)
Do not put 28 teams on night rotation. That is what engineers are rejecting.
Staff command centrally. Page engineers only for code their team owns.
- Layer A, company command: Duty IC plus secondary, Duty Comms, and a scribe pool. Target 25–35 certified ICs and 15–20 comms leads. Roughly one week per person every 6–8 months. Five-minute ack SLA, then secondary, then on-call director.
- Layer B, Tier 0/1 domains: group 28 teams into 8–12 product and platform domains (ledger app, Postgres platform, payments orchestration, auth, API edge, Kubernetes/platform, settlement). Primary plus secondary 24x7. Minimum six trained people. Target one week in six, never worse than one in four.
- Layer C, Tier 2/3: business-hours on-call. After hours the IC pages the EM, who holds a written escalation list.
- No engineer joins another team's responder pool without training, access, runbooks, shadow shifts, and both teams' acceptance.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner.
- Evaluate follow-the-sun coverage as a 12-month option, not a year-1 dependency.
9. Paid on-call, NY labor compliance, and fatigue rules (depends on: 4, 8)
Unpaid on-call is a retention risk and a New York legal exposure. Pay must be live in payroll before any mandatory night rotation starts.
Treat interrupted nights and recovery time as compensable work.
- Weekly stipend by layer, higher for 24x7 Tier 0/1 and Duty IC, lower for business-hours, with holiday and weekend premiums.
- After-hours call-out pay or TOIL. A paid recovery day after prolonged overnight work, a SEV1, or a qualifying SEV2. Managers cover the next day.
- HR, Legal, and Finance publish dollar amounts, FLSA exempt/non-exempt treatment, NY wage-hour rules, tax treatment, and payroll timing within 14 days of charter.
- Load rules: no primary on two rotations; no consecutive primary weeks; no on-call the week after a SEV1 you commanded.
- A person may declare temporarily unfit after overnight work with no performance penalty.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- If a rotation cannot staff six people, merge domains or hire. Do not run two-person 24x7.
- Model annual cost against $1.3M in credits and get it as a CFO/board line item.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts a month at 85% noise is why detection takes 22 minutes. Make quality a condition of paging a human.
A page is a product with a quality bar, not a dump of host metrics.
- Every paging alert must have a named owner, customer or SLO impact, runbook, tested threshold, severity mapping, dashboard, expected action, and dedup key. Fail any of these and it becomes a ticket or is deleted.
- Page on customer symptoms: SLO burn, error budget, settlement-queue age. Cause-based CPU and memory alerts become dashboards or tickets.
- Set a **page budget** of at most two out-of-hours pages per person per week. A breach triggers a mandatory tuning sprint and blocks new paging alerts for that team.
- Auto-quarantine alerts that fire more than five times a month with no action, or that have more than 70% no-action acks. Never silently disable without compensating detection and a recorded decision.
- Run new alerts in shadow for seven days unless an emergency exception is approved.
- Monthly per-team kill, tune, or keep review. Target under 500 pages a month and actionability above 75% within six months.
11. Customer-journey detection and ledger assurance (depends on: 5, 10)
Stop customers being the monitoring system. Detect on money-path outcomes, not host metrics.
Every postmortem will ask why a customer saw it first.
- Define SLOs per Tier 0/1 capability: initiation success, auth latency, settlement timeliness, API availability, reporting freshness. Set an internal target stricter than 99.95%.
- Run synthetic full-lifecycle payments from outside the platform, in both regions, every 60 seconds. Include a small real-value canary where legally feasible.
- Add ledger assurance: continuous double-entry reconciliation, replication lag, failover-readiness, unexpected balances, duplicate identifiers, settlement-window countdown.
- Add per-customer anomaly detection for the top 100 accounts.
- Auto-create a triage incident within five minutes when support tickets or AM reports match impact keywords.
12. Single incident and paging platform (depends on: 6, 8, 10)
Collapse six alerting tools into one queue, one timeline, and one audit record.
The platform must still work if one AWS region is down. Verify SMS and phone fallback, plus an offline runbook.
- Select one paging product, one incident layer, and one hosted status page.
- One Slack command declares the incident, creates the channel and bridge, pages the Duty IC, sets severity, and starts the clock.
- Ingest the six existing tools first. Deduplicate and route. Retire a legacy path only after named owners and two weeks of verified operation.
- Auto-capture timestamps, acknowledgements, role assignments, severity changes, comms, mitigated time, and resolved time. Export for SOC 2.
- Integrate the service catalog, Jira, Salesforce or CS tooling for affected-customer lists, and Zoom or Slack Huddle.
- Test paging, escalation, status publication, and conference access every week.
- After a team's wave, old paging paths are disabled, not left as a fallback.
13. Detection-to-command escalation path (depends on: 7, 8, 12)
Write the single path from something looks wrong to someone is in charge. Target: a commander in under five minutes, any hour.
If nobody claims IC in five minutes, the platform assigns the Duty IC. The assignee can hand over, not decline.
- Converge every entry point on the same declare command: alert, engineer, support, account manager, partner bank, SEV hotline.
- SEV1: page the owning primary immediately; secondary at 5 minutes unacked; domain manager and Duty IC at 10; exec liaison at 15.
- SEV2: primary ack in 10 minutes; IC assigned in 15.
- Human acknowledgement is required. Delivery to a device does not count.
- The IC can page any team's on-call, with a 10-minute ack obligation. This reciprocity makes single-team ownership viable.
- Unowned alerts go to Layer A command, then the missing owner record is a control defect.
- Pre-authorise regional failover, ledger read-only mode, and partner-bank notice so the IC does not wait for an executive. Dual-control still applies to ledger writes.
14. Live execution and major-incident playbooks (depends on: 7, 13)
Limit customer and financial harm before proving root cause. One procedure from the first minute to handback.
A service cannot page at night until it meets the readiness bar.
- Open channel, bridge, record, and timeline immediately for SEV1 and SEV2. The IC states severity, known impact, hypothesis, objective, roles, and next update time.
- Freeze unrelated production changes during SEV1. Record exceptions the IC approves.
- Prefer reversible mitigation: rollback, feature flag, traffic isolation, rate limit, partner reroute.
- Write playbooks first for Postgres ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank outage, settlement-window breach, suspected security compromise, and suspected duplicate payments.
- Guard against split-brain, replay, and duplication during regional or database recovery. Reconcile and process backlog before calling the incident resolved.
- Mitigation means customer impact ended. Resolution means stable, backlog processed, and ledger reconciled.
- Formal IC handover after 4 hours. Reopen if impact recurs during the stability window.
- Readiness bar per Tier 0/1 service: diagram, dependencies, dashboard, runbook, rollback, kill switch, escalation contacts, RPO/RTO. Untested runbooks are marked stale.
15. Internal, customer, and regulatory communications (depends on: 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved process. The Communications Lead is the single author.
State impact and the next update time. Never speculate on cause.
- Internal: one working channel, one bridge, one read-only broadcast for execs, Support, and Sales. SEV1 update every 30 minutes even if unchanged; SEV2 every 60 minutes.
- Executives ask questions only of the Executive Liaison. Publish this as a signed exec behaviour rule.
- Status page: SEV1 within 15 minutes, SEV2 within 30 minutes; then 30/60-minute updates; resolve notice within 30 minutes of mitigation. Templates pre-approved with Legal.
- Top 100 accounts: named AM contact within 30 minutes of SEV1, with a briefing pack from Comms. Long tail gets the status page plus email or webhook.
- Customer-facing summary within 5 business days for SEV1.
- Mandatory regulatory checkpoint on every SEV1 and every security SEV2, within 2 hours, recorded even when not reportable. Cover NYDFS Part 500 72-hour clock, state breach laws, GLBA/FTC, PCI if in scope, sponsor-bank and card-network windows, FinCEN/OFAC if relevant.
- Legal owns outbound regulatory letters. The IC owns facts. Security or Legal may limit public detail during an active threat, with the reason recorded.
- Encode bespoke customer-contract notice SLAs into account tiering.
16. SLA credit and financial-impact workflow (depends on: 6, 15)
Link incidents to money so severity, credits, and investment stay consistent. Finance should not learn about outages from invoices.
Make credit calculation an output of the incident record, not a negotiation.
- Agree availability measurement per contract and component with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and capability telemetry. Produce a proposed credit schedule within 5 business days of resolution.
- Posture: proactive credits for top-tier accounts; claims-based for the rest. Document the approval chain.
- Track credits per incident and root-cause family. Quarterly report which reliability investments would have prevented which credits.
- Target: cut credits from $1.3M to under $400k in 12 months. Use that delta as the ongoing business case.
17. Blameless postmortems and action tracking (depends on: 6, 12)
Eleven of 64 actions closed is the clearest process failure. Make learning mandatory and actions as binding as customer commitments.
Closing a ticket without evidence of effectiveness does not close the action.
- Mandatory for every SEV1 and SEV2, every customer-first detection, every incident over 2 hours, every repeat of a known cause, and every ledger near-miss.
- Draft within 3 business days, review within 5, published internally within 10. The IC owns the draft. The owning EM is accountable.
- One template: timeline, customer and financial impact, why detection was late, why mitigation took that long, contributing factors, what went well, actions.
- Blameless in writing. No individual named as a cause. HR and management commit that postmortems are never used in performance reviews.
- Every action gets a named person, priority, due date, Jira ticket, and verification method. P0 (prevents SEV1 recurrence) due in 30 days and committed into the next sprint before roadmap work. P1 in 60 days. P2 in 90 days.
- Teams reserve 15–20% of sprint capacity for reliability and incident actions.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Weekly Incident Review Board with engineering directors. Target 90% of P0/P1 actions closed on time within two quarters.
18. Training, certification, and commander academy (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Nobody holds independent duty untrained. Training happens in paid working time.
The certification register is an audit artefact.
- All employees, 1 hour: recognise impact, declare, find the channel and status page.
- Responder, half day: severity, escalation, runbooks, comms hygiene. Mandatory before joining a rotation.
- Scribe, 2 hours: timeline discipline. The entry point.
- Incident Commander, two days plus shadowing: command presence, decisions under uncertainty, severity, handover, exec management. Two shadowed incidents and one simulated SEV1 before certification.
- Communications Lead, one day: status writing, customer tiering, legal boundaries, regulator triggers.
- Certification valid 12 months, renewed via simulation.
- New joiners shadow two shifts and do not hold primary in their first 90 days.
- Appoint a champion in each of the 28 teams.
19. Simulations and game days (depends on: 14, 15, 18)
Rehearse before the next real SEV1. Use a standing calendar, not a one-off exercise.
Do not inject uncontrolled changes into the production ledger.
- Monthly 60-minute tabletop per engineering group, using a real incident from the 31.
- Quarterly full-scale game day: regional failover, ledger replica promotion, dependency failure. Whole role structure, timed.
- Twice-yearly unannounced paging drill, including nights, to measure real acknowledgement times.
- One security incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Also exercise status-page failure and loss of the primary chat or pager.
- Use replicas, staging, or tightly governed tests for ledger scenarios.
- Every exercise produces actions in the same tracker as real incidents.
- Complete two cross-company exercises before the SOC 2 audit.
20. Critical-path pilot (depends on: 9, 11, 12, 18, 19)
Prove the process on the highest-risk surface with willing teams before asking 28 teams to adopt it.
Publish a one-page result company-wide. That result is the adoption argument.
- Six-week Wave 0: ledger app, Postgres platform, payments orchestration, Kubernetes/platform, API gateway, Support intake, plus two of the 12 teams already on-call.
- Activate the full stack: severity scale, Duty IC, one pager, page budget, status page, mandatory postmortems, paid on-call.
- Parallel-run old paths for one week, then cut over. The program lead coaches every SEV2+ and does not secretly take command.
- Weekly retro. Expect 20–40 process defects and fix them in the standard before rollout.
- Exit gate: MTTD under 10 minutes on pilot services; IC assigned within 5 minutes in 95% of incidents; page volume down 50%; all postmortems on time; no unpaid pages; sentiment not worse.
21. Metrics, reviews, and error budgets (depends on: 12, 17, 20)
Instrument the process itself. Do not reward hiding incidents or suppressing pages.
Every metric has a target and a named owner. Dashboards are public inside the company.
- Response: MTTD, time to declare, time to IC, MTTA, MTTM, MTTR, incidents by severity, percent customer-first. Report median and p90, by tier, journey, and region.
- Quality: pages per person per week, actionability, budget breaches, postmortem on-time rate, action closure and age, missed acknowledgements.
- Business: credits, 99.95% per capability, error-budget burn, failed-payment volume, reconciliation breaks, repeat-incident rate.
- People: rotation size, frequency, after-hours pages, recovery days, sentiment, attrition among on-call staff.
- Cadence: weekly Incident Review Board; monthly Reliability Review; quarterly exec and board review; annual policy review.
- Error budgets on Tier 0/1 SLOs. Burn too fast and the team pauses features to pay down reliability.
- Reconcile dashboards monthly against a sample of incident records and customer cases so missing incidents cannot hide.
22. Wave rollout with readiness gates (depends on: 20, 21)
Roll out in four waves by criticality, every three weeks. Gates keep the standard credible. A missed gate is rescheduled, not waived.
Finish all 28 teams by week 24 so roughly three months of operating evidence remain before audit fieldwork.
- Wave 1: remaining Tier 0. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team kit: catalog complete, alerts migrated and inside budget, runbooks at the readiness bar, rotation of 6+ certified responders, one IC nominee, one drill passed, paid rotation in payroll.
- Named coach for three weeks. The director signs the gate.
- Freeze legacy paging per wave.
- Publish a live adoption scoreboard.
- Any production service that cannot provide sustainable ownership gets an executive-reviewed deadline, compensating control, and expiry date. No indefinite verbal exceptions.
23. Change management, incentives, and the deal (depends on: 1, 4, 9)
Run this from day one in parallel with design. Engineers will judge fairness. Executives will judge visible results.
Repeat the deal until it is muscle memory: paid on-call, paged only for what you own, trained commander, real sprint capacity for actions.
- Launch via CTO all-hands, per-team roadshows, a one-page laptop card, an internal wiki, and a Slack help channel with a 4-hour answer SLA.
- Put incident-response contribution into promotion criteria. Award best postmortem and biggest noise cut quarterly. Thank people publicly after every SEV1.
- Put adoption, alert hygiene, action closure, and on-call load fairness into every engineering manager's quarterly objectives.
- Write an exception path for engineers who cannot do nights because of caring responsibilities or health, covered by stipended volunteers.
- Prohibit retaliation for good-faith declaration or escalation.
- Pulse-survey at 60 and 120 days. If fairness or load is red, pause expansion until fixed.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
24. SOC 2 evidence, mock audit, and sustainability (depends on: 17, 21, 22)
Evidence is a by-product of doing the work, not a reconstruction the week before the auditor. Protect the process after SOC 2 is signed.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to CC7.2–CC7.4, CC2.2/CC2.3, CC5, and availability A1.2. Confirm interpretation early.
- Maintain versioned, signed policies for incident response, severity, on-call, communications, and postmortems, reviewed annually.
- Automate evidence: incident records, paging and ack logs, status history, postmortem library, action closure, training register, drill records, and reportability decisions including not reportable.
- Internal dry-run at month 6 on a 15-incident evidence walkthrough. Mock audit at month 7. Fix with weeks to spare.
- Assign permanent owners for policy, pager, status page, catalog, training, and metrics.
- Year-two roadmap: follow-the-sun cell, self-healing for the top recurring causes, error-budget release gates, and blast-radius reduction on the shared ledger.
- Quarterly board summary of severe incidents, credits, overdue P0s, and resilience investment so attention does not die after the audit.
Previous Proposal 5 (ID: 0ab2b1c8-46af-4001-ab14-0c074d30a026, Agent: deepseek-v4-pro_refine_5, LLM: deepseek/deepseek-v4-pro):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 6 months of full rollout.
- Customer-first detection falls from 40% to below 10% within 6 months and below 5% within 12.
- Median time to mitigate SEV0/SEV1 falls from 3h10 to under 60 minutes within 12 months; SEV2 under 2 hours.
- Incident Commander assigned and announced within 5 minutes for 95% of SEV0/SEV1 incidents; zero incidents with unclear command beyond 10 minutes.
- Status page updated within policy time for 95% of SEV0/SEV1/SEV2 (10/15/30 minutes).
- Monthly pages fall from 3,400 to under 500 with actionability above 75%.
- All six legacy alerting tools decommissioned by week 16.
- 100% of SEV0/SEV1 incidents have a blameless postmortem published within 10 business days.
- Postmortem action closure rises from 17% to 90% for P0/P1 actions on time.
- SLA credits fall from $1.3M to under $400K in first 12 months.
- Customer-impacting incidents decline to under 15 per year; repeat root causes under 10%.
- 100% of 180 services have a named owning team and criticality tier.
- All 28 teams onboarded by week 24; every Tier 0/1 team has 24x7 primary+secondary coverage with 6+ certified responders.
- At least 30 certified Incident Commanders and 20 certified Communications Leads active.
- Paid on-call policy is approved and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among on-call engineers.
- Quarterly game days and twice-yearly unannounced paging drills executed on schedule, each with tracked action items.
Steps (26):
1. Executive mandate, program governance, and interim incident command
Convert the CEO's outage complaint into a company-level improvement program with a named owner, budget, and deadlines. In the first week, establish an interim command process so no incident remains unowned while the permanent process is designed.
- Appoint Head of Reliability as program owner and CTO/CEO as executive sponsor.
- Form steering group with Engineering, SRE, Support, CS, Legal, Compliance, HR, Finance, Security.
- Approve non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Fund tooling, on-call compensation, training, and 3 dedicated program FTEs; anchor to $1.3M credits.
- Set timeline: design weeks 1-4, tooling and pilot weeks 5-10, rollout weeks 11-20, audit evidence collection from week 8, dry-run month 6.
- Stand up interim 24x7 duty officer and single declaration path within 48 hours, compensated retroactively.
2. Baseline data and alert estate analysis (depends on: 1)
Re-open the last 31 incidents and profile the current alert estate so every later design decision is evidence-based.
- Re-code each incident: detection source, timestamps, owner, severity, credits, root-cause family.
- Quantify customer-first detections and missing signals.
- Analyze two 'nobody in charge' incidents minute-by-minute.
- Inventory six alerting tools: volume, noise, owner, runbook coverage, top 50 noisy rules.
- Freeze baseline metrics: MTTD 22m, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed.
3. Stakeholder listening and resistance mapping (depends on: 1, 2)
Treat engineer pushback as design input. Interview all 28 teams plus Support, CS, and Sales to understand real objections and find early adopters.
- Test objection: unpaid work, night load, unfamiliar code, poor runbooks, or blame.
- Document current informal practices from 12 on-call teams.
- Recruit 10-15 credible engineers as co-design working group.
- Survey baseline sentiment on on-call and alerts.
- Publish promise: paid on-call, page only for owned services, trained commander, reserved reliability capacity.
4. Service ownership catalog and criticality tiering (depends on: 2, 3)
Build a machine-readable service catalog that assigns each of the 180 services a single owning team, escalation policy, and criticality tier.
- Define fields: owner team, manager, Slack channel, escalation policy, dependencies, dashboards, runbooks.
- Tier 0: ledger, money movement, auth, shared PostgreSQL; Tier 1 customer-facing; Tier 2 internal; Tier 3 non-critical.
- Map customer-visible capabilities and dependencies across regions.
- Identify orphan and shared services; force ownership decision within 30 days or schedule decommissioning.
- Publish coverage gaps by tier; Tier 0 gaps become executive escalations.
5. Define severity levels and declaration triggers (depends on: 4)
Adopt a five-level severity model with objective payment-specific triggers, so declaration is a lookup rather than debate. Anyone may declare; only the Incident Commander may downgrade.
- SEV0: unauthorized, lost, duplicated, or corrupted money movement; ledger integrity loss; confirmed data breach; both regions failed. Page all roles, exec, legal, and consider payment pause.
- SEV1: widespread payment failure, no workaround, severe SLA breach, or single-region total loss. Full role activation and public status page.
- SEV2: significant degradation, multiple customers, or workaround available but SLA risk. IC and SMEs paged; status page if customer-visible.
- SEV3: limited impact with workaround; team-led, business-hours response.
- SEV4: internal/no customer impact; ticket only.
- Auto-escalation: unresolved SEV3 >2h becomes SEV2; unresolved SEV2 >1h becomes SEV1; ledger or security always at least SEV2.
- Provide decision tree and 12 worked examples from actual incidents.
6. Define incident roles, decision authority, and handover rules (depends on: 5)
Codify five roles with written responsibilities and explicit authority, so there is never confusion about who is in charge.
- Incident Commander: owns severity, priorities, cross-team pulls, mitigation decisions; does not code.
- Communications Lead: owns status page, internal updates, account manager briefs, regulator coordination.
- Scribe: maintains timeline, decisions, and evidence for postmortem/audit.
- Subject-Matter Responders: diagnose and remediate only their owned services.
- Executive Liaison (SEV0/SEV1): handles exec communications and external stakeholders.
- Rule: IC identified within 5 minutes and announced in channel; handover announced and logged; roles may combine below SEV2, never at SEV0/SEV1.
7. Design 24x7 incident command and comms staffing model (depends on: 6, 4)
Create a central, trained incident command rotation instead of making each of 28 teams field its own commander.
- Recruit 24-30 certified ICs and 12-16 Comms Leads from a cross-team volunteer pool with manager approval.
- Weekly rotations, primary and secondary; 5-minute acknowledgement SLA with auto-escalation.
- Scribe pool as entry-level rotation.
- Night coverage: paid command rotation in New York timezone initially; evaluate follow-the-sun coverage later.
- Eligibility: certification required; commanders can leave with 30 days notice.
- Ensure distinct persons for IC, CL, and primary SME on SEV0/SEV1.
8. Design team on-call rotations and cross-team escalation policies (depends on: 4, 6)
Define tiered team on-call obligations so engineers are only paged for services they own, and cross-team pages go through the IC.
- Tier 0/1 services: 24x7 primary+secondary, at least 6 trained responders per rotation, one-week shifts.
- Tier 2/3: business-hours on-call; after-hours escalation list to engineering manager.
- Platform, infrastructure, database: 24x7 due shared ledger and Kubernetes.
- Routing: every page resolves via service catalog to owning team's escalation policy; cross-team pages only by IC.
- Guardrails: max one primary week in four, no consecutive weeks, no on-call in week after leading SEV0/1, protected recovery time after night work.
9. Design paid on-call compensation and fatigue safeguards (depends on: 8, 3)
Make on-call paid and legally compliant before any new rotation starts, and convert unpaid pager culture into a fair employment condition.
- Weekly stipends: primary 24x7 $800-1,200, secondary 30-50%, business-hours $300-500; holidays premium.
- Out-of-hours incident pay: $150 per night page plus hourly beyond one hour; time-off-in-lieu after overnight work.
- Additional command rotation stipend and SEV0/1 bonus for active responders.
- Verify FLSA/NY wage rules with HR, Legal, Finance; document exemption treatment.
- Base annual budget against current $1.3M credits.
- Track load and trigger staffing review if >2 after-hours pages per person per week sustained.
10. Enforce alert quality standards and noise budget (depends on: 2, 4, 8)
Replace 3,400 monthly alerts at 85% noise with a contractual paging standard that makes on-call sustainable.
- Every paging alert must have owning team, customer impact statement, runbook link, severity mapping, and tested threshold.
- Page only on customer-visible symptoms or SLO burn; cause-based alerts become tickets or dashboards.
- Set page budget: max 2 out-of-hours pages per person per week; breach triggers mandatory alert-tuning sprint.
- Auto-quarantine alerts with >5 firings/month without action or >70% no-action acknowledgements.
- Target <500 actionable pages/month and >75% actionability within 6 months.
- Weekly per-team alert review, monthly cross-team review.
11. Build detection uplift: synthetics, SLOs, and support intake (depends on: 4, 10)
Shift detection from host metrics to customer outcomes so the company stops hearing about outages from clients first.
- Define SLOs per Tier 0/1 capability: payment initiation, auth, settlement timeliness, API availability, ledger consistency.
- Deploy external synthetic transactions from both regions every 60 seconds, covering full payment flow and ledger write.
- Add ledger assurance checks: replication lag, double-entry balance, settlement window countdown.
- Top-100 customer anomaly detection to catch single-tenant outages.
- Auto-create triage incident from support tickets or account manager keywords within 5 minutes.
- Track customer-detected-first as a defect and require a postmortem action.
12. Consolidate alerting and incident tooling (depends on: 5, 8, 10, 11)
Collapse six alerting tools into one integrated paging and incident management platform to create a single system of record for people and audit.
- Select paging/on-call platform and incident management layer (e.g., PagerDuty + incident.io/FireHydrant).
- Implement one-command Slack declaration that auto-creates channel, bridge, pages roles, sets severity, starts timeline.
- Migrate all monitoring sources into the one tool; decommission legacy paging only after two weeks verified.
- Integrate service catalog, status page, Jira action tracking, Salesforce/CS customer lists, and conference bridge.
- Ensure out-of-band paging and offline fallback if a region or chat tool is down.
- Automate evidence capture for SOC2: timestamps, role assignments, severity changes, comms sent.
13. Define acknowledgement and escalation paths (depends on: 12, 6, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts at first signal.
- For SEV0/SEV1: page primary on-call; after 5 minutes unacked page secondary; at 10 page manager and duty IC; at 15 page executive officer.
- For SEV2: primary ack within 10 minutes, IC assigned within 15; escalate on miss.
- For SEV3: ack within 30 minutes or create work item.
- Automatic duty IC page for ledger/security, cross-team, or unresolved ownership.
- If impact unknown after 15 minutes, raise severity.
- Cross-team responders summoned by IC have 10-minute acknowledgement obligation.
- Route alerts with no owner to command rotation, then treat missing ownership as control defect.
- Human acknowledgement required; delivery confirmation is not sufficient.
14. Standardize internal communications (depends on: 6, 5)
Separate the working war room from executive and stakeholder updates, with fixed cadence and pre-approved templates.
- Auto-create incident channel and one read-only broadcast channel for execs, support, sales.
- SEV0: internal update every 15 minutes; SEV1 every 30; SEV2 every 60; SEV3 on state change.
- Template: impact, what we know, what we are doing, ETA/next update, current IC and CL.
- Executives questions through Executive Liaison only; IC is not interrupted.
- Support/CS receives affected-customer list and holding statement within 15 minutes for SEV0/SEV1 and 30 for SEV2.
- Handover protocol for incidents lasting >4 hours: formal IC handover and fatigue check.
15. Standardize customer and status page communications (depends on: 5, 14)
Replace ad-hoc status page updates with a timed, role-owned, template-driven process, including account manager outreach.
- Status page timing: SEV0 initial post within 10 minutes, SEV1 within 15, SEV2 within 30; updates every 15/30/60 minutes until resolved.
- Resolution notice within 30 minutes of mitigation; customer-facing summary in 5 business days for SEV0/SEV1.
- Pre-approve 12-15 templates with Legal and Comms.
- Tiered outreach: top 100 accounts get direct call/email from AM within 30 minutes of SEV0/SEV1; long tail via subscription.
- Use factual language: state impact and next update; never speculate cause or blame vendor.
- Comms Lead is sole author for customer language.
16. Regulatory, legal, and account manager notification playbook (depends on: 5, 15)
Build a notification decision tree and contact matrix so legal/regulatory obligations are assessed early and never forgotten.
- Map obligations: NYDFS Part 500 72-hour cybersecurity event notification, state breach laws, GLBA/FTC, PCI, sponsor bank/card network contractual windows, FinCEN/OFAC if relevant.
- Add regulatory assessment checkpoint for every SEV0 and security SEV1 within 2 hours, even if not reportable.
- Maintain 24x7 contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel with named backups.
- Pre-draft notification templates and test quarterly.
- Encode customer-specific notification SLAs from enterprise contracts into customer tiering.
- Account managers receive legal-approved script and affected-customer list.
17. SLA credit workflow and financial impact measurement (depends on: 5, 15)
Link every incident to money automatically so severity, credits, and prioritization stay consistent and Finance is never surprised.
- Define availability measurement per contract and capability with Legal/Finance.
- Auto-compute affected minutes per customer from incident record and telemetry; generate credit proposal within 5 business days.
- Decide proactive credits for top tier vs claims-based for others; document approval chain.
- Track credits by root-cause family and incident; feed quarterly reliability investment decisions.
- Target reducing annual credits from $1.3M to under $400K in first year.
18. Standardize blameless postmortems (depends on: 5, 6)
Make postmortems mandatory with fixed deadlines, a single format, and a blameless review forum, replacing 'various formats'.
- Mandatory: all SEV0/SEV1; SEV2 with customer impact, credits, >2h, repeat cause, or detected by customer; near-miss involving ledger.
- Draft within 3 business days, peer review within 5, publish within 10.
- One template: timeline, customer/financial impact, detection gap, response gap, contributing factors, what went well.
- Blameless rules: context-based, no individual blame, never used in performance reviews.
- Weekly Incident Review Board reviews all postmortems and challenges quality.
- Publish searchable postmortem library and quarterly recurring causes report.
19. Track postmortem actions with owner and due date (depends on: 18, 12)
Fix the 11/64 completion rate by giving every action the same status as customer commitments, with capacity and escalation.
- Every action gets named owner, priority, due date, and Jira ticket auto-created from postmortem.
- P0 actions prevent SEV0 recurrence, due 30 days; P1 60; P2 90.
- Reserve 15-20% team sprint capacity for incident actions.
- Escalation ladder: manager at 7 days overdue, director at 14, CTO dashboard at 30; P0 overdue blocks features.
- Monthly reporting of closure rate in engineering leadership.
- Target 90% closure of P0/P1 on time in two quarters.
20. Create runbooks and service readiness bar (depends on: 4, 8, 11)
Ensure every service is prepared for 3 a.m. response before it is allowed to page anyone.
- Readiness checklist for Tier 0/1: architecture diagram, dependencies, dashboards, rollback, feature flags, escalation contacts, data-loss statement.
- Write major-incident playbooks: shared PostgreSQL failure, cross-region failover, Kubernetes control plane loss, partner outage, settlement breach, security compromise.
- Prioritize ledger playbooks: documented failover, read-only mode, reconciliation, signed RPO/RTO.
- Runbooks must be tested twice yearly; stale runbooks marked in catalog.
- No paging alerts without readiness sign-off; gap reported to director.
21. Train, certify, and simulate incident response (depends on: 6, 13, 14, 15, 18, 20)
Build a certification path so 24x7 roles are staffed by people who have practiced, and validate the process through drills.
- Scribe course (2 hrs); Responder (half-day); Communications Lead (1 day); Incident Commander (2 days + shadowing/tabletop).
- Certify ICs and CLs; recertify annually.
- New responders shadow two shifts before primary; no primary in first 90 days.
- Monthly tabletop using real past incidents; quarterly game day including regional failover or ledger scenario.
- Twice-yearly unannounced paging drill; one regulator/legal exercise annually.
- Track drill metrics and action items.
22. Pilot on critical services and iterate (depends on: 12, 13, 15, 19, 21)
Prove the process on the highest-risk surface before full rollout. Run a six-week pilot with tight measurement and public results.
- Select 5-6 teams: payments core, ledger/database, platform/Kubernetes, API gateway, plus two existing on-call teams.
- Activate full stack: severity, roles, command rotation, paid on-call, single tooling, alert budget, status page, postmortems.
- Weekly pilot retro; fix defects within 48 hours.
- Exit criteria: MTTD <10 min, IC assigned within 5 min in 95% incidents, pages down 50%, postmortems on time, positive sentiment.
- Publish one-page pilot result as adoption argument.
23. Wave rollout to all teams and decommission legacy paths (depends on: 22, 20)
Roll out to all 28 teams in four waves by criticality, with explicit readiness gates and legacy tool shutdown.
- Wave 1 remaining Tier 0; Wave 2 Tier 1; Wave 3 Tier 2; Wave 4 Tier 3/internal.
- Per-team onboarding: catalog complete, alerts pruned, runbooks ready, rotation staffed with 6+ certified responders, one IC candidate, one drill passed.
- Coach per wave 3 weeks.
- Gate signed by director; failures rescheduled, not waived.
- After onboarding, freeze legacy alerting paths; no fallback to old tools.
- Publish adoption scoreboard.
24. Establish metrics, dashboards, and review cadence (depends on: 12, 19, 22)
Make process performance visible with a small set of metrics and fixed review meetings, so the program is owned by data.
- Response: MTTD, time to IC, MTTA, MTTM, MTTR, % customer-detected-first.
- Quality: page volume per person, alert actionability, postmortem on-time, action closure rate.
- Business: SLA credits, availability vs 99.95, repeat incidents.
- People: on-call load, page per engineer, sentiment, attrition.
- Cadence: weekly Incident Review Board, monthly Reliability Review, quarterly Executive/Board review.
- Dashboards self-serve with targets and named owners.
25. SOC 2 readiness and internal dry-run audit (depends on: 19, 23, 24)
Design evidence as a by-product and test it with an internal walkthrough before the external auditor arrives.
- Map process to SOC2 CC7.3/CC7.4, CC7.2, CC2.2/2.3, CC5, availability criteria.
- Publish versioned policy documents: Incident Response, Severity, On-Call, Communication, Postmortem.
- Automate evidence: incident records with timestamps, role assignments, paging logs, status history, postmortems, action board, training register, drill records.
- Ensure process operates at least 3 months before fieldwork.
- Run internal dry-run month 6; sample 15 incidents; fix gaps with 8 weeks to spare.
- Keep remediation log for process deviations.
26. Continuous improvement, culture, and sustainability (depends on: 23, 25)
Prevent the process from decaying after the audit by embedding review, feedback, and roadmap ownership.
- Quarterly process retrospective with IC pool and responders.
- Re-baseline metrics every six months; raise targets.
- Year-two roadmap: follow-the-sun coverage, self-healing top 3 causes, error budgets gating releases, blast-radius reduction for shared ledger.
- Annual policy review and certification renewal.
- Quarterly on-call sentiment survey with published actions.
- Board quarterly report on availability, credits, and incident trends.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Establish the mandate, owner, funding, and schedule", "description": "Launch incident management as a **company operating program within 48 hours**. Give one leader authority to set standards across all 28 teams.\n\n- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.\n- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.\n- Form a decision group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.\n- Approve the non-negotiables: one severity scale, one human-paging platform, one incident record, paid on-call, mandatory postmortems, and tracked corrective actions.\n- Reserve 10%–15% of engineering capacity for detection, runbooks, and incident actions.\n- Fund tooling, compensation, training, exercises, and reliability work. Use the $1.3M in credits as the minimum financial comparison.\n- Set milestones: interim process by day 7, policy and critical-path pilot by week 6, Tier 0/1 coverage by week 10, company rollout by week 18, internal audit in month 6, and mock audit in month 7.", "dependencies": []}, {"step_id": "S2", "title": "Install a seven-day incident-response floor", "description": "Do not wait for policy design or tooling migration. Put a minimum process into operation immediately and begin retaining evidence.\n\n- Publish a one-page provisional severity guide, declaration procedure, role card, and communication schedule.\n- Create one monitored declaration path through chat, telephone, and the current paging environment.\n- Use one incident channel, bridge, timeline document, and naming convention for every suspected major incident.\n- Staff temporary primary and backup Duty Incident Commanders from experienced on-call engineers and engineering leaders.\n- Require a named commander within 10 minutes. The duty engineering director assumes command if nobody else does.\n- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.\n- Compensate interim duty retroactively under the permanent policy.\n- Hold a 15-minute daily operational review until the permanent process is live.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Create the factual, legal, and audit baseline", "description": "Build one defensible baseline for process design, executive decisions, and SOC 2 testing. Confirm the auditor's expected Type II observation period immediately.\n\n- Reconstruct all 31 incidents from first customer impact through resolution, communications, credits, postmortem, and corrective action.\n- Analyze the two incidents with unclear command minute by minute.\n- Identify the missing detection signal for every customer-first incident.\n- Inventory the six alert sources, 3,400 monthly alert events, noisy rules, duplicates, missing owners, and missing runbooks.\n- Record the current rotations, unpaid work, after-hours load, and teams without coverage.\n- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M credits, 85% noise, and 11 of 64 actions closed.\n- Map incident-response controls to applicable SOC 2 criteria with Compliance and the auditor.\n- Establish retention, confidentiality, legal-hold, and access requirements for incident evidence.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Co-design the fairness contract with engineers", "description": "Treat pager resistance as a valid design constraint. Make the employment and ownership bargain explicit before expanding on-call.\n\n- Interview representatives from all 28 teams and the Support, Customer Success, Security, and Operations groups.\n- Distinguish objections involving unpaid work, unfamiliar code, bad alerts, weak runbooks, sleep disruption, or blame.\n- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.\n- Publish the core promise: responders are paid, paged only for services they own or have formally accepted, supported by trained command, and given capacity to remove recurring defects.\n- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.\n- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.\n- Measure baseline trust, fairness, fatigue, and psychological safety. Repeat the survey at days 60 and 120, then quarterly.", "dependencies": ["S1"]}, {"step_id": "S5", "title": "Build the service catalog and criticality model", "description": "Make a machine-readable catalog the source of truth for routing, escalation, customer impact, and audit evidence. Every production service must have one accountable owner.\n\n- Record the owning team, manager, product capability, escalation policy, communication channel, dashboard, runbook, dependencies, regions, and data stores for all 180 services.\n- Map payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, and reporting to their dependencies.\n- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.\n- Classify Tier 1 as customer-facing services whose failure can create material contractual impact.\n- Classify Tier 2 as internal or deferrable services, and Tier 3 as non-critical systems.\n- Record SLO, RTO, RPO, regional mode, recovery method, and contractual obligations for Tier 0 and Tier 1.\n- Give orphan services an owner or approved decommission date within 30 days.\n- Maintain separate but coordinated ownership for the ledger application and PostgreSQL platform.", "dependencies": ["S3"]}, {"step_id": "S6", "title": "Adopt one severity scale and incident lifecycle", "description": "Use four impact-based levels. Classify on actual or credible customer, financial, security, regulatory, and contractual harm rather than organizational seniority.\n\n- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; both regions impaired; or a core payment journey broadly unavailable. Page every role immediately, open a bridge, notify executives, publish customer status within 15 minutes when applicable, begin legal assessment, and require a postmortem.\n- **SEV2 — major:** material payment degradation; settlement deadline at risk; a critical customer or material cohort unavailable; regional impairment with reduced resilience; or likely SLA breach. Page command and technical roles, publish customer status within 30 minutes when customer-visible, and require a postmortem.\n- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads; page only when immediate action is necessary.\n- **SEV4 — operational event:** no current customer impact and no urgent risk. Create a ticket and handle during normal operations.\n- Treat percentages as supporting guardrails, never reasons to under-classify integrity, settlement, security, or contractual risk.\n- Anyone may declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.\n- Automatically use at least SEV2 posture for suspected ledger-integrity events, cross-team incidents with unknown ownership, and materially unknown impact lasting 15 minutes.\n- Escalate a customer-visible SEV3 that remains unmitigated for two hours.\n- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed. Resolution requires stability, backlog recovery, and necessary reconciliation.", "dependencies": ["S3", "S5"]}, {"step_id": "S7", "title": "Define roles, authority, and handoffs", "description": "Separate command, communications, recordkeeping, and technical repair. One named person must hold command throughout every SEV1 and SEV2.\n\n- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not serve as the primary technical operator.\n- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language.\n- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.\n- **Subject-Matter Responders:** diagnose and mitigate only services for which they have ownership, access, training, or formally accepted support responsibility.\n- **Executive Duty Officer:** removes organizational barriers and makes exceptional business decisions without displacing the commander.\n- Security, Legal, Compliance, Vendor Management, and Finance join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.\n- Require distinct commander, communications, scribe, and primary technical lead for SEV1. Communications and scribe may combine temporarily for bounded SEV2 incidents.\n- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.\n- Preserve dual control, privileged-access restrictions, and reconciliation requirements for all ledger operations.\n- Announce every role assignment and handoff verbally and in writing, with the exact transfer time.", "dependencies": ["S6"]}, {"step_id": "S8", "title": "Create sustainable 24x7 coverage across 28 teams", "description": "Use central command coverage and risk-based technical coverage rather than creating 28 fragile night rotations. All services receive a response path, but only critical domains maintain direct overnight technical rotations.\n\n- Create a 24x7 Incident Command corps of 24–30 certified people, with primary and secondary coverage at all times.\n- Create 18–24 trained Communications Leads and a similarly sized scribe pool using Support, Customer Operations, Engineering Operations, and qualified managers.\n- Maintain a surge roster for simultaneous incidents and a 24x7 Executive Duty Officer schedule.\n- Group Tier 0 and Tier 1 ownership into approximately 8–12 coherent domains only where responders have training, access, runbooks, and explicit acceptance.\n- Staff each critical domain with primary and secondary responders and at least six qualified people. Target no more than one primary week in six.\n- Give Tier 2 and Tier 3 services business-hours coverage plus a maintained manager and director escalation path.\n- When lower-tier impact becomes SEV1 or SEV2, central command activates the manager escalation path and obtains the necessary owner.\n- Route unknown-owner pages to Duty Command and platform triage temporarily. Log each one as a catalog control failure.\n- Do not merge small teams merely for schedule convenience. Provide training, reassign services, add staffing, or decommission unsupported services.", "dependencies": ["S4", "S5", "S7"]}, {"step_id": "S9", "title": "Approve compensation and fatigue protections", "description": "End unpaid on-call before expanding mandatory night coverage. HR, Finance, Payroll, and employment counsel should approve the policy within 14 days.\n\n- Use market-validated weekly bands, initially budgeting approximately $800–$1,200 for Tier 0/1 primary duty and $300–$500 for secondary duty.\n- Budget approximately $900–$1,300 for Duty Incident Commander weeks and $400–$800 for Communications Lead or scribe duty, adjusted for actual burden.\n- Pay holiday premiums and compensate active after-hours work according to exempt or non-exempt status and applicable federal and New York rules.\n- Provide a protected paid recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.\n- Reduce normal sprint commitments by approximately 15% during a primary on-call week.\n- Prohibit simultaneous primary assignments, consecutive primary weeks, on-call during leave, and invisible schedule swaps.\n- Allow responders to declare themselves temporarily unfit after disruptive night work without career penalty.\n- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.\n- Budget roughly $0.9M–$1.2M annually, then refine using actual rotation count, employment classification, and activation data.", "dependencies": ["S4", "S8"]}, {"step_id": "S10", "title": "Establish one paging and incident system of record", "description": "Monitoring tools may remain specialized, but all human pages must enter one controlled platform. This removes conflicting schedules and creates one evidence trail.\n\n- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page.\n- Ingest events from the six existing monitoring tools before disabling their direct paging paths.\n- Route pages through service-catalog ownership and deduplicate related events.\n- Provide a single declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.\n- Capture declarations, acknowledgements, escalations, role assignments, severity changes, decisions, communications, mitigation, resolution, and postmortem linkage automatically.\n- Integrate ticketing, customer-success data, status-page publishing, and conference facilities.\n- Require role-based access, MFA, access reviews, and immutable or tamper-evident history.\n- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, or an AWS region.\n- Retire each legacy human-paging path only after ownership review, end-to-end tests, and two weeks of verified operation.", "dependencies": ["S3", "S5", "S8"]}, {"step_id": "S11", "title": "Enforce an alert-quality contract and page budget", "description": "Treat every page as a production interface with an owner and an expected action. Noise reduction must not create detection gaps.\n\n- Require every paging rule to identify the service, owning team, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.\n- Page only when prompt human judgment or intervention can materially reduce customer, financial, security, or contractual risk.\n- Route informational, capacity, and non-urgent infrastructure conditions to dashboards or ticket queues.\n- Prefer payment-outcome, settlement-risk, queue-age, and error-budget-burn alerts over raw CPU, memory, pod, or log thresholds.\n- Run new paging rules in shadow mode for at least seven days unless an emergency exception is approved.\n- Review rules with repeated no-action acknowledgements, low actionability, or excessive firing within two business days.\n- Set a load budget of no more than two after-hours pages per responder per week, measured over four weeks.\n- Require an owner, compensating detection, and recorded approval before suppressing or deleting a rule.\n- Burn down the top 50 noisy rules first. Review missed detections and noise together so teams cannot improve metrics by becoming blind.", "dependencies": ["S3", "S5", "S10"]}, {"step_id": "S12", "title": "Detect payment and ledger failures before customers", "description": "Move detection from host health to customer journeys and financial outcomes. Use internal SLOs with enough headroom to protect the contractual 99.95% SLA.\n\n- Define SLIs and SLOs for payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, API access, webhooks, and reporting freshness.\n- Run external synthetic transactions through critical payment journeys at least every minute and from paths independent of the production platform.\n- Validate each AWS region and expose dependencies that undermine nominal regional redundancy.\n- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.\n- Alert on queue age and delayed value relative to settlement deadlines, not only queue depth.\n- Add tenant and cohort anomaly detection for high-value customers and major payment methods.\n- Convert high-priority Support, account-manager, processor, sponsor-bank, and network reports into incident candidates within five minutes.\n- Record the detection source for every incident. Treat customer-first detection as a mandatory missed-detection review.", "dependencies": ["S5", "S11"]}, {"step_id": "S13", "title": "Codify detection, escalation, and live execution", "description": "Create one time-bound path from the first credible signal to named command and mitigation. Notification delivery does not count as human acknowledgement.\n\n- Page the owning critical-domain primary and Duty Incident Commander immediately for suspected SEV1 or SEV2.\n- Escalate an unacknowledged technical page to secondary at five minutes, manager at 10, and director at 15.\n- Escalate an unclaimed command page to backup commander at five minutes. The Executive Duty Officer assumes command at 10 minutes until a certified handoff occurs.\n- Require Support and account managers to use the same declaration path as automated monitoring and engineers.\n- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.\n- Begin with a standard statement of severity, known impact, assigned roles, current objective, workstreams, and next update time.\n- Freeze unrelated changes during SEV1 and normally during SEV2. Record any exception.\n- Prefer reversible mitigation such as rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, or partner rerouting.\n- Separate mitigation and diagnosis workstreams when staffing permits.\n- Require formal command handoff for long incidents, shift changes, or fatigue. Do not leave an incident unowned during transfer.\n- Reconcile the ledger and safely drain backlogs before resolving payment or ledger incidents.", "dependencies": ["S7", "S10"]}, {"step_id": "S14", "title": "Standardize internal incident communications", "description": "Give responders one working room and stakeholders one controlled information source. Executives must not interrupt the technical command path.\n\n- Create one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.\n- Issue an initial internal brief within 10 minutes for SEV1 and 15 minutes for SEV2.\n- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.\n- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.\n- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within 15 minutes for SEV1 and 30 minutes for SEV2.\n- Route executive questions through the Executive Duty Officer or Communications Lead.\n- Record material decisions and outbound messages in the incident timeline.", "dependencies": ["S7", "S10", "S13"]}, {"step_id": "S15", "title": "Standardize status-page and customer communications", "description": "Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate impact before root cause is known.\n\n- Publish customer status within 15 minutes of declaring a customer-visible SEV1 and within 30 minutes for customer-visible SEV2.\n- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.\n- Publish a monitoring notice promptly after mitigation and a resolution notice within 30 minutes of verified recovery.\n- Map status-page components to customer capabilities rather than internal service names.\n- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, data-integrity investigation, security restriction, monitoring, and resolution.\n- State observed impact, affected capabilities, known workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.\n- Give affected strategic accounts direct account-manager outreach within 30 minutes for SEV1 and 60 minutes for SEV2.\n- Require account managers to use the approved briefing and prohibit independent technical explanations.\n- Provide a customer-facing incident summary within five business days for SEV1 and qualifying SEV2 events.\n- Record any legally necessary delay or restriction of public detail, its approver, and the alternative communication plan.", "dependencies": ["S6", "S10", "S14"]}, {"step_id": "S16", "title": "Operationalize legal, regulatory, contractual, and credit decisions", "description": "Some payments incidents start external notification clocks. Make assessment mandatory without assuming that every operational incident is reportable.\n\n- Build a counsel-validated matrix covering applicable NYDFS rules, state breach laws, GLBA or FTC obligations, PCI requirements, money-transmitter obligations, sponsor-bank and network contracts, cyber insurance, and customer contracts.\n- Engage Legal and Compliance immediately for every SEV1, security event, and suspected ledger-integrity event.\n- Record an initial reportability assessment within one hour for SEV1 and within two hours for other potentially reportable events.\n- Record non-reportable decisions as evidence, including facts considered, approver, timestamp, and reassessment trigger.\n- Maintain tested 24x7 contacts for counsel, regulators, sponsor banks, networks, insurers, and critical vendors.\n- Encode customer-specific notice deadlines and channels in the customer record used by the Communications Lead.\n- Have Legal own regulatory text and submission. Keep technical command with the Incident Commander.\n- Have Finance calculate affected minutes, delayed value, likely credits, and contractual exposure within five business days.\n- Track credits by incident and recurring cause to support reliability investment decisions.", "dependencies": ["S3", "S6", "S15"]}, {"step_id": "S17", "title": "Make postmortems mandatory, consistent, and blameless", "description": "Use one learning standard with fixed deadlines. Keep postmortems separate from performance, misconduct, and disciplinary processes.\n\n- Require a postmortem for every SEV1 and SEV2.\n- Also require one for customer-first detection, impact over two hours, SLA credits, contractual breach, repeated contributing factors, control failures, and ledger-integrity near misses.\n- Produce the factual draft within three business days, hold the review within five, and publish the approved version within 10.\n- Make the Incident Commander responsible for the timeline and the owning engineering director accountable for completion.\n- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.\n- Require explicit analysis of why detection was not earlier and why mitigation took as long as it did.\n- Use a trained facilitator who was not the incident commander.\n- Describe decisions in the context and information available at the time. Do not name an individual as the root cause.\n- Publish broadly useful findings internally while restricting security, privacy, personnel, or privileged material appropriately.", "dependencies": ["S6", "S7", "S10"]}, {"step_id": "S18", "title": "Give corrective actions enforceable ownership", "description": "Treat incident actions as risk commitments, not suggestions. A ticket is not complete until the expected risk reduction is verified.\n\n- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.\n- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with milestones.\n- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.\n- Reserve 10%–15% of team capacity for approved reliability and incident work.\n- Escalate overdue high-risk actions to the manager after seven days, director after 14, and CTO after 30.\n- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.\n- Permit overdue actions to block related releases where the unrepaired condition could reproduce severe impact.\n- Verify effectiveness through tests, telemetry, exercises, or production evidence before closure.\n- Triage all 53 historical open actions within 30 days. Complete, re-plan, or formally accept each risk.", "dependencies": ["S10", "S17"]}, {"step_id": "S19", "title": "Set the service-readiness bar and critical playbooks", "description": "A critical service must be supportable at 3 a.m. before its team is placed on direct overnight coverage. Existing critical detection must remain active while readiness gaps are repaired.\n\n- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and data-integrity constraints.\n- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional failure, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.\n- Define stop-processing and read-only modes for credible financial-integrity risk.\n- Document controlled failover, split-brain prevention, replay protection, backlog recovery, and post-recovery reconciliation.\n- Exercise critical runbooks at least twice per year and after material changes.\n- Block new Tier 0/1 releases and new paging rules when readiness requirements are missing.\n- Handle existing gaps through a named owner, compensating control, executive-approved expiry date, and remediation plan rather than disabling detection.", "dependencies": ["S5", "S8", "S11", "S12"]}, {"step_id": "S20", "title": "Publish policy and certify every response role", "description": "Convert the operating design into concise, signed documents and practical training. Training and exercises occur during paid working time.\n\n- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, and exception standards.\n- Provide one-page severity, role, authority, escalation, and communication cards inside the incident tool.\n- Train all employees to recognize and declare incidents.\n- Train all engineers in severity, acknowledgement, evidence preservation, handoff, and financial-integrity precautions.\n- Certify responders only after they demonstrate access, dashboards, runbooks, rollback, and escalation competence.\n- Certify Incident Commanders through formal instruction, simulation, and at least two shadowed incidents or exercises.\n- Train Communications Leads in status writing, account segmentation, contractual clocks, and legal escalation.\n- Train scribes in timeline quality and fact-versus-hypothesis labeling.\n- Require two shadow shifts before independent primary duty and annual recertification.\n- Maintain the training and certification register as operational and audit evidence.", "dependencies": ["S6", "S7", "S8", "S9", "S11", "S13", "S14", "S15", "S16", "S17", "S18"]}, {"step_id": "S21", "title": "Pilot on the payment critical path", "description": "Run a four-to-six-week pilot across the highest-risk journey before expanding. Use real incidents and exercises to correct the model quickly.\n\n- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, API edge, authentication, settlement, and Support intake.\n- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking.\n- Run old paging paths in parallel for no more than one week, then make the new platform authoritative.\n- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.\n- Have the program team coach incidents without silently assuming command.\n- Correct critical process or tool defects within 48 hours and update the standard visibly.\n- Exit only after 95% timely command assignment, 95% communication compliance, no unpaid pages, complete required postmortems, and at least 50% lower noise.", "dependencies": ["S9", "S10", "S12", "S19", "S20"]}, {"step_id": "S22", "title": "Roll out by customer journey and risk", "description": "Expand in controlled waves while the interim process remains active company-wide. Complete rollout early enough to accumulate operating evidence before the audit.\n\n- Roll out remaining Tier 0 domains first, then Tier 1, Tier 2, and Tier 3.\n- Use two-to-three-week waves with a named program coach and director sign-off.\n- Gate each team on catalog ownership, appropriate coverage, compensation, trained responders, tested escalation, alert quality, runbooks, access, and a passed tabletop.\n- Require six or more responders only for direct 24x7 technical rotations. Apply business-hours coverage and manager escalation to lower tiers.\n- Reschedule failed gates or approve a time-limited executive exception with a compensating control.\n- Disable legacy human-paging paths after verified cutover for each wave.\n- Publish an internal adoption dashboard by team, service tier, and control gap.\n- Finish critical coverage by week 10 and all 28 teams by week 18.", "dependencies": ["S21"]}, {"step_id": "S23", "title": "Exercise command, communications, regional recovery, and fallbacks", "description": "Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.\n\n- Run the first company-wide command and communications tabletop within 30 days.\n- Run monthly domain tabletops using scenarios from the 31 historical incidents.\n- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous security and operational events.\n- Test loss of chat, identity, paging, conference, and status-page providers.\n- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors where appropriate.\n- Conduct at least one unannounced after-hours paging test before the audit.\n- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.\n- Measure time to acknowledgement, command, customer notice, mitigation decision, handoff, and recovery.\n- Create tracked actions for every material exercise finding.", "dependencies": ["S19", "S20"]}, {"step_id": "S24", "title": "Measure performance and operate fixed review forums", "description": "Use a balanced scorecard that exposes weak controls without rewarding hidden incidents or suppressed alerts. Report medians and 90th percentiles rather than averages alone.\n\n- Measure time from first impact to detection, declaration, acknowledgement, command assignment, mitigation, recovery, and resolution.\n- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.\n- Track pages, actionability, duplicates, after-hours load, missed detections, and page-budget breaches.\n- Track postmortem timeliness, action age, closure by due date, verified effectiveness, and recurring contributing factors.\n- Track availability by customer journey, error-budget burn, failed or delayed payment value, reconciliation breaks, and SLA credits.\n- Track rotation size, duty frequency, overnight activations, recovery days, schedule exceptions, sentiment, and attrition.\n- Hold a weekly Incident Review Board for incidents, postmortems, control failures, noisy alerts, and overdue actions.\n- Hold a monthly executive reliability review for trends, funding, contractual exposure, and accepted risks.\n- Hold a quarterly control and resilience review with Product, Security, Compliance, Risk, and Internal Audit.\n- Reconcile incident records monthly against support cases, status history, credits, and customer complaints to detect under-reporting.", "dependencies": ["S3", "S10", "S17", "S18"]}, {"step_id": "S25", "title": "Prove SOC 2 operating effectiveness before fieldwork", "description": "Generate audit evidence through normal operation rather than reconstructing it later. Test both control design and consistent execution.\n\n- Maintain approved and versioned policies, exceptions, catalog records, schedules, compensation activation, access reviews, training, incidents, communications, postmortems, actions, and exercises.\n- Sample incidents monthly from first signal through verified corrective action.\n- Include customer-reported events, downgraded incidents, non-reportable legal decisions, missed timelines, and exercises in the population.\n- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.\n- Run a formal mock audit in month 7 using the populations, evidence requests, and interviews expected from the external auditor.\n- Record deviations honestly with owners, remediation dates, and compensating controls. Never rewrite historical records.\n- Confirm that evidence retention covers the full auditor-defined observation period.\n- Brief commanders, responders, Support, and Compliance on the real process without scripting inaccurate answers.", "dependencies": ["S20", "S22", "S23", "S24"]}, {"step_id": "S26", "title": "Institutionalize improvement and reduce structural risk", "description": "Prevent the program from decaying after the audit. Use incident evidence to drive permanent ownership and architectural investment.\n\n- Assign permanent owners for policy, service catalog, paging, status page, training, metrics, evidence, and exercise scheduling.\n- Review severity thresholds, staffing, compensation, and communication timings annually and after material process failures.\n- Recertify command and communications personnel annually.\n- Review recurring failure families quarterly and require executive decisions where corrective work repeatedly loses priority.\n- Maintain quarterly responder surveys and publish actions addressing fairness, fatigue, and tool friction.\n- Report severe incidents, credits, overdue high-risk actions, and resilience investment to the board or risk committee quarterly.\n- Treat the shared ledger cluster as a strategic concentration risk. Fund failover assurance, blast-radius reduction, isolation of non-critical readers, and stronger regional independence.\n- Evaluate follow-the-sun command or technical coverage using one year of page-load and staffing data.\n- Build year-two plans for automated mitigation, deployment safety, graceful degradation, and error-budget release controls.", "dependencies": ["S22", "S24", "S25"]}], "estimated_complexity": "high", "success_metrics": "- By day 7, every suspected major incident uses one incident record, one coordination channel, and a named commander.\n- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes; at least 95% receive command within 5 minutes.\n- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.\n- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and technical coverage.\n- By week 18, all 180 services have an owner, tier, tested escalation path, and coverage appropriate to their risk.\n- No mandatory night rotation starts before compensation, training, access, runbooks, and staffing controls are active.\n- Every direct 24x7 technical rotation has at least six qualified responders or a documented, expiring executive exception.\n- No responder is routinely assigned primary duty more often than one week in six or simultaneously assigned to two primary rotations.\n- At least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes by month 3.\n- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.\n- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.\n- Median mitigation time falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.\n- At least 95% of qualifying SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.\n- At least 95% of incidents meet their required internal and customer update cadence by month 3.\n- Monthly human pages fall from 3,400 alert events to no more than 1,500 by day 90 and no more than 500 actionable pages by month 6.\n- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, with no decline in Tier 0/1 detection coverage.\n- Average after-hours load remains at or below two pages per responder per week; every sustained breach produces a remediation plan.\n- All six monitoring sources route human pages through the single paging platform by week 12; direct legacy paging paths are disabled by week 18.\n- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.\n- All 53 historical open actions are triaged within 30 days.\n- At least 90% of high-priority corrective actions are completed by their approved due date with effectiveness evidence by month 6.\n- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.\n- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.\n- Monthly contracted availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.\n- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.\n- Quarterly on-call surveys reach at least 75% favorable responses on fairness, ownership boundaries, compensation, and sustainability by month 6.\n- The month-7 mock audit finds no unowned high-risk control gap, and at least 95% of sampled incidents contain complete operating evidence."}Effectively re-issues P1's round-1 plan: 21 of its 32 steps match P1 and only 5 match its own previous version. It gains the machinery it lacked (Policy v1, noise campaign, risk register, 90-day inspect) plus a day-one compliance step from P2, but contributes nothing new and leaves two internal inconsistencies.
- New step 5 maps TSC criteria, evidence, retention, legal hold and starts monthly sampling in week 1 — better audit sequencing than P1's round-2 step 30.
- Adds machinery its round-1 lacked: Policy v1 with exception register (24), quota-driven noise burn-down (26), risk register (31), 90-day inspect-and-adapt (32), anti-gaming reconciliation (27), separate live-execution doctrine (15).
- Keeps P4's "missing ownership or runbooks is a release blocker for Tier 0/1" (step 6) and P2's check that nominal two-region services do not hide single-region dependencies (step 12).
- Cleaner dependency chain than round 1: baseline → catalog → severity → roles → coverage → compensation → tooling → escalation → pilot → waves.
- Dropped the ledger blast-radius reduction workstream as a defined item, yet step 31 still promises to escalate "if the blast-radius workstream slips" — a dangling reference.
- Folded the readiness bar into the tenth bullet of step 15 rather than keeping a runbook/readiness step, where it is easy to lose during rollout gates.
- Dropped its own round-1 metrics for monthly 99.95% availability and "<10% of engineers unwilling to join their owned-service rotation".
- Step 28 still claims completion "at least five months before the audit report date" against a weeks 12–22 rollout and a month-8 audit — arithmetically impossible.
- Keeps "postmortem on request" for SEV3, looser than P1/P4's customer-detected/2-hour/repeat trigger.
- Adds no idea the round did not already have; it is now strictly behind P1's round-2 version.
- Proposal 1 : P1's overall architecture: Layer A/B/C coverage, Policy v1 with exception register, noise burn-down campaign, risk register, 90-day inspect-and-adapt.
- Proposal 2 : Design compliance and evidence controls on day one, with retention, legal hold and monthly sampling.
- Proposal 2 : Validate that nominally two-region services do not depend on single-region databases, identities, queues or third parties.
- Proposal 4 : Missing ownership or runbooks is a release blocker for Tier 0/1 services.
- Proposal 2 : A SEV0 financial/security crisis tier and numeric declaration guardrails.
- Proposal 1 : A standalone runbooks/readiness/ledger blast-radius step with a parallel architecture workstream.
+ Establish executive mandate, program office, and funding+ Map compliance, evidence, and audit requirements from day one+ Build the service ownership catalog and criticality tiers+ Design the three-layer 24x7 coverage model+ Build payment-outcome detection and ledger assuranceExecutive mandate, single owner, funding and non-negotiablesService catalog, ownership and money-path tiering24x7 coverage model: central command corps, local expertiseDetection uplift on the money pathRunbooks, readiness bar and ledger blast-radius reduction
The plan produced
1. Establish executive mandate, program office, and funding
Turn the CEO email into a chartered company program with one accountable owner and authority over all 28 teams.
- Appoint the CTO as executive sponsor and a Head of Reliability as program owner with full-time authority.
- Create a permanent program office: one program lead, one platform engineer, one analyst.
- Form an eight-person steering group: Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours.
- Lock the non-negotiables: one severity scale, one paging platform, one postmortem format, paid on-call, mandatory action tracking, named service ownership.
- Approve budget anchored against the $1.3M in credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs.
- Reserve 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Publish a one-page charter stating incident response is a company operating process, not a per-team choice.
2. Install a seven-day interim command bridge (after 1) from P1 step 2
Do not leave the company unprotected while the permanent process is designed. Put a crude but real command structure in place within one week.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Direct Support to escalate credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical actions: complete, re-plan, or formally risk-accept, prioritizing ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Build the forensic baseline of incidents, alerts, and money lost (after 1) from P1 step 3
Rebuild the facts before designing anything. This becomes both the design input and the before picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific missing signal. This list drives the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Freeze baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the on-call fairness deal (after 1) from P4 step 4
Engineer pushback is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS, and Sales in two weeks. Separate the real objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey covering fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
5. Map compliance, evidence, and audit requirements from day one (after 1) from P2 step 4
Design evidence as a by-product of operations, not a reconstruction before the auditor arrives. Confirm the SOC 2 observation window immediately.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for communications; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure, training register, drill records, reportability decisions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve all policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (after 3) from P4 step 5
You cannot route a page correctly across 180 services until every service has a named owner. Build a machine-readable catalog as the single source of truth.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact: Tier 0 (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for each Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
7. Adopt the severity scale, declaration rules, and lifecycle (after 3, 6) from P1 step 6
Replace judgment calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter.
- SEV1 (crisis): money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of all roles and executive; bridge in 5 minutes; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 (critical): material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- SEV3 (major, contained): narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected, Declared, Triaged, Mitigated (customer impact ends), Monitoring, Resolved (backlog processed and ledger reconciled), Reviewed.
- Publish a decision tree with 12 worked examples from the real 31 incidents.
8. Define incident roles, authority, and handover discipline (after 7) from P1 step 7
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging.
- Incident Commander: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates for SEV1/SEV2.
- Subject-Matter Responders: engineers of the owning team; they mitigate, they do not run the room.
- Executive Duty Officer (SEV1): removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Design the three-layer 24x7 coverage model (after 6, 8) from P4 step 8
Do not create 28 night rotations. Centralise coordination in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- Layer A, Incident Command corps: approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of approximately 18 from Support, CS, and engineering management. A scribe pool used as the training entry point.
- Layer B, critical-path domain rotations: consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- Layer C, everyone else: business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing: a US-only paid night rotation now, a follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
10. Approve on-call compensation, labor compliance, and fatigue safeguards (after 4, 9) from P1 step 9
Unpaid on-call in New York is both a retention problem and a legal exposure. Pay must be live in payroll before any mandatory night rotation starts.
- Indicative scheme: approximately $1,000 per primary 24x7 week, $400 secondary, $250 for business-hours rotations, a separate $1,200 Duty Commander stipend, holiday premiums, approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- HR, Finance, and employment counsel publish amounts, eligibility, tax treatment, and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a 15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
11. Set the alert quality standard, page budget, and noise burn-down (after 3, 6) from P1 step 10
3,400 alerts at 85% noise is why detection takes 22 minutes. Make alert quality a condition of being allowed to page a human.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly.
- Target 3,400 to under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
12. Build payment-outcome detection and ledger assurance (after 6, 11) from P4 step 11
Stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer detected first becomes a named defect class with a mandatory tracked action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Consolidate to one paging platform, one incident record, one status page (after 7, 9, 11) from P1 step 12
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline, and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is unavailable. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
14. Codify the escalation ladder and five-minute command rule (after 8, 9, 13)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately, secondary at 5 minutes, domain manager and Duty Commander at 10, Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension, and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
15. Write the live incident execution doctrine and major-incident playbooks (after 8, 13, 14) from P4 step 14
Give responders one short operating procedure from the first minutes through closure. Priority is limiting customer and financial harm, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the single largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window, an explicit handback to the owning team, Support, and Customer Success, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
16. Standardize internal communications protocol (after 8, 13)
Standardise the internal picture so executives, Support, and Sales are informed without interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
17. Build customer communications, status page, and account-manager outreach (after 7, 16)
Replace whoever is around with a timed, owned, pre-approved process. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook, and RSS. All 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity, or blame.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
18. Create the regulatory, partner, and legal notification playbook (after 7, 17) from P1 step 17
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded even when the answer is not reportable, with evidence, decision-maker, and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
19. Operationalize SLA credit and financial impact workflow (after 7, 17) from P1 step 18
Tie incidents to money so severity, credits, and investment decisions stay honest, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one. Use the delta as the standing business case for on-call pay and reliability capacity.
20. Establish the blameless postmortem standard and Incident Review Board (after 7, 8) from P1 step 19
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly top-five recurring causes analysis.
21. Enforce action ownership, capacity reservation, and tracking (after 13, 20) from P1 step 20
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
22. Build the training and certification academy (after 8, 14, 16, 20) from P1 step 22
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days + shadowing): command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Comms Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum; never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Run the exercise programme: tabletops, game days, and unannounced drills (after 13, 15, 22) from P1 step 23
The process must meet a simulated SEV1 before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger. Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (after 7, 8, 9, 10, 11, 14, 16, 17, 18, 20, 21) from P1 step 24
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (after 10, 12, 13, 15, 22, 24) from P1 step 25
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment.
- Publish a one-page result to the whole company.
26. Run the alert noise burn-down campaign (after 11, 13, 25) from P1 step 28
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness. Review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 to 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
27. Establish metrics, dashboards, review cadence, and anti-gaming (after 13, 20, 25) from P1 step 26
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- Response: time to detect, declare, commander, acknowledgement, mitigate, resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
28. Wave rollout to all 28 teams with readiness gates (after 25, 27) from P1 step 27
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. Failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Publish a live adoption scoreboard.
29. Run change management, fairness, and pager culture programme (after 4, 10, 25) from P1 step 29
Run this from day one in parallel. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. Build SOC 2 evidence by design, internal testing, and mock audit (after 24, 27, 28) from P1 step 30
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any change as an exception.
31. Maintain the program risk register and contingencies (after 1)
Name the ways this programme fails and pre-commit the response. Review it monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (after 27, 28, 30) from P1 step 32
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
- A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 3efb9ff1-28a5-40aa-8f73-9853e91aa095, Agent: opus5_refine_1, LLM: anthropic/claude-opus-5):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (32):
1. Executive mandate, single owner, funding and non-negotiables
Convert the CEO email into a chartered program with one accountable owner and authority over all 28 teams. Incident response becomes a **company operating process**, not a per-team choice.
- Name an executive sponsor (CTO) and one accountable owner (Head of Reliability / Incident Management) with a small permanent office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group with Engineering, SRE/Platform, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours. It is not a 28-person committee.
- Fix the non-negotiables now: one severity scale, one paging tool, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.
- Set the clock deliberately earlier than the audit: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready week 28.
- Approve budget against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises.
2. Seven-day interim command bridge (depends on: 1)
Do not let design work leave the company unprotected for six weeks. Put a crude but real process in place within seven days and improve it later.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster from engineering managers and senior SREs, primary plus backup, 24x7. Pay it retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Tell Support to escalate credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items: complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This becomes both the design input and the "before" picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start / detect / declare / commander-assigned / mitigate / resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. This list drives the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome. Identify the top 50 rules producing most of the noise and every rule with no owner or runbook.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS and Sales in two weeks. Test what the objection actually is: unpaid work, lost sleep, unfamiliar code, missing runbooks, or fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey (fairness, trust in alerts, willingness, burnout) to re-measure at 60, 120 and 365 days.
5. Service catalog, ownership and money-path tiering (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that is the single source of truth for paging, impact and audit.
- One accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not by technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions and third parties, including sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Every orphan service gets an owner in 30 days or a decommission date approved by the sponsor. Tier 0 without an owner is an executive escalation.
6. Severity scale, declaration rules and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- **SEV3 (major, contained):** narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- **SEV4:** no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring; only the commander may downgrade, with the evidence recorded.
- Lifecycle: Detected → Declared → Triaged → **Mitigated** (customer impact ends) → Monitoring → **Resolved** (backlog processed and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Incident roles, authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, roles, cadence and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit spend. Retains command when a VP joins.
- **Communications Lead:** single voice for status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates for SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team; they mitigate, they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance and Vendor Management join on defined triggers.
- Rules: command claimed within 5 minutes, stated in channel ("I am IC"), distinct people for command, comms and technical lead at SEV1/SEV2, and every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation or privileged-access rules.
8. 24x7 coverage model: central command corps, local expertise (depends on: 5, 7)
Do not create 28 night rotations. Centralise coordination in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- **Layer A — Incident Command corps:** ~30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of ~18 from Support, CS and engineering management; a scribe pool used as the training entry point.
- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- **Layer C — everyone else:** business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing, with cost and timeline: a US-only paid night rotation now, a Lisbon or APAC follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
9. On-call compensation, labour compliance and fatigue safeguards (depends on: 8, 4)
Unpaid on-call in New York is both a retention problem and a legal exposure. Pay for it before asking anyone to sign up, and publish the numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 for business-hours rotations, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover rather than expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, tax treatment and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a ~15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
10. Alert quality standard, page budget and noise burn-down (depends on: 3, 5)
3,400 alerts at 85% noise is the reason detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human, and burn the backlog down deliberately rather than by mass silencing.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a **page budget** of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements; return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
- Target 3,400 → under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
11. Detection uplift on the money path (depends on: 10, 5)
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls; five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrating without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack or the identity provider is unavailable. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live incident execution doctrine (depends on: 7, 12, 13)
Give responders one short operating procedure for the first minutes through closure. Priority is limiting customer and financial harm, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; controlled backlog drain; **reconciliation completed before any payment or ledger incident is declared resolved**.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and Customer Success, and reopening if impact recurs.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook and RSS; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack; the long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity or blame; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded **even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review; Legal owns outbound regulatory text, the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
18. SLA credit and financial impact workflow (depends on: 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one, and use the delta as the standing business case for on-call pay and reliability capacity.
19. Blameless postmortem standard and Incident Review Board (depends on: 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first?** and **why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, capacity reservation and enforcement (depends on: 19, 12)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
21. Runbooks, readiness bar and ledger blast-radius reduction (depends on: 5, 8)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the single largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel workstream on blast-radius reduction — tenant or function partitioning, read replicas, and isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (depends on: 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days + shadowing):** command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence for their own domain before independent primary duty; two shadow shifts minimum, never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (depends on: 22, 12, 21)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 11, 12, 21, 22, 9)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect (from first impact), time to declare, time to commander, acknowledgement, mitigate, resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; team scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Wave rollout to all 28 teams with readiness gates (depends on: 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director; failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
28. Alert noise burn-down campaign (depends on: 10, 12, 25)
Run the noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
29. Change management, fairness and pager culture (depends on: 4, 9, 25)
Run this from day one in parallel. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (depends on: 24, 26, 27)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and contingencies (depends on: 1)
Name the ways this programme fails and pre-commit the response. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 26, 27, 30)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 2 (ID: 2e38c07e-bf01-4b70-ba95-0ca1ef39e2d8, Agent: gpt5.6-sol_refine_2, LLM: openai/gpt-5.6-sol):
Estimated Complexity: high
Success Metrics: - By day 7, every suspected SEV0–SEV2 uses one incident record, one coordination channel, and one named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named commander for more than 10 minutes; at least 95% are assigned within 5 minutes.
- By day 30, 100% of Tier 0 services have a named owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, every Tier 0 and Tier 1 customer journey has compensated 24×7 command and subject-matter coverage.
- By day 120, all 180 production services have an owner, response tier, tested escalation path, and appropriate coverage model.
- No mandatory night rotation starts before its compensation, training, access, and staffing controls are active.
- Every critical primary rotation has at least six qualified responders or a documented, expiring executive exception by day 120.
- No responder is routinely scheduled for primary duty more often than one week in six by day 120.
- At least 95% of critical pages are acknowledged within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by day 120 and under 60 minutes by month 8.
- At least 95% of SEV0 and SEV1 customer notices are issued within 15 minutes of declaration by month 3.
- At least 95% of customer-visible SEV2 notices are issued within 30 minutes by month 3.
- At least 95% of incidents meet their required update cadence by month 3.
- Monthly paging volume falls from 3,400 to no more than 1,500 by day 90 and no more than 700 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6 without loss of critical detection coverage.
- All new paging alerts satisfy the owner, impact, action, dashboard, runbook, deduplication, and escalation standard by day 60.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 currently open historical actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date by month 6, with effectiveness evidence.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Two end-to-end cross-company exercises, including regional failure and ledger recovery, are completed before the audit.
- Monthly availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- SLA credits decline by at least 50% on an annualized trailing basis within 12 months.
- Quarterly on-call surveys show improving fairness and sustainability, with at least 75% favorable responses by month 6.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (24):
1. Create the mandate, ownership, and funding
Launch incident management as a company operating program within 48 hours. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable program owner.
- Form a small steering group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Fund two to three implementation staff, paging and incident tooling, observability work, training, exercises, and on-call compensation.
- Reserve 10%–15% of engineering capacity for alert remediation, runbooks, and incident actions.
- Authorize incident commanders to freeze deployments, order rollback, disable features, shift traffic, and invoke continuity plans.
- Preserve financial controls. Incident commanders may coordinate ledger recovery but may not bypass dual approval, privileged-access controls, or reconciliation.
- Set the delivery target: critical controls operational within 60 days, enterprise rollout within 120 days, and a mock audit in month 7.
2. Install an interim process in seven days (depends on: 1)
Do not wait for new tools or the final policy. Put a minimum viable incident process into operation immediately and start collecting evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communications schedule.
- Establish one monitored declaration path through chat, telephone, and the existing paging tools.
- Create a standard incident channel, bridge, incident document, and naming convention.
- Staff temporary primary and backup incident commanders 24×7 from the existing on-call teams and engineering leadership.
- Compensate interim duties retroactively under the final compensation policy.
- Require a named incident commander within 10 minutes for every suspected major incident.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Hold a daily 15-minute control review until the permanent process is live.
3. Build the baseline, ownership catalog, and risk map (depends on: 1)
Establish the facts behind the current failures and assign every production component an owner. Use the resulting catalog as the source for routing, escalation, and audit evidence.
- Reconstruct all 31 customer-impacting incidents, including first impact, detection source, declaration, command assignment, mitigation, resolution, customer communications, and credits.
- Analyze the two incidents with no clear leader and every case detected first by customers.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, databases, queues, regional dependencies, payment processors, and banking partners.
- Assign one accountable team, engineering manager, product owner, primary escalation, secondary escalation, dashboard, and runbook to each service.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 based on financial integrity, customer impact, contractual exposure, and dependency centrality.
- Map critical customer journeys to their application, PostgreSQL, Kubernetes, regional, and third-party dependencies.
- Inventory the six alert sources and all 3,400 monthly alerts by owner, volume, actionability, and duplication.
- Interview representatives from all 28 teams and baseline on-call sentiment, fatigue, and objections.
- Give orphan services an owner or decommissioning decision within 30 days.
4. Design compliance and evidence controls from day one (depends on: 1)
Map the operating process to audit, legal, contractual, and record-retention requirements before finalizing it. Confirm the expected SOC 2 observation period with the auditor immediately.
- Map controls to applicable SOC 2 criteria for monitoring, incident identification, response, recovery, communications, corrective action, access, and availability.
- Define evidence required for declarations, pages, acknowledgements, role assignments, decisions, status updates, postmortems, actions, training, drills, and exceptions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve the Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, and Postmortem Standard.
- Record control exceptions with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing evidence before the audit.
5. Adopt one severity and lifecycle standard (depends on: 2, 3, 4)
Use four impact-based severity levels across operational, security, data, and third-party incidents. Start at the highest credible severity when facts are uncertain, then downgrade with recorded evidence.
- **SEV0 — financial or security crisis:** suspected ledger corruption, unauthorized or duplicated funds movement, material data compromise, both-region loss, or a decision to suspend payment processing. Page all roles immediately; engage executives, Security, Legal, Compliance, and Risk; assess regulatory duties within one hour; use distinct role holders; require a postmortem.
- **SEV1 — critical availability event:** a core payment journey is unavailable, payment failures exceed an initial 10% guardrail for five minutes, regional loss has impaired failover, a settlement deadline is at imminent risk, or rapid error-budget burn makes a material SLA breach likely. Assign all roles; notify internal stakeholders within 10 minutes; publish customer status within 15 minutes; update every 30 minutes; require a postmortem.
- **SEV2 — major bounded event:** approximately 1%–10% of payment attempts fail, a material customer subset or critical customer is down, degradation has a workaround, or contractual impact is likely. Assign an incident commander and responders; add communications and scribe roles for customer impact; publish status within 30 minutes; update every 60 minutes; require a postmortem for customer-visible events.
- **SEV3 — limited event:** localized impact, a safe workaround, and no financial-integrity, security, regulatory, or material contractual risk. The owning team leads; page only if immediate action is necessary; use a ticket otherwise.
- Treat the percentage thresholds as declaration guardrails, not reasons to under-classify integrity, settlement, security, or strategic-customer risk.
- Permit any employee to declare an incident. Only the incident commander may lower severity, with the rationale logged.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of active customer harm. Declare resolution only after stability, backlog recovery, and required ledger reconciliation.
6. Define roles, authority, and handoffs (depends on: 5)
Separate command, communication, recordkeeping, and technical repair. One named person must hold command at every moment of a major incident.
- **Incident Commander:** owns severity, priorities, role assignment, escalation, decision cadence, mitigation coordination, and closure. The commander does not act as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language. Legal or Compliance retains ownership of regulatory submissions.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, and status changes.
- **Subject-Matter Responders:** diagnose and mitigate services for which they have accepted ownership, access, training, and runbooks.
- **Executive Duty Officer:** removes organizational obstacles and approves exceptional business decisions without displacing the incident commander.
- Require separate commander, communications, scribe, and primary technical lead for SEV0 and SEV1.
- For customer-visible SEV2, keep the commander separate from the primary technical responder; communications and scribe may be combined if workload permits.
- Announce every role assignment and transfer in the incident channel. Require verbal and written handoff with current impact, decisions, risks, and next actions.
- Keep executives and account managers out of the technical command path; questions flow through the Executive Duty Officer or Communications Lead.
7. Create sustainable 24×7 coverage across the service estate (depends on: 3, 6)
Use central command coverage and service-domain responder coverage rather than creating 28 fragile night rotations. Engineers remain responsible only for services they own or have formally trained to support.
- Build a company incident-command pool of approximately 18–24 certified senior engineers and managers, with a primary and backup scheduled at all times.
- Build 12–16-person communications and scribe pools using Support, Customer Operations, Engineering Operations, and qualified engineering managers.
- Schedule an Engineering Director or equivalent as the 24×7 Executive Duty Officer.
- Group related service owners into approximately 8–12 coherent responder domains only where members have training, access, and explicit acceptance.
- Require Tier 0 and Tier 1 domains to provide 24×7 primary and secondary responders, normally with at least six trained people in each sustainable rotation.
- Give Tier 2 services business-hours coverage plus a maintained manager escalation path. Treat Tier 3 conditions as tickets unless their impact changes.
- Maintain distinct but coordinated coverage for the ledger application and the PostgreSQL platform.
- Route unknown-owner events to the duty commander and platform triage temporarily. Treat every such event as an ownership control defect.
- If a small team cannot staff a fair rotation, merge coverage only after training or provide headcount, service reassignment, or decommissioning.
8. Approve compensation and fatigue protections (depends on: 7)
End unpaid on-call before expanding mandatory coverage. Publish the policy through HR, Legal, Finance, and Payroll within 14 days.
- Pay fixed weekly stipends for primary and secondary service rotations.
- Pay separate stipends for duty commander, communications, and scribe assignments.
- Provide additional call-out compensation or equivalent paid recovery time for material after-hours work.
- Apply overtime and reporting rules correctly for non-exempt employees under federal and New York requirements.
- Pay higher rates for company holidays and provide a protected recovery day after qualifying overnight work, SEV0 events, or prolonged SEV1 response.
- Target no more than one primary week in six and prohibit simultaneous primary assignments.
- Avoid consecutive primary weeks and make all swaps visible in the paging system.
- Reduce sprint commitments for people carrying primary duty rather than expecting normal delivery capacity.
- Provide a documented accommodation path for health, disability, or caregiving constraints without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
9. Set the service readiness and runbook standard (depends on: 3, 5, 7)
A team cannot respond effectively at night without ownership, access, telemetry, and rehearsed recovery procedures. Apply a formal readiness gate to every Tier 0 and Tier 1 service.
- Require a current architecture diagram, dependency map, dashboards, SLOs, runbooks, rollback method, feature-control method, contacts, and tested escalation path.
- Record RTO, RPO, data-integrity requirements, regional mode, and customer-facing capabilities in the service catalog.
- Link every paging alert to the exact runbook step expected from the responder.
- Prevent new paging alerts for services that fail readiness review. Preserve existing critical detection through a documented exception until a safe replacement exists.
- Write priority playbooks for PostgreSQL failure, ledger-integrity investigation, regional failover, Kubernetes control-plane degradation, payment-processor failure, queue backlog, credential compromise, and payment suspension.
- For the ledger, document read-only or stop-processing modes, failover controls, replay and duplicate protections, backlog handling, and post-recovery reconciliation.
- Require runbook review after material incidents and at least twice per year through exercises.
10. Establish one paging and incident system of record (depends on: 4, 5, 6, 7)
Consolidate paging and incident coordination without requiring an unsafe big-bang replacement of every monitoring system. Monitoring sources may remain specialized, but all human pages must enter one controlled platform.
- Select one enterprise paging platform and one integrated incident record, with chat, telephone, SMS, conference, status-page, ticketing, and service-catalog integrations.
- Initially ingest events from all six current tools, then deduplicate, correlate, and route by service ownership.
- Automate incident channel and bridge creation, role paging, timeline capture, severity changes, communications reminders, and postmortem creation.
- Preserve immutable records of declarations, acknowledgements, role assignments, decisions, and messages.
- Use role-based access, multifactor authentication, break-glass controls, and periodic access reviews.
- Provide telephone and offline fallback procedures for loss of chat, identity, the paging vendor, or an AWS region.
- Test paging and fallback paths weekly.
- Retire a legacy paging route only after its signals have owners, quality review, successful end-to-end tests, and at least two weeks of verified operation in the new path.
11. Enforce alert quality and burn down noise safely (depends on: 3, 10)
Treat paging alerts as production products with owners and quality requirements. Do not reduce noise by silently disabling detection.
- Require every page to identify the service, owner, customer or SLO risk, urgency, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Define an actionable page as one requiring prompt human judgment or intervention that materially reduces customer, financial, security, or contractual risk.
- Route informational, capacity-planning, and non-urgent conditions to dashboards or ticket queues.
- Prefer symptom and error-budget burn alerts over raw CPU, memory, pod, or log-volume thresholds.
- Run new alerts in shadow mode for at least seven days unless an emergency exception is approved.
- Review alerts with less than 50% actionability, more than three firings in seven days, or repeated no-action acknowledgements within two business days.
- Require compensating detection and an owner before suppressing or removing an alert.
- Set a responder load target of no more than two after-hours pages per person per week. A breach creates a mandatory alert-remediation plan.
- Review alert actionability, duplication, missed detection, and page load monthly by domain.
- Prioritize the small number of rules producing most of the current 85% noise.
12. Detect payment and ledger failures before customers (depends on: 3, 11)
Shift detection from infrastructure symptoms to customer journeys and financial outcomes. Set internal objectives stricter than the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, webhooks, and reporting freshness.
- Run external synthetic transactions from outside the production boundary and through both regions at least every minute for critical paths.
- Monitor payment failure rates, latency, queue age, delayed value, settlement-window risk, regional asymmetry, and third-party response quality.
- Add continuous ledger controls for reconciliation breaks, unexpected balances, duplicate identifiers, replication lag, backup health, and failover readiness.
- Add tenant or cohort anomaly detection for high-value customers and common payment methods.
- Convert high-priority Support, account-manager, bank, and processor reports into incident candidates within five minutes.
- Review every customer-first incident as a missed-detection defect and create a corrective action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Codify escalation and live incident execution (depends on: 6, 10)
Create one time-bound path from signal to ownership and mitigation. Delivery of a notification does not count as acknowledgement.
- Page the owning Tier 0 or Tier 1 primary immediately; page the secondary after five unacknowledged minutes; page the domain manager at 10 minutes; escalate to the engineering director at 15 minutes.
- For SEV0–SEV2, page the duty commander immediately. Page the backup after five minutes and require the Engineering Director to assume command if no certified commander owns the event by 10 minutes.
- Automatically involve command for integrity concerns, security concerns, regional events, cross-team impact, critical customer journeys, or unresolved ownership.
- If impact remains materially unknown after 15 minutes, increase response posture rather than waiting for certainty.
- Open one incident channel, bridge, and system record. State severity, known impact, assigned roles, current objective, and next update time.
- Freeze unrelated changes during SEV0 and SEV1 unless the commander records an exception.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require controlled backlog processing and reconciliation before resolving payment or ledger incidents.
- Require a formal command handoff for incidents extending beyond four hours or when fatigue impairs a role holder.
14. Standardize internal and customer communications (depends on: 5, 6, 10)
Communicate known impact early without waiting for root cause. The Communications Lead uses approved facts and always states the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish customer status within 15 minutes when customer-visible and legally safe, then update at least every 30 minutes.
- For SEV1, brief internal stakeholders and publish status within 15 minutes, then update every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send notice within 30 minutes, then update every 60 minutes.
- For SEV3, communicate directly to affected customers only when impact or contract terms require it.
- Give account managers an approved statement and affected-customer list within 30 minutes for SEV0 or SEV1 and within 60 minutes for SEV2.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. Describe affected capabilities, symptoms, workarounds, and the next update.
- Do not speculate about root cause, blame, data integrity, security scope, or recovery time.
- Post a Monitoring update within 30 minutes of mitigation. Post Resolved only after stability and required reconciliation.
- Provide a customer-facing incident summary within five business days for qualifying incidents.
- Record and approve any delay or restriction of public details during an active security threat.
15. Operationalize regulatory, partner, contract, and credit decisions (depends on: 4, 5, 14)
Give Legal and Compliance a timed decision process while keeping technical command with the incident commander. Record both reportable and non-reportable determinations.
- Build a jurisdiction and obligation matrix covering applicable NYDFS requirements, state breach laws, GLBA or FTC requirements, PCI obligations, money-transmitter rules, sponsor banks, payment networks, cyber insurance, and customer contracts.
- Validate applicability and deadlines with counsel rather than assuming every operational incident is reportable.
- Begin reportability assessment immediately for every SEV0, security event, suspected ledger-integrity event, and relevant SEV1.
- Target a documented initial legal assessment within one hour for SEV0 and within two hours for other potentially reportable events.
- Record the decision, evidence, approver, legal deadline, submission owner, and confirmation of delivery.
- Maintain tested 24×7 contacts for regulators, banks, networks, insurers, outside counsel, and critical vendors.
- Encode customer-specific notification clocks and channels in the customer record used by the Communications Lead.
- Have Finance calculate affected minutes, likely credits, and contractual exposure from the incident record within five business days.
- Review whether proactive credits or claims-based handling applies by customer segment and contract.
16. Make postmortems mandatory, consistent, and blameless (depends on: 5, 6, 10)
Use one review standard to learn from incidents and test whether controls worked. Keep learning reviews separate from performance or misconduct processes.
- Require a postmortem for every SEV0 and SEV1.
- Require one for customer-visible SEV2, customer-first detection, incidents lasting more than two hours, contractual breaches, repeat failures, control gaps, and ledger-integrity near misses.
- Produce the factual draft within three business days, conduct the review within five, and publish the approved version within 10.
- Use one template covering summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, lessons, and actions.
- Analyze why detection was not earlier and why mitigation took as long as it did.
- Examine technical, organizational, process, testing, dependency, and incentive factors rather than forcing a single root cause.
- Use a trained facilitator and written blameless rules. Describe decisions in the context and information available at the time.
- Publish broadly useful findings internally while maintaining restricted versions for security, privacy, personnel, or privileged material.
- Review postmortem quality and recurring factors in a weekly Incident Review Board.
17. Enforce action ownership and effectiveness tracking (depends on: 16)
Treat corrective actions as risk commitments rather than suggestions. Closing a ticket is insufficient without evidence that the control or system behavior improved.
- Give every action one named individual owner, manager, priority, due date, expected risk reduction, verification method, and linked work item.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with intermediate milestones.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Place recurrence-prevention actions for SEV0 and SEV1 ahead of discretionary feature work unless an executive accepts the residual risk.
- Escalate overdue high-risk actions to the manager after seven days, director after 14 days, and CTO after 30 days.
- Require written residual-risk acceptance, compensating controls, and a new review date when high-risk work is deferred.
- Verify completed actions through tests, telemetry, drills, or production evidence.
- Triage the 53 currently open historical actions within 30 days. Complete, re-plan, or formally accept the risk, prioritizing ledger, regional, security, and detection items.
18. Measure outcomes, controls, business impact, and human load (depends on: 10, 11, 14, 17)
Use a balanced scorecard so teams are not rewarded for suppressing alerts or avoiding incident declarations. Report medians and 90th percentiles, not averages alone.
- Measure time from first impact to internal detection, declaration, acknowledgement, command assignment, mitigation, and resolution.
- Track customer-first detection, missed escalations, status-page timeliness, update-cadence compliance, and role conflicts.
- Track incident count, recurrence, availability, error-budget burn, affected payment value, delayed transactions, reconciliation breaks, and SLA credits.
- Track page volume, actionability, duplicates, after-hours pages, missed acknowledgements, and load per responder.
- Track postmortem timeliness, action completion, action age, verified effectiveness, and repeated contributing factors.
- Track rotation size, duty frequency, recovery days, swaps, attrition signals, and quarterly responder sentiment.
- Hold a weekly Incident Review Board for incidents, actions, missed controls, and noisy alerts.
- Hold a monthly executive reliability review for trends, investment decisions, contractual exposure, and accepted risks.
- Hold a quarterly resilience and controls review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile dashboard data against customer cases and sampled incident records monthly to identify missing incidents or metric gaming.
19. Train participants and address pager resistance (depends on: 6, 7, 8, 13, 14, 16)
Introduce the operating model as a fair exchange, not an audit mandate. The message is that people are paid, paged only for accepted services, supported by trained command, and given capacity to fix defects.
- Train all employees to recognize impact, declare an incident, and find the incident channel and customer status page.
- Give all 260 engineers role-based training in severity, escalation, evidence preservation, handoff, and financial-integrity precautions.
- Certify incident commanders through instruction, simulation, two shadowed events or exercises, and observed performance.
- Train Communications Leads in status writing, account-manager briefing, contractual clocks, and legal escalation.
- Train scribes in timeline quality, decision capture, fact-versus-hypothesis labeling, and evidence handling.
- Require responders to demonstrate access, dashboards, rollback, runbooks, and escalation competence before primary duty.
- Require two shadow shifts before independent primary on-call.
- Appoint one adoption champion in each team and hold weekly office hours during rollout.
- Publish compensation, fatigue protections, service boundaries, and the accommodation process before assigning shifts.
- Use paid working time for training, exercises, shadowing, runbook work, and certification.
- Survey engineers at baseline, day 60, day 120, and quarterly thereafter.
20. Pilot on the payment critical path (depends on: 8, 9, 11, 12, 13, 14, 17, 19)
Run a four-to-six-week pilot across the highest-risk customer journey. Use real incidents and exercises to correct the process before wider rollout.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, authentication or API edge, settlement, and Support intake.
- Activate paid primary and secondary rotations, central command, communications, the incident record, status templates, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new system authoritative.
- Review every pilot page within one business day for routing, actionability, responder load, and missing context.
- Have the program team coach incidents without silently taking command from assigned role holders.
- Hold a weekly pilot retrospective and correct critical process or tool defects within 48 hours.
- Require 95% command assignment within five minutes, 95% timely communications, no unpaid pages, complete postmortems, and at least a 50% page-noise reduction before expansion.
21. Roll out by customer journey and risk (depends on: 18, 20)
Expand in fixed waves rather than waiting for every team to become perfect. Apply explicit readiness gates and time-limited exceptions.
- Days 0–7: operate the interim declaration, command, and communications process.
- By day 30: complete Tier 0 ownership, activate the first certified command roster, approve compensation, and begin the pilot.
- By day 60: provide compensated 24×7 coverage for every Tier 0 and Tier 1 customer journey and route all critical pages through the new platform.
- By day 90: complete critical detection upgrades, customer communications, regulatory playbooks, and the first major alert-noise reduction.
- By day 120: assign every production service a response tier, owner, tested escalation, and appropriate coverage model.
- Roll out remaining teams in two-to-three-week waves ordered by customer and dependency risk.
- Require each wave to pass ownership, alert, runbook, access, training, compensation, and tabletop gates.
- Disable legacy paging paths after verified cutover rather than leaving ambiguous parallel obligations.
- Publish a weekly adoption dashboard by team and escalate failed gates as business risks.
- Never start mandatory night coverage before compensation, staffing, training, and access are ready.
22. Exercise command, regional resilience, and ledger recovery (depends on: 9, 13, 14, 15, 19)
Validate the process under realistic conditions before depending on it during a crisis. Use the same action-tracking rules for exercises and real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops covering payment failure, customer-first detection, third-party failure, and ambiguous ownership.
- Exercise loss of one AWS region, Kubernetes degradation, PostgreSQL failure, suspected duplicate payments, queue backlog, and settlement risk.
- Exercise simultaneous operational and security events to test command and disclosure boundaries.
- Exercise loss of chat, status-page, identity, or paging providers using telephone and offline fallbacks.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors.
- Validate backups, restoration, RPO, RTO, failover prerequisites, financial controls, and post-recovery reconciliation.
- Avoid uncontrolled production-ledger experiments; use staging, replicas, simulations, or tightly governed production tests.
- Complete at least two cross-company exercises before the audit, including one overnight or unannounced paging test.
23. Test SOC 2 operating effectiveness before the auditor (depends on: 4, 18, 21, 22)
Demonstrate that the controls operate consistently, not merely that policies exist. Correct failures through tracked remediation rather than rewriting historical records.
- Preserve policy approvals, service ownership, schedules, compensation activation, access reviews, training, certifications, incidents, communications, postmortems, actions, and exercises.
- Sample evidence monthly from initial signal through verified action closure.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Use exceptions to document missed acknowledgements, late communications, incomplete records, and compensating controls.
- Conduct a formal mock audit in month 7 using the populations and interview roles expected by the external auditor.
- Trace at least one SEV0 or SEV1, one SEV2, a customer-reported event, and an exercise end to end.
- Remediate evidence and operating gaps before external fieldwork.
- Brief commanders, responders, Support, and Compliance for auditor interviews without scripting inaccurate answers.
24. Institutionalize continuous improvement (depends on: 18, 21, 23)
Keep incident management active after the audit by assigning permanent owners, budgets, and review cycles. Use incident trends to drive architectural investment.
- Assign permanent owners for the policy, service catalog, paging platform, status page, training program, metrics, and evidence repository.
- Review severity thresholds, communications timing, staffing, and compensation annually and after material process failures.
- Recertify commanders and communications leads annually through observed exercises.
- Review recurring failure families quarterly and require executive action when remediation repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fatigue, fairness, tool friction, and psychological safety.
- Report severe incidents, SLA exposure, overdue risks, resilience investments, and customer-first detection to the board or risk committee quarterly.
- Prioritize reduction of the shared-ledger concentration risk, stronger regional independence, deployment safety, graceful degradation, and automated mitigation.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
Previous Proposal 3 (ID: 2437bc94-f488-46b5-866e-01e1d4afbf3b, Agent: qwen3.8-max_refine_3, LLM: alibaba/qwen3.8-max):
Estimated Complexity: high
Success Metrics: - Median time to detect customer-impacting incidents falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median mitigation time falls from 3h10 to under 90 minutes by day 120 and under 60 minutes by month 9.
- A named incident commander is assigned within 5 minutes in at least 95% of SEV1 and SEV2 incidents; zero incidents remain unowned for more than 10 minutes.
- First status-page update occurs within 15 minutes of SEV1 declaration and 30 minutes of SEV2 declaration in at least 95% of qualifying incidents.
- Monthly paging volume falls from 3,400 to under 500 actionable pages within six months, with alert actionability above 80%.
- Out-of-hours pages average no more than 2 per responder per week; sustained breaches trigger mandatory alert remediation.
- Six legacy alerting tools are consolidated into one paging and incident platform, with legacy paging paths disabled by week 16.
- 100% of Tier 0 and Tier 1 services have a named owning team, escalation path, dashboard, and runbook by day 60.
- All 28 teams are onboarded with readiness gates by week 22; every 24x7 critical-path rotation has at least six trained responders.
- 30 or more certified incident commanders and 20 or more certified communications leads provide continuous primary and secondary coverage.
- Paid on-call is approved by HR, legal, and finance and active in payroll before any new mandatory night rotation begins.
- 100% of required SEV1 and SEV2 postmortems are drafted within 3 business days and published within 10 business days in the standard template.
- Postmortem action closure rises from 17% to at least 90% of high-priority actions completed by their due date within two quarters.
- All 53 open historical actions are triaged within 30 days; high-risk unaccepted items are completed or formally risk-accepted within 90 days.
- Repeat incidents from a known unaddressed contributing factor decline by at least 50% within six months.
- Annualized SLA credits fall from $1.3M to under $400k within 12 months.
- Monthly availability meets or exceeds 99.95% by month 6, with exceptions reviewed at the executive reliability meeting.
- At least two cross-company exercises, including regional and ledger scenarios, are completed before the SOC 2 audit, with critical findings tracked.
- The month-6 internal dry-run audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete evidence.
- SOC 2 Type II incident-response controls pass with zero exceptions at the month-8 audit.
- On-call sentiment improves quarter over quarter; fewer than 10% of engineers report unwillingness to participate in their owned-service rotation by month 6.
- No increase in attrition among engineers on active rotations compared with baseline.
Steps (27):
1. Executive mandate, program funding, and governance
Convert the CEO email into a **written mandate within 48 hours**. Name one accountable program owner and a small decision group. Time-box design to five weeks so rollout starts well before the audit.
- Appoint the CTO as executive sponsor and a Head of Reliability or Incident Management as program owner with full-time authority.
- Create an 8–10 person design working group: SRE/platform lead, payments and ledger engineering managers, support lead, compliance, legal, HR, finance, and two rotating engineering managers.
- Approve budget lines: tooling consolidation, on-call compensation, training, coaching, and 2–3 dedicated program staff.
- State the non-negotiables: paid on-call, named service ownership, one severity scale, one paging platform, mandatory postmortems, and protected engineering capacity for reliability actions.
- Fix the master timeline: interim controls week 1, design weeks 1–5, pilot weeks 6–11, full rollout weeks 12–22, internal audit rehearsal month 6, audit-ready month 7.
- Publish a one-page charter to the company that says incident response is a company operating process, not an optional team practice.
2. Evidence baseline from incidents, alerts, and coverage gaps (depends on: 1)
Before changing anything, create an auditable **before picture** from the last 12 months. This baseline drives the severity design, staffing model, and executive reporting.
- Reconstruct all 31 customer-impacting incidents: detection source, first owner, severity, mitigation time, credits paid, and whether command was clear.
- Write a specific case review of the two incidents with unclear ownership for more than one hour.
- Inventory the six alert tools, alert volume by team, noise rate, rules with no owner, and rules with no runbook.
- Map the 16 teams without on-call and the 12 teams with unpaid on-call.
- Review the 53 open postmortem actions and triage the highest-risk items first.
- Freeze baseline metrics: 22-minute detection, 40% customer-first detection, 3h10 mitigation, $1.3M credits, 3,400 alerts, 85% noise, and 17% action closure.
3. Stakeholder listening, resistance mapping, and co-design group (depends on: 1)
Treat engineer pushback as **design input, not an attitude problem**. The objection to carrying a pager for other teams' code must shape ownership, routing, compensation, and command staffing.
- Interview leads from all 28 teams plus support, customer success, sales, compliance, and HR within two weeks.
- Separate the real causes of resistance: unpaid night work, unfamiliar systems, poor runbooks, unfair load, fear of blame, or unclear authority.
- Recruit 10–15 respected engineers and managers as co-designers so the model is built with the teams.
- Survey baseline on-call sentiment, trust in alerts, and psychological safety; repeat at 90, 180, and 365 days.
- Document the central promise: responders are paged for services they own, trained incident commanders coordinate, and all on-call work is compensated.
4. Service ownership catalog, criticality tiers, and dependency map (depends on: 2)
No alert can page the right team until every service has a **named owner**. Build the catalog as the routing foundation for on-call, severity impact mapping, status-page components, and audit evidence.
- Assign one accountable team to each of the 180 services, with an engineering manager, Slack channel, escalation policy, dashboard, and runbook link.
- Tier services: Tier 0 for money movement, ledger integrity, authentication, settlement, and shared PostgreSQL; Tier 1 for customer-facing degradable services; Tier 2 for internal or batch services; Tier 3 for non-critical services.
- Map customer journeys to services, databases, regions, third parties, and contractual SLA components.
- Mark orphan and shared services; require ownership, reassignment, or decommissioning within 30 days.
- Document cross-region dependencies, failover constraints, and services that defeat nominal regional redundancy.
- Treat missing ownership for Tier 0 or Tier 1 services as an executive escalation and a release-blocking risk.
5. Severity scale, declaration rights, and automatic triggers (depends on: 2, 4)
Adopt one severity scale so declaration is a lookup, not a debate. **Anyone may declare; only the incident commander may downgrade.** When uncertain, start higher.
- SEV1: money movement stopped or incorrect, ledger integrity in doubt, confirmed security or data event, both regions impaired, or broad SLA-credit exposure.
- SEV2: major degradation, settlement window at risk, one or more critical customers fully down, or likely contractual breach.
- SEV3: limited impact with a workaround, no financial-integrity or security risk.
- SEV4: no customer impact; handled by ticket during business hours.
- Define fixed triggers for each level: pages, command staffing, bridge call, status-page timing, account-manager outreach, executive notification, and postmortem obligation.
- Add automatic escalation: SEV3 open more than 2 hours becomes SEV2; any incident touching the shared ledger or unclear ownership becomes SEV2 unless the commander documents otherwise.
- Publish a decision tree and 10–12 worked examples from the actual 31 incidents.
6. Incident roles, authority, and handover rules (depends on: 5)
Solve the **nobody-was-in-charge failure** by making command explicit, trained, and transferable. Separate coordination from technical remediation so responders are not asked to debug unfamiliar code.
- Incident Commander: owns severity, priorities, escalation, mitigation strategy, role assignment, handoffs, and closure. Does not write code during the incident. May freeze deploys, invoke pre-approved failover, and pull any on-call responder.
- Communications Lead: owns status page, internal updates, account-manager briefs, and coordination with legal or compliance.
- Scribe: maintains the timestamped record of decisions, actions, role changes, and customer communications.
- Subject-matter responders: engineers from owning teams who diagnose and remediate within their accepted domain.
- Executive liaison: required for SEV1; shields the commander from executive questions and owns board or regulator escalation.
- Require the commander to claim the role within 5 minutes of SEV1 or SEV2 declaration, state it in the incident channel, and record every handover.
- Allow role combination only for SEV3 or SEV4; require separate people for command, communications, and primary technical work at SEV1.
7. Paid on-call, fatigue safeguards, and HR compliance (depends on: 3)
Unpaid on-call is a retention, fairness, and legal risk in New York. **Pay must be approved before any new mandatory rotation starts.** This is the fastest way to reduce pager resistance.
- Create a weekly stipend for primary and secondary on-call, differentiated by tier and night responsibility.
- Add per-page or per-incident pay for out-of-hours activation, plus guaranteed recovery time after night work or major incidents.
- Create a separate stipend for duty incident commanders and communications leads because command is a heavier burden.
- Verify FLSA, New York wage-hour, overtime, holiday, and payroll treatment with HR, legal, and finance.
- Cap rotation frequency: no more than one primary week in four or one week in six depending on staffing; no simultaneous primary assignments.
- Require at least six trained responders for any 24x7 rotation; fund hiring or service reassignment where teams are too small.
- Publish the compensation package and payroll start date before asking engineers to join rotations.
8. On-call staffing model and rotation rules (depends on: 4, 6, 7)
Do not force 28 identical night rotations. Use a **layered model**: central command coverage, critical-path team coverage, and business-hours coverage for lower-tier services.
- Company incident-command and communications rotation: 24x7 pool of 25–35 trained volunteers and designated senior staff, giving primary plus secondary coverage at all times.
- Critical-path teams: 24x7 primary and secondary on-call for Tier 0 and Tier 1 domains, including payments, ledger, authentication, API edge, Kubernetes platform, and shared PostgreSQL.
- Other teams: business-hours on-call with a documented night escalation list owned by the engineering manager.
- Group services into 8–12 coherent domains so rotations are sustainable; no rotation with fewer than six people is approved without an executive exception.
- Require handoff overlap, shadow shifts before first primary duty, and no on-call during approved PTO.
- Define cross-team pull rules: the commander may page another team's on-call with a 10-minute acknowledgement obligation; this is coordinated by command, not pushed onto responders.
- Publish schedules, swap rules, and load limits in the paging tool.
9. Alert quality standard, page budget, and noise controls (depends on: 2, 4)
With 3,400 alerts and 85% noise, detection fails because people stop trusting pages. Make alert quality a **condition of paging anyone**.
- Every page must have a named owning team, customer or SLO impact, severity, dashboard, runbook, expected action, and escalation policy.
- Page on customer-visible symptoms: payment success rate, latency, settlement deadlines, error-budget burn, ledger integrity, and replication health.
- Demote cause-based CPU, memory, or infrastructure-only alerts to dashboards or tickets unless they map to a customer journey.
- Set a page budget: maximum 2 out-of-hours pages per person per week; breach triggers a mandatory alert-tuning sprint.
- Auto-flag alerts that fire repeatedly without action or have high no-action acknowledgement rates.
- Require shadow mode for new alerts before they page humans, except for documented emergencies.
- Review alert quality monthly by team and publish a noise leaderboard.
- Target fewer than 500 actionable pages per month and above 80% actionability within six months.
10. Detection uplift across payments, ledger, and customer signals (depends on: 4, 9)
Customers detected 40% of incidents first. Detection must shift to **payment outcomes, ledger integrity, and inbound customer signals**, not host metrics alone.
- Define SLIs and SLOs for payment initiation, authorization, settlement, reconciliation, refunds, API availability, and reporting freshness.
- Run synthetic end-to-end payment tests from outside the platform in both AWS regions every 60 seconds.
- Add continuous ledger assurance: double-entry reconciliation, replication lag, failover readiness, disk pressure, and settlement-window countdown alerts.
- Monitor per-customer anomalies for top accounts so a single-tenant outage is detected before the account manager calls.
- Convert high-priority support tickets and account-manager reports into incident candidates within 5 minutes.
- Track detection source for every incident; make customer-first detection a reviewed defect with a corrective action.
- Add partner, banking, and card-network notification intake into the same declaration path.
11. Single incident platform and alert-tool consolidation (depends on: 5, 8, 9)
Collapse six tools into **one paging and incident-management system** with one queue, one timeline, and one audit trail. Do not create a parallel audit process.
- Select an integrated stack: paging and on-call scheduling, incident workflow, Slack or chat integration, bridge calling, status-page API, and ticketing integration.
- Implement one-command declaration that creates the incident record, channel, bridge, severity label, role prompts, and clock.
- Migrate alert sources by wave; retire a legacy paging path only after routing, ownership, and acknowledgement tests pass.
- Automate evidence capture: timestamps, acknowledgements, role assignments, severity changes, communications, and postmortem links.
- Integrate with the service catalog, on-call schedules, Jira or equivalent action tracker, customer-success tooling, and the status page.
- Verify the platform works during a single-region failure, including SMS, phone, and offline fallback paths.
- Set a hard date after which pages outside the chosen platform are not valid on-call obligations.
12. Escalation paths and five-minute command rule (depends on: 6, 8, 11)
Create one path from **signal to named commander in under five minutes**, any hour of the day. Make escalation automatic and time-bound.
- Accept declarations from automated alerts, engineers, support, account managers, partners, and customers through the same command.
- For SEV1 and SEV2, page the duty incident commander and owning-team primary immediately.
- Acknowledgement ladder: primary 5 minutes, secondary 10 minutes, manager 15 minutes, director or executive 20 minutes.
- If no commander claims the incident within 5 minutes, the platform assigns and announces one; the assignee may hand over but cannot leave the incident unowned.
- For SEV3, require team acknowledgement within 30 minutes; otherwise create a tracked work item.
- Give the commander pre-approved authority to invoke regional failover, ledger read-only mode, feature kill switches, and partner notifications without waiting for executive sign-off.
- Record every missed acknowledgement and escalation failure for weekly review.
13. Internal communications protocol (depends on: 6, 11)
Separate the **working incident channel from the audience channel** so responders can work and executives, support, and sales stay informed without interrupting the commander.
- Create one incident channel and one bridge per SEV1–SEV3 incident; use a read-only broadcast channel for executives and support.
- Set update cadence: every 15 minutes for SEV1, 30 minutes for SEV2, and at state changes for SEV3.
- Use a fixed update template: impact, customer-visible symptoms, current action, ETA or next update, commander, and communications lead.
- Brief support and customer success with a live affected-customer list and approved holding statements within 15 minutes of SEV1 or SEV2.
- Require executives to route questions through the executive liaison; the commander is not interrupted.
- Define fatigue rules and formal handover for incidents lasting more than 4 hours.
14. Customer status page, account-manager outreach, and SLA credit workflow (depends on: 5, 13)
Replace ad-hoc status updates with a **timed, owned, template-driven customer communication process**. Link incidents to SLA credits so finance and customer success are not surprised.
- Publish status-page updates within 15 minutes of SEV1 declaration and 30 minutes for SEV2; update every 30 or 60 minutes until resolution.
- Assign the communications lead as the single author; use pre-approved templates reviewed by legal and communications.
- Map status-page components to customer capabilities: payments, settlement, reporting, API, onboarding, and regional availability.
- For top accounts, require direct account-manager outreach within 30 minutes of SEV1 with an approved briefing pack.
- Send proactive email or webhook notifications to subscribed customers for SEV1 and SEV2.
- Publish a resolution notice within 30 minutes of mitigation and a customer-facing incident summary within 5 business days for SEV1.
- Automate SLA credit calculation from incident duration and affected capability; review with finance and legal within 5 business days.
- Track credits by incident and root-cause family to guide reliability investment.
15. Regulatory, legal, and partner notification playbook (depends on: 5, 14)
In payments, some incidents are reportable and the clock starts at detection. Make regulatory assessment a **mandatory step in the incident process**, not an afterthought.
- Map obligations: NYDFS cybersecurity-event rules, state breach laws, GLBA safeguards, PCI DSS if in scope, money-transmitter duties, card-network and sponsor-bank contracts, and cyber-insurer notice.
- Require a legal or compliance reportability assessment within 2 hours for every SEV1 and every security-related SEV2, even when the answer is not reportable.
- Maintain a 24x7 contact matrix for regulators, sponsor banks, card networks, outside counsel, insurers, and law enforcement.
- Pre-draft notification templates and preserve legal review.
- Encode customer-contract notification deadlines into account tiers and the communications workflow.
- Record every reportability decision, approver, deadline, and submission confirmation in the incident record.
16. Postmortem standard and blameless review (depends on: 5, 6)
Replace inconsistent postmortems with a **mandatory, blameless, single-format process**. The discipline comes from deadlines, facilitation, and action tracking.
- Require postmortems for all SEV1 and SEV2 incidents, customer-detected incidents, incidents over 2 hours, repeat failures, and ledger near-misses.
- Draft within 3 business days, peer review within 5, publish within 10 for SEV1 and SEV2.
- Use one template: summary, impact, timeline, detection analysis, response analysis, contributing factors, what worked, what failed, and action items.
- Make blamelessness explicit: focus on systems and decisions, not individual fault; never use postmortems in performance discipline.
- Hold a weekly incident review board to review postmortems, ratify severity, and challenge weak actions.
- Maintain a searchable postmortem library and quarterly recurring-cause analysis.
- Require a trained facilitator for major reviews; the incident commander attends but does not facilitate.
17. Action-item tracking, ownership, and delivery gates (depends on: 16, 11)
Only 11 of 64 actions were closed. Give postmortem actions the same status as **customer commitments**, with named owners and visible escalation.
- Create every action as a ticket with one named individual owner, priority, due date, and verification method.
- Use delivery classes: containment within 7 days, corrective work within 30 days, strategic work within 90 days.
- Reserve 15–20% of team sprint capacity for reliability and incident actions.
- Escalate overdue items: manager at 7 days, director at 14 days, CTO dashboard at 30 days.
- Block related feature releases when overdue P0 actions prevent recurrence of a severe incident.
- Require director approval and documented residual risk acceptance for overdue high-risk items.
- Verify effectiveness after completion; closing a ticket without evidence does not close the action.
- Target 90% of high-priority actions completed on time within two quarters.
18. Runbooks, critical-incident playbooks, and readiness bar (depends on: 4, 8)
Poor runbooks are a real cause of pager resistance and slow mitigation. Define a **minimum readiness bar** before a service is allowed to page anyone at night.
- Require for every Tier 0 and Tier 1 service: architecture summary, dependencies, dashboards, alert-to-runbook map, rollback procedure, feature flags, escalation contacts, and customer-impact statement.
- Write major playbooks for shared PostgreSQL ledger failure, regional failover, Kubernetes control-plane loss, payment-processor outage, settlement-window breach, duplicate-payment suspicion, and security compromise.
- Define ledger recovery rules: failover procedure, read-only degraded mode, reconciliation, RPO/RTO, and data-loss tolerance approved by executives.
- Test runbooks in drills at least twice a year; mark untested runbooks stale.
- Prevent paging alerts for services without readiness sign-off unless the engineering manager accepts the gap in writing.
- Keep runbooks linked from every alert and incident template.
19. Training, certification, and role readiness (depends on: 6, 12, 13, 16)
Command and communications are skills. Build a **tiered certification path** so rotations are staffed by people who have practiced, not by whoever is around.
- All employees: 1-hour module on declaring incidents, finding the incident channel, and reading the status page.
- All responders: half-day training on severity, acknowledgement, escalation, runbooks, and evidence hygiene.
- Incident commanders: 2-day course plus two shadowed incidents and one simulation before certification.
- Communications leads: training on status-page writing, customer language, account-manager briefs, and regulatory triggers.
- Scribes: training on timeline discipline and audit evidence.
- Certify for 12 months; renew through a simulation.
- Require shadow shifts before independent primary duty; no new hire holds primary within 90 days.
- Publish the certification register as an audit artifact.
20. Simulation program and game days (depends on: 19, 11, 18)
Rehearse the process before it meets a real SEV1. Simulations build commander confidence, expose runbook gaps, and produce audit evidence.
- Run monthly 60-minute tabletops using real incidents from the 31-incident baseline.
- Run quarterly game days covering regional failover, ledger replica promotion, dependency failure, partner outage, and security event.
- Run twice-yearly unannounced paging drills to measure night acknowledgement times.
- Include support, account managers, legal, compliance, and executives in at least one exercise per quarter.
- Produce tracked action items from every exercise using the same board as real incidents.
- Measure time to commander, time to first status update, and time to mitigation decision.
21. Critical-path pilot and gate review (depends on: 7, 10, 11, 14, 18, 19)
Prove the model on the highest-risk services with willing teams before full rollout. Run a **six-week pilot with daily feedback and public exit criteria**.
- Pilot with payments, ledger/database, platform/Kubernetes, API edge, authentication, and support intake.
- Activate severity scale, duty commanders, paid rotations, single paging platform, alert budget, status-page policy, and postmortem process.
- Hold a weekly pilot retrospective and fix process defects quickly.
- Validate night acknowledgement, cross-team pull response, severity clarity, and compensation payroll.
- Exit gate: commander assigned within 5 minutes in 95% of incidents, status page on time, page noise down at least 50%, postmortems on time, and positive on-call sentiment.
- Publish pilot results to the whole company as the main adoption argument.
22. Wave rollout across all 28 teams (depends on: 21)
Roll out by criticality and dependency, not by calendar alone. Use **readiness gates** so teams are not forced live without coverage.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1 teams. Wave 3: Tier 2 teams. Wave 4: Tier 3 and internal platforms.
- Gate per team: catalog entry complete, alerts migrated, runbooks ready, six-person rotation staffed, one commander candidate nominated, compensation in payroll, and one tabletop passed.
- Assign a program coach to each wave for three weeks.
- Freeze legacy paging paths for each team after successful onboarding.
- Publish a live adoption scoreboard by team.
- Complete all teams by week 22, leaving months of operating evidence before the audit.
23. Metrics, dashboards, and review cadence (depends on: 11, 16, 21)
Instrument the process itself. Leadership must see whether the program is working, and auditors must see **operating evidence, not retrospective paperwork**.
- Track response metrics: MTTD, time to declare, time to commander, acknowledgement time, mitigation time, resolution time, and customer-first detection rate.
- Track quality metrics: page volume, alert actionability, missed pages, postmortem timeliness, action closure rate, and status-page compliance.
- Track business metrics: availability against 99.95%, SLA credits, repeat incidents, and error-budget consumption.
- Track people metrics: on-call load, night pages per person, recovery time usage, sentiment, and attrition signals.
- Hold weekly incident review, monthly reliability review, quarterly executive review, and annual policy review.
- Publish dashboards internally and give every metric a target and owner.
24. Change management, incentives, and culture (depends on: 3, 7, 21)
The process will be judged on fairness. Communicate the deal repeatedly and make participation **recognized, compensated, and safe**.
- Core message: you are paid, paged only for what you own, supported by a trained commander, and given real capacity for actions.
- Run CTO all-hands, team roadshows, office hours, FAQ, and an internal incident-management hub.
- Add incident response and reliability work to promotion criteria and manager objectives.
- Recognize good postmortems, alert-noise reduction, and calm incident leadership.
- Provide a written path for engineers who cannot do nights; cover those shifts with paid volunteers or adjusted staffing.
- Prohibit retaliation for good-faith declaration or escalation.
- Publish sentiment survey results, including bad news, to maintain credibility.
25. SOC 2 evidence design and internal dry-run audit (depends on: 16, 17, 22, 23)
Design the process so audit evidence is a **by-product of normal operations**. Test it internally before the external auditor does.
- Map the process to SOC 2 criteria: incident identification, response, recovery, monitoring, communication, control activities, availability, and corrective action.
- Approve versioned policies: incident response policy, severity standard, on-call policy, communications policy, postmortem standard, and exception process.
- Retain incident records, paging logs, status-page history, postmortems, action tickets, training records, drill records, and access reviews for the audit period.
- Record regulatory reportability decisions even when no notification is required.
- Run an internal dry-run audit at month 6: sample at least 15 incidents and trace evidence end to end.
- Fix gaps with at least four weeks before the audit.
- Maintain an exception and remediation log instead of claiming perfection.
26. Program risk register and contingency planning (depends on: 1)
Name the likely failure modes now and pre-commit responses. Review the register monthly with the executive sponsor.
- Commander volunteer shortfall: roster command duty for senior engineers and managers until the pool reaches 25 trained people.
- Compensation delay: use immediate time-off-in-lieu plus phased stipend, but do not start mandatory night rotations without pay.
- Tool migration slip: preserve single paging and routing scope; defer advanced automation if needed.
- Alert pruning causing missed incidents: move noisy alerts to ticket first, observe 30 days, then delete.
- Burnout in experienced teams: monitor load weekly and cap pages per person.
- Major SEV1 during rollout: program owner shifts to incident support, wave schedule slips one wave, sponsor is informed same day.
- Shared ledger concentration: track blast-radius reduction and failover improvements as top strategic actions.
27. Continuous improvement, maturity roadmap, and post-audit sustainability (depends on: 23, 25)
Prevent the classic failure where the process decays after the audit. Build the second-year operating plan before the first year ends.
- Hold quarterly process retrospectives with commanders, responders, support, and communications leads.
- Re-baseline targets every six months; tighten goals once current targets are met.
- Move from lagging metrics to leading indicators: error-budget burn, near-miss rate, drill performance, and action ageing.
- Plan year-two improvements: follow-the-sun coverage, automated mitigation, ledger blast-radius reduction, error-budget release gates, and per-customer real-time impact reporting.
- Keep annual policy review, certification renewal, drills, and board reporting on a permanent calendar independent of the audit cycle.
- Report quarterly to the board or risk committee on availability, credits, severe incidents, overdue actions, and resilience investment.
Previous Proposal 4 (ID: 185429cf-5cd1-451d-ab7e-34d64ca403df, Agent: grok4.6_refine_4, LLM: xai/grok-4.6):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 10 minutes by month 3 and under 5 minutes by month 9.
- Customer-first detection falls from 40% to under 20% by month 3 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by month 4 and under 60 minutes by month 12.
- A named Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes after week 4.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases by month 3, with required update cadence met in 95% of cases.
- Monthly pages fall from 3,400 to under 500 within 6 months, with actionability above 75%; out-of-hours pages at or under 2 per person per week.
- All six legacy alerting tools route through one paging platform by week 16; legacy paging paths disabled per wave.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by due date within two quarters; all 53 currently open historical actions triaged within 30 days of S2.
- SLA credits fall from $1.3M to under $400k in the first 12 months; customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known unaddressed cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar; every Layer B rotation has 6+ certified responders or a time-limited executive exception.
- Paid on-call policy approved by HR, Legal, and Finance and in payroll before any mandatory night rotation starts.
- All 28 teams onboarded by week 24 with 24x7 command cover from 25+ certified ICs and 15+ certified Comms Leads.
- Internal dry-run at month 6 passes a 15-incident evidence walkthrough; month-7 mock audit finds no unowned high-risk control gap; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment improves quarter over quarter; no increase in attrition among engineers on rotation; pulse at 60 and 120 days shows at least 70% agree rotations are fair and limited to services they own.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked actions; two cross-company exercises completed before the audit.
Steps (24):
1. Program charter, mandate, and funding
Convert the CEO email into a named program with one owner, a budget, and a deadline earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, 28 teams will opt out.
- Appoint a Program Lead (Head of Reliability) with direct CTO and CEO sponsorship.
- Form a small steering group: CTO, VP Eng, Head of CS/Support, CISO, Legal, Finance, HR. Not a 28-team committee.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Timeline: operating floor in 7 days; design weeks 1–6; pilot weeks 7–12; rollout weeks 13–24; mock audit week 28; SOC 2 at month 8.
- Fund tooling ($150–250k/yr), on-call pay (~$600k–1M/yr), 2–3 program FTEs, and reserved engineering capacity. Anchor the ask against $1.3M in credits plus unmeasured incident cost.
- Freeze baselines now: 31 incidents, MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, $1.3M credits, 3,400 alerts/month at 85% noise, 11 of 64 actions closed.
2. Immediate 7-day operating floor (depends on: 1)
Do not wait for tooling, compensation, or the audit. Put a minimum process in place this week so the next outage has a named commander.
Start retaining artifacts on day one. This week becomes the first audit evidence.
- Publish a one-page interim severity guide and a single declaration path through Slack, phone, and pager.
- Staff interim primary and backup Incident Commander 24x7 from existing on-call veterans and engineering managers. Compensate this duty retroactively.
- Require a named IC within 10 minutes for every suspected major incident. If nobody claims it, the duty manager is IC.
- Use one channel, one bridge, one timeline doc, and one naming convention for every major incident.
- Direct Support to escalate credible customer reports immediately. They do not wait for engineering confirmation.
- Triage all 53 open historical actions. Close, re-plan, or formally risk-accept. Ledger integrity, payment duplication, regional resilience, security, and detection come first.
- Hold a daily 15-minute ops review until the permanent process is live.
3. Forensic baseline of incidents and alerts (depends on: 1)
Rebuild the facts before locking design. This is the before picture for the CEO and the auditor.
Every later design choice should trace to this evidence.
- Re-code each of the 31 incidents: trigger, service, detection source, timestamps, who led, credits paid, root-cause family.
- Quantify the 40% customer-first detections and name the missing signal in each case.
- Reconstruct the two nobody-in-charge incidents minute by minute. Use them as the burning-platform story.
- Audit the six alerting tools: volume per tool and team, top 50 noisy rules, rules with no owner or runbook.
- Freeze the baseline numbers. Do not let them drift during design.
4. Listening tour and the fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the main delivery risk. Treat it as a design constraint, not an attitude problem.
The answer that will actually stick is **the deal**: you are paid; you are paged only for services you own; a trained commander runs the room; your postmortem actions get real sprint capacity.
- Interview all 28 teams plus Support, CS, and Sales in two weeks.
- Separate the objections: unpaid work, nights, unfamiliar code, bad runbooks, fear of blame. Each needs a different fix.
- Collect informal practices from the 12 teams already on-call. They are the pilot candidates and the veteran pool.
- Recruit 10–15 credible engineers as a design working group so the process is co-authored.
- Baseline sentiment on alert trust, on-call willingness, and burnout. Re-measure at 6 and 12 months.
5. Service ownership catalog and criticality tiers (depends on: 3)
You cannot page the right person across 180 services until each one has a named owner. This is the foundation of fairness, routing, and audit evidence.
Build a machine-readable catalog as the single source of truth.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, the shared PostgreSQL cluster, queues, partner banks, and customer-facing endpoints.
- Assign one owning team, named engineering manager, Slack channel, escalation policy, and dependency list per service.
- Tier 0: money movement, ledger, auth, shared Postgres, regional control plane. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map Tier 0/1 services to customer capabilities: initiation, authorization, settlement, reporting, onboarding.
- Assign coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
6. Severity scale and declaration rules (depends on: 3, 5)
Replace in-the-moment debate with a lookup. Four payments-specific levels, with worked examples drawn from the 31 real incidents.
Anyone may declare. Only the Incident Commander may downgrade, with recorded rationale. When unsure, start high.
- **SEV1**: money movement stopped or incorrect; ledger integrity in doubt; material security or data exposure; both regions impaired; or more than 10% of customers impacted. Pages IC, comms, scribe, SMEs, and exec liaison. Bridge in 5 minutes. Status page in 15 minutes. Mandatory postmortem and regulator assessment.
- SEV2: severe degradation; settlement window at risk; a strategic customer fully down; SLA breach likely. IC and SMEs paged. Status page in 30 minutes. Mandatory postmortem.
- SEV3: partial impact with a workaround; no credit exposure. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, or a repeat.
- SEV4: internal or minor. Ticket only. No page.
- Auto-escalate: any SEV3 open more than 2 hours, or any incident touching the shared ledger, becomes SEV2. Unknown impact after 15 minutes is raised, not sat on.
- Nobody is punished for over-declaring. Publish that rule in writing and repeat it.
7. Roles, authority, and ledger dual-control (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, transferable, and logged.
The Incident Commander owns the incident, not the fix, and **never types** in a production terminal.
- Incident Commander: declares severity, pulls anyone, freezes deploys, invokes failover, authorises spend. One IC at a time. Assumed within 5 minutes and announced in channel.
- Communications Lead: single voice to customers, status page, account managers, and the exec summary.
- Scribe: timestamped timeline, decisions, and open questions. Feeds the postmortem and the audit trail. Required for SEV1 and SEV2.
- Subject-matter responders: diagnose and mitigate only services they own, with access and runbooks.
- Executive Liaison (SEV1): shields the IC from exec questions; owns regulator and board escalation.
- The IC may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access rules.
- Handover is verbal and written, with time and new owner recorded. Roles may combine below SEV2; never at SEV1. The IC stays in charge if a VP joins.
8. Three-layer 24x7 coverage model (depends on: 5, 7)
Do not put 28 teams on night rotation. That is what engineers are rejecting.
Staff command centrally. Page engineers only for code their team owns.
- Layer A, company command: Duty IC plus secondary, Duty Comms, and a scribe pool. Target 25–35 certified ICs and 15–20 comms leads. Roughly one week per person every 6–8 months. Five-minute ack SLA, then secondary, then on-call director.
- Layer B, Tier 0/1 domains: group 28 teams into 8–12 product and platform domains (ledger app, Postgres platform, payments orchestration, auth, API edge, Kubernetes/platform, settlement). Primary plus secondary 24x7. Minimum six trained people. Target one week in six, never worse than one in four.
- Layer C, Tier 2/3: business-hours on-call. After hours the IC pages the EM, who holds a written escalation list.
- No engineer joins another team's responder pool without training, access, runbooks, shadow shifts, and both teams' acceptance.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner.
- Evaluate follow-the-sun coverage as a 12-month option, not a year-1 dependency.
9. Paid on-call, NY labor compliance, and fatigue rules (depends on: 4, 8)
Unpaid on-call is a retention risk and a New York legal exposure. Pay must be live in payroll before any mandatory night rotation starts.
Treat interrupted nights and recovery time as compensable work.
- Weekly stipend by layer, higher for 24x7 Tier 0/1 and Duty IC, lower for business-hours, with holiday and weekend premiums.
- After-hours call-out pay or TOIL. A paid recovery day after prolonged overnight work, a SEV1, or a qualifying SEV2. Managers cover the next day.
- HR, Legal, and Finance publish dollar amounts, FLSA exempt/non-exempt treatment, NY wage-hour rules, tax treatment, and payroll timing within 14 days of charter.
- Load rules: no primary on two rotations; no consecutive primary weeks; no on-call the week after a SEV1 you commanded.
- A person may declare temporarily unfit after overnight work with no performance penalty.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- If a rotation cannot staff six people, merge domains or hire. Do not run two-person 24x7.
- Model annual cost against $1.3M in credits and get it as a CFO/board line item.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts a month at 85% noise is why detection takes 22 minutes. Make quality a condition of paging a human.
A page is a product with a quality bar, not a dump of host metrics.
- Every paging alert must have a named owner, customer or SLO impact, runbook, tested threshold, severity mapping, dashboard, expected action, and dedup key. Fail any of these and it becomes a ticket or is deleted.
- Page on customer symptoms: SLO burn, error budget, settlement-queue age. Cause-based CPU and memory alerts become dashboards or tickets.
- Set a **page budget** of at most two out-of-hours pages per person per week. A breach triggers a mandatory tuning sprint and blocks new paging alerts for that team.
- Auto-quarantine alerts that fire more than five times a month with no action, or that have more than 70% no-action acks. Never silently disable without compensating detection and a recorded decision.
- Run new alerts in shadow for seven days unless an emergency exception is approved.
- Monthly per-team kill, tune, or keep review. Target under 500 pages a month and actionability above 75% within six months.
11. Customer-journey detection and ledger assurance (depends on: 5, 10)
Stop customers being the monitoring system. Detect on money-path outcomes, not host metrics.
Every postmortem will ask why a customer saw it first.
- Define SLOs per Tier 0/1 capability: initiation success, auth latency, settlement timeliness, API availability, reporting freshness. Set an internal target stricter than 99.95%.
- Run synthetic full-lifecycle payments from outside the platform, in both regions, every 60 seconds. Include a small real-value canary where legally feasible.
- Add ledger assurance: continuous double-entry reconciliation, replication lag, failover-readiness, unexpected balances, duplicate identifiers, settlement-window countdown.
- Add per-customer anomaly detection for the top 100 accounts.
- Auto-create a triage incident within five minutes when support tickets or AM reports match impact keywords.
12. Single incident and paging platform (depends on: 6, 8, 10)
Collapse six alerting tools into one queue, one timeline, and one audit record.
The platform must still work if one AWS region is down. Verify SMS and phone fallback, plus an offline runbook.
- Select one paging product, one incident layer, and one hosted status page.
- One Slack command declares the incident, creates the channel and bridge, pages the Duty IC, sets severity, and starts the clock.
- Ingest the six existing tools first. Deduplicate and route. Retire a legacy path only after named owners and two weeks of verified operation.
- Auto-capture timestamps, acknowledgements, role assignments, severity changes, comms, mitigated time, and resolved time. Export for SOC 2.
- Integrate the service catalog, Jira, Salesforce or CS tooling for affected-customer lists, and Zoom or Slack Huddle.
- Test paging, escalation, status publication, and conference access every week.
- After a team's wave, old paging paths are disabled, not left as a fallback.
13. Detection-to-command escalation path (depends on: 7, 8, 12)
Write the single path from something looks wrong to someone is in charge. Target: a commander in under five minutes, any hour.
If nobody claims IC in five minutes, the platform assigns the Duty IC. The assignee can hand over, not decline.
- Converge every entry point on the same declare command: alert, engineer, support, account manager, partner bank, SEV hotline.
- SEV1: page the owning primary immediately; secondary at 5 minutes unacked; domain manager and Duty IC at 10; exec liaison at 15.
- SEV2: primary ack in 10 minutes; IC assigned in 15.
- Human acknowledgement is required. Delivery to a device does not count.
- The IC can page any team's on-call, with a 10-minute ack obligation. This reciprocity makes single-team ownership viable.
- Unowned alerts go to Layer A command, then the missing owner record is a control defect.
- Pre-authorise regional failover, ledger read-only mode, and partner-bank notice so the IC does not wait for an executive. Dual-control still applies to ledger writes.
14. Live execution and major-incident playbooks (depends on: 7, 13)
Limit customer and financial harm before proving root cause. One procedure from the first minute to handback.
A service cannot page at night until it meets the readiness bar.
- Open channel, bridge, record, and timeline immediately for SEV1 and SEV2. The IC states severity, known impact, hypothesis, objective, roles, and next update time.
- Freeze unrelated production changes during SEV1. Record exceptions the IC approves.
- Prefer reversible mitigation: rollback, feature flag, traffic isolation, rate limit, partner reroute.
- Write playbooks first for Postgres ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank outage, settlement-window breach, suspected security compromise, and suspected duplicate payments.
- Guard against split-brain, replay, and duplication during regional or database recovery. Reconcile and process backlog before calling the incident resolved.
- Mitigation means customer impact ended. Resolution means stable, backlog processed, and ledger reconciled.
- Formal IC handover after 4 hours. Reopen if impact recurs during the stability window.
- Readiness bar per Tier 0/1 service: diagram, dependencies, dashboard, runbook, rollback, kill switch, escalation contacts, RPO/RTO. Untested runbooks are marked stale.
15. Internal, customer, and regulatory communications (depends on: 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved process. The Communications Lead is the single author.
State impact and the next update time. Never speculate on cause.
- Internal: one working channel, one bridge, one read-only broadcast for execs, Support, and Sales. SEV1 update every 30 minutes even if unchanged; SEV2 every 60 minutes.
- Executives ask questions only of the Executive Liaison. Publish this as a signed exec behaviour rule.
- Status page: SEV1 within 15 minutes, SEV2 within 30 minutes; then 30/60-minute updates; resolve notice within 30 minutes of mitigation. Templates pre-approved with Legal.
- Top 100 accounts: named AM contact within 30 minutes of SEV1, with a briefing pack from Comms. Long tail gets the status page plus email or webhook.
- Customer-facing summary within 5 business days for SEV1.
- Mandatory regulatory checkpoint on every SEV1 and every security SEV2, within 2 hours, recorded even when not reportable. Cover NYDFS Part 500 72-hour clock, state breach laws, GLBA/FTC, PCI if in scope, sponsor-bank and card-network windows, FinCEN/OFAC if relevant.
- Legal owns outbound regulatory letters. The IC owns facts. Security or Legal may limit public detail during an active threat, with the reason recorded.
- Encode bespoke customer-contract notice SLAs into account tiering.
16. SLA credit and financial-impact workflow (depends on: 6, 15)
Link incidents to money so severity, credits, and investment stay consistent. Finance should not learn about outages from invoices.
Make credit calculation an output of the incident record, not a negotiation.
- Agree availability measurement per contract and component with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and capability telemetry. Produce a proposed credit schedule within 5 business days of resolution.
- Posture: proactive credits for top-tier accounts; claims-based for the rest. Document the approval chain.
- Track credits per incident and root-cause family. Quarterly report which reliability investments would have prevented which credits.
- Target: cut credits from $1.3M to under $400k in 12 months. Use that delta as the ongoing business case.
17. Blameless postmortems and action tracking (depends on: 6, 12)
Eleven of 64 actions closed is the clearest process failure. Make learning mandatory and actions as binding as customer commitments.
Closing a ticket without evidence of effectiveness does not close the action.
- Mandatory for every SEV1 and SEV2, every customer-first detection, every incident over 2 hours, every repeat of a known cause, and every ledger near-miss.
- Draft within 3 business days, review within 5, published internally within 10. The IC owns the draft. The owning EM is accountable.
- One template: timeline, customer and financial impact, why detection was late, why mitigation took that long, contributing factors, what went well, actions.
- Blameless in writing. No individual named as a cause. HR and management commit that postmortems are never used in performance reviews.
- Every action gets a named person, priority, due date, Jira ticket, and verification method. P0 (prevents SEV1 recurrence) due in 30 days and committed into the next sprint before roadmap work. P1 in 60 days. P2 in 90 days.
- Teams reserve 15–20% of sprint capacity for reliability and incident actions.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Weekly Incident Review Board with engineering directors. Target 90% of P0/P1 actions closed on time within two quarters.
18. Training, certification, and commander academy (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Nobody holds independent duty untrained. Training happens in paid working time.
The certification register is an audit artefact.
- All employees, 1 hour: recognise impact, declare, find the channel and status page.
- Responder, half day: severity, escalation, runbooks, comms hygiene. Mandatory before joining a rotation.
- Scribe, 2 hours: timeline discipline. The entry point.
- Incident Commander, two days plus shadowing: command presence, decisions under uncertainty, severity, handover, exec management. Two shadowed incidents and one simulated SEV1 before certification.
- Communications Lead, one day: status writing, customer tiering, legal boundaries, regulator triggers.
- Certification valid 12 months, renewed via simulation.
- New joiners shadow two shifts and do not hold primary in their first 90 days.
- Appoint a champion in each of the 28 teams.
19. Simulations and game days (depends on: 14, 15, 18)
Rehearse before the next real SEV1. Use a standing calendar, not a one-off exercise.
Do not inject uncontrolled changes into the production ledger.
- Monthly 60-minute tabletop per engineering group, using a real incident from the 31.
- Quarterly full-scale game day: regional failover, ledger replica promotion, dependency failure. Whole role structure, timed.
- Twice-yearly unannounced paging drill, including nights, to measure real acknowledgement times.
- One security incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Also exercise status-page failure and loss of the primary chat or pager.
- Use replicas, staging, or tightly governed tests for ledger scenarios.
- Every exercise produces actions in the same tracker as real incidents.
- Complete two cross-company exercises before the SOC 2 audit.
20. Critical-path pilot (depends on: 9, 11, 12, 18, 19)
Prove the process on the highest-risk surface with willing teams before asking 28 teams to adopt it.
Publish a one-page result company-wide. That result is the adoption argument.
- Six-week Wave 0: ledger app, Postgres platform, payments orchestration, Kubernetes/platform, API gateway, Support intake, plus two of the 12 teams already on-call.
- Activate the full stack: severity scale, Duty IC, one pager, page budget, status page, mandatory postmortems, paid on-call.
- Parallel-run old paths for one week, then cut over. The program lead coaches every SEV2+ and does not secretly take command.
- Weekly retro. Expect 20–40 process defects and fix them in the standard before rollout.
- Exit gate: MTTD under 10 minutes on pilot services; IC assigned within 5 minutes in 95% of incidents; page volume down 50%; all postmortems on time; no unpaid pages; sentiment not worse.
21. Metrics, reviews, and error budgets (depends on: 12, 17, 20)
Instrument the process itself. Do not reward hiding incidents or suppressing pages.
Every metric has a target and a named owner. Dashboards are public inside the company.
- Response: MTTD, time to declare, time to IC, MTTA, MTTM, MTTR, incidents by severity, percent customer-first. Report median and p90, by tier, journey, and region.
- Quality: pages per person per week, actionability, budget breaches, postmortem on-time rate, action closure and age, missed acknowledgements.
- Business: credits, 99.95% per capability, error-budget burn, failed-payment volume, reconciliation breaks, repeat-incident rate.
- People: rotation size, frequency, after-hours pages, recovery days, sentiment, attrition among on-call staff.
- Cadence: weekly Incident Review Board; monthly Reliability Review; quarterly exec and board review; annual policy review.
- Error budgets on Tier 0/1 SLOs. Burn too fast and the team pauses features to pay down reliability.
- Reconcile dashboards monthly against a sample of incident records and customer cases so missing incidents cannot hide.
22. Wave rollout with readiness gates (depends on: 20, 21)
Roll out in four waves by criticality, every three weeks. Gates keep the standard credible. A missed gate is rescheduled, not waived.
Finish all 28 teams by week 24 so roughly three months of operating evidence remain before audit fieldwork.
- Wave 1: remaining Tier 0. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team kit: catalog complete, alerts migrated and inside budget, runbooks at the readiness bar, rotation of 6+ certified responders, one IC nominee, one drill passed, paid rotation in payroll.
- Named coach for three weeks. The director signs the gate.
- Freeze legacy paging per wave.
- Publish a live adoption scoreboard.
- Any production service that cannot provide sustainable ownership gets an executive-reviewed deadline, compensating control, and expiry date. No indefinite verbal exceptions.
23. Change management, incentives, and the deal (depends on: 1, 4, 9)
Run this from day one in parallel with design. Engineers will judge fairness. Executives will judge visible results.
Repeat the deal until it is muscle memory: paid on-call, paged only for what you own, trained commander, real sprint capacity for actions.
- Launch via CTO all-hands, per-team roadshows, a one-page laptop card, an internal wiki, and a Slack help channel with a 4-hour answer SLA.
- Put incident-response contribution into promotion criteria. Award best postmortem and biggest noise cut quarterly. Thank people publicly after every SEV1.
- Put adoption, alert hygiene, action closure, and on-call load fairness into every engineering manager's quarterly objectives.
- Write an exception path for engineers who cannot do nights because of caring responsibilities or health, covered by stipended volunteers.
- Prohibit retaliation for good-faith declaration or escalation.
- Pulse-survey at 60 and 120 days. If fairness or load is red, pause expansion until fixed.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
24. SOC 2 evidence, mock audit, and sustainability (depends on: 17, 21, 22)
Evidence is a by-product of doing the work, not a reconstruction the week before the auditor. Protect the process after SOC 2 is signed.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to CC7.2–CC7.4, CC2.2/CC2.3, CC5, and availability A1.2. Confirm interpretation early.
- Maintain versioned, signed policies for incident response, severity, on-call, communications, and postmortems, reviewed annually.
- Automate evidence: incident records, paging and ack logs, status history, postmortem library, action closure, training register, drill records, and reportability decisions including not reportable.
- Internal dry-run at month 6 on a 15-incident evidence walkthrough. Mock audit at month 7. Fix with weeks to spare.
- Assign permanent owners for policy, pager, status page, catalog, training, and metrics.
- Year-two roadmap: follow-the-sun cell, self-healing for the top recurring causes, error-budget release gates, and blast-radius reduction on the shared ledger.
- Quarterly board summary of severe incidents, credits, overdue P0s, and resilience investment so attention does not die after the audit.
Previous Proposal 5 (ID: 0ab2b1c8-46af-4001-ab14-0c074d30a026, Agent: deepseek-v4-pro_refine_5, LLM: deepseek/deepseek-v4-pro):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 6 months of full rollout.
- Customer-first detection falls from 40% to below 10% within 6 months and below 5% within 12.
- Median time to mitigate SEV0/SEV1 falls from 3h10 to under 60 minutes within 12 months; SEV2 under 2 hours.
- Incident Commander assigned and announced within 5 minutes for 95% of SEV0/SEV1 incidents; zero incidents with unclear command beyond 10 minutes.
- Status page updated within policy time for 95% of SEV0/SEV1/SEV2 (10/15/30 minutes).
- Monthly pages fall from 3,400 to under 500 with actionability above 75%.
- All six legacy alerting tools decommissioned by week 16.
- 100% of SEV0/SEV1 incidents have a blameless postmortem published within 10 business days.
- Postmortem action closure rises from 17% to 90% for P0/P1 actions on time.
- SLA credits fall from $1.3M to under $400K in first 12 months.
- Customer-impacting incidents decline to under 15 per year; repeat root causes under 10%.
- 100% of 180 services have a named owning team and criticality tier.
- All 28 teams onboarded by week 24; every Tier 0/1 team has 24x7 primary+secondary coverage with 6+ certified responders.
- At least 30 certified Incident Commanders and 20 certified Communications Leads active.
- Paid on-call policy is approved and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among on-call engineers.
- Quarterly game days and twice-yearly unannounced paging drills executed on schedule, each with tracked action items.
Steps (26):
1. Executive mandate, program governance, and interim incident command
Convert the CEO's outage complaint into a company-level improvement program with a named owner, budget, and deadlines. In the first week, establish an interim command process so no incident remains unowned while the permanent process is designed.
- Appoint Head of Reliability as program owner and CTO/CEO as executive sponsor.
- Form steering group with Engineering, SRE, Support, CS, Legal, Compliance, HR, Finance, Security.
- Approve non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Fund tooling, on-call compensation, training, and 3 dedicated program FTEs; anchor to $1.3M credits.
- Set timeline: design weeks 1-4, tooling and pilot weeks 5-10, rollout weeks 11-20, audit evidence collection from week 8, dry-run month 6.
- Stand up interim 24x7 duty officer and single declaration path within 48 hours, compensated retroactively.
2. Baseline data and alert estate analysis (depends on: 1)
Re-open the last 31 incidents and profile the current alert estate so every later design decision is evidence-based.
- Re-code each incident: detection source, timestamps, owner, severity, credits, root-cause family.
- Quantify customer-first detections and missing signals.
- Analyze two 'nobody in charge' incidents minute-by-minute.
- Inventory six alerting tools: volume, noise, owner, runbook coverage, top 50 noisy rules.
- Freeze baseline metrics: MTTD 22m, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed.
3. Stakeholder listening and resistance mapping (depends on: 1, 2)
Treat engineer pushback as design input. Interview all 28 teams plus Support, CS, and Sales to understand real objections and find early adopters.
- Test objection: unpaid work, night load, unfamiliar code, poor runbooks, or blame.
- Document current informal practices from 12 on-call teams.
- Recruit 10-15 credible engineers as co-design working group.
- Survey baseline sentiment on on-call and alerts.
- Publish promise: paid on-call, page only for owned services, trained commander, reserved reliability capacity.
4. Service ownership catalog and criticality tiering (depends on: 2, 3)
Build a machine-readable service catalog that assigns each of the 180 services a single owning team, escalation policy, and criticality tier.
- Define fields: owner team, manager, Slack channel, escalation policy, dependencies, dashboards, runbooks.
- Tier 0: ledger, money movement, auth, shared PostgreSQL; Tier 1 customer-facing; Tier 2 internal; Tier 3 non-critical.
- Map customer-visible capabilities and dependencies across regions.
- Identify orphan and shared services; force ownership decision within 30 days or schedule decommissioning.
- Publish coverage gaps by tier; Tier 0 gaps become executive escalations.
5. Define severity levels and declaration triggers (depends on: 4)
Adopt a five-level severity model with objective payment-specific triggers, so declaration is a lookup rather than debate. Anyone may declare; only the Incident Commander may downgrade.
- SEV0: unauthorized, lost, duplicated, or corrupted money movement; ledger integrity loss; confirmed data breach; both regions failed. Page all roles, exec, legal, and consider payment pause.
- SEV1: widespread payment failure, no workaround, severe SLA breach, or single-region total loss. Full role activation and public status page.
- SEV2: significant degradation, multiple customers, or workaround available but SLA risk. IC and SMEs paged; status page if customer-visible.
- SEV3: limited impact with workaround; team-led, business-hours response.
- SEV4: internal/no customer impact; ticket only.
- Auto-escalation: unresolved SEV3 >2h becomes SEV2; unresolved SEV2 >1h becomes SEV1; ledger or security always at least SEV2.
- Provide decision tree and 12 worked examples from actual incidents.
6. Define incident roles, decision authority, and handover rules (depends on: 5)
Codify five roles with written responsibilities and explicit authority, so there is never confusion about who is in charge.
- Incident Commander: owns severity, priorities, cross-team pulls, mitigation decisions; does not code.
- Communications Lead: owns status page, internal updates, account manager briefs, regulator coordination.
- Scribe: maintains timeline, decisions, and evidence for postmortem/audit.
- Subject-Matter Responders: diagnose and remediate only their owned services.
- Executive Liaison (SEV0/SEV1): handles exec communications and external stakeholders.
- Rule: IC identified within 5 minutes and announced in channel; handover announced and logged; roles may combine below SEV2, never at SEV0/SEV1.
7. Design 24x7 incident command and comms staffing model (depends on: 6, 4)
Create a central, trained incident command rotation instead of making each of 28 teams field its own commander.
- Recruit 24-30 certified ICs and 12-16 Comms Leads from a cross-team volunteer pool with manager approval.
- Weekly rotations, primary and secondary; 5-minute acknowledgement SLA with auto-escalation.
- Scribe pool as entry-level rotation.
- Night coverage: paid command rotation in New York timezone initially; evaluate follow-the-sun coverage later.
- Eligibility: certification required; commanders can leave with 30 days notice.
- Ensure distinct persons for IC, CL, and primary SME on SEV0/SEV1.
8. Design team on-call rotations and cross-team escalation policies (depends on: 4, 6)
Define tiered team on-call obligations so engineers are only paged for services they own, and cross-team pages go through the IC.
- Tier 0/1 services: 24x7 primary+secondary, at least 6 trained responders per rotation, one-week shifts.
- Tier 2/3: business-hours on-call; after-hours escalation list to engineering manager.
- Platform, infrastructure, database: 24x7 due shared ledger and Kubernetes.
- Routing: every page resolves via service catalog to owning team's escalation policy; cross-team pages only by IC.
- Guardrails: max one primary week in four, no consecutive weeks, no on-call in week after leading SEV0/1, protected recovery time after night work.
9. Design paid on-call compensation and fatigue safeguards (depends on: 8, 3)
Make on-call paid and legally compliant before any new rotation starts, and convert unpaid pager culture into a fair employment condition.
- Weekly stipends: primary 24x7 $800-1,200, secondary 30-50%, business-hours $300-500; holidays premium.
- Out-of-hours incident pay: $150 per night page plus hourly beyond one hour; time-off-in-lieu after overnight work.
- Additional command rotation stipend and SEV0/1 bonus for active responders.
- Verify FLSA/NY wage rules with HR, Legal, Finance; document exemption treatment.
- Base annual budget against current $1.3M credits.
- Track load and trigger staffing review if >2 after-hours pages per person per week sustained.
10. Enforce alert quality standards and noise budget (depends on: 2, 4, 8)
Replace 3,400 monthly alerts at 85% noise with a contractual paging standard that makes on-call sustainable.
- Every paging alert must have owning team, customer impact statement, runbook link, severity mapping, and tested threshold.
- Page only on customer-visible symptoms or SLO burn; cause-based alerts become tickets or dashboards.
- Set page budget: max 2 out-of-hours pages per person per week; breach triggers mandatory alert-tuning sprint.
- Auto-quarantine alerts with >5 firings/month without action or >70% no-action acknowledgements.
- Target <500 actionable pages/month and >75% actionability within 6 months.
- Weekly per-team alert review, monthly cross-team review.
11. Build detection uplift: synthetics, SLOs, and support intake (depends on: 4, 10)
Shift detection from host metrics to customer outcomes so the company stops hearing about outages from clients first.
- Define SLOs per Tier 0/1 capability: payment initiation, auth, settlement timeliness, API availability, ledger consistency.
- Deploy external synthetic transactions from both regions every 60 seconds, covering full payment flow and ledger write.
- Add ledger assurance checks: replication lag, double-entry balance, settlement window countdown.
- Top-100 customer anomaly detection to catch single-tenant outages.
- Auto-create triage incident from support tickets or account manager keywords within 5 minutes.
- Track customer-detected-first as a defect and require a postmortem action.
12. Consolidate alerting and incident tooling (depends on: 5, 8, 10, 11)
Collapse six alerting tools into one integrated paging and incident management platform to create a single system of record for people and audit.
- Select paging/on-call platform and incident management layer (e.g., PagerDuty + incident.io/FireHydrant).
- Implement one-command Slack declaration that auto-creates channel, bridge, pages roles, sets severity, starts timeline.
- Migrate all monitoring sources into the one tool; decommission legacy paging only after two weeks verified.
- Integrate service catalog, status page, Jira action tracking, Salesforce/CS customer lists, and conference bridge.
- Ensure out-of-band paging and offline fallback if a region or chat tool is down.
- Automate evidence capture for SOC2: timestamps, role assignments, severity changes, comms sent.
13. Define acknowledgement and escalation paths (depends on: 12, 6, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts at first signal.
- For SEV0/SEV1: page primary on-call; after 5 minutes unacked page secondary; at 10 page manager and duty IC; at 15 page executive officer.
- For SEV2: primary ack within 10 minutes, IC assigned within 15; escalate on miss.
- For SEV3: ack within 30 minutes or create work item.
- Automatic duty IC page for ledger/security, cross-team, or unresolved ownership.
- If impact unknown after 15 minutes, raise severity.
- Cross-team responders summoned by IC have 10-minute acknowledgement obligation.
- Route alerts with no owner to command rotation, then treat missing ownership as control defect.
- Human acknowledgement required; delivery confirmation is not sufficient.
14. Standardize internal communications (depends on: 6, 5)
Separate the working war room from executive and stakeholder updates, with fixed cadence and pre-approved templates.
- Auto-create incident channel and one read-only broadcast channel for execs, support, sales.
- SEV0: internal update every 15 minutes; SEV1 every 30; SEV2 every 60; SEV3 on state change.
- Template: impact, what we know, what we are doing, ETA/next update, current IC and CL.
- Executives questions through Executive Liaison only; IC is not interrupted.
- Support/CS receives affected-customer list and holding statement within 15 minutes for SEV0/SEV1 and 30 for SEV2.
- Handover protocol for incidents lasting >4 hours: formal IC handover and fatigue check.
15. Standardize customer and status page communications (depends on: 5, 14)
Replace ad-hoc status page updates with a timed, role-owned, template-driven process, including account manager outreach.
- Status page timing: SEV0 initial post within 10 minutes, SEV1 within 15, SEV2 within 30; updates every 15/30/60 minutes until resolved.
- Resolution notice within 30 minutes of mitigation; customer-facing summary in 5 business days for SEV0/SEV1.
- Pre-approve 12-15 templates with Legal and Comms.
- Tiered outreach: top 100 accounts get direct call/email from AM within 30 minutes of SEV0/SEV1; long tail via subscription.
- Use factual language: state impact and next update; never speculate cause or blame vendor.
- Comms Lead is sole author for customer language.
16. Regulatory, legal, and account manager notification playbook (depends on: 5, 15)
Build a notification decision tree and contact matrix so legal/regulatory obligations are assessed early and never forgotten.
- Map obligations: NYDFS Part 500 72-hour cybersecurity event notification, state breach laws, GLBA/FTC, PCI, sponsor bank/card network contractual windows, FinCEN/OFAC if relevant.
- Add regulatory assessment checkpoint for every SEV0 and security SEV1 within 2 hours, even if not reportable.
- Maintain 24x7 contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel with named backups.
- Pre-draft notification templates and test quarterly.
- Encode customer-specific notification SLAs from enterprise contracts into customer tiering.
- Account managers receive legal-approved script and affected-customer list.
17. SLA credit workflow and financial impact measurement (depends on: 5, 15)
Link every incident to money automatically so severity, credits, and prioritization stay consistent and Finance is never surprised.
- Define availability measurement per contract and capability with Legal/Finance.
- Auto-compute affected minutes per customer from incident record and telemetry; generate credit proposal within 5 business days.
- Decide proactive credits for top tier vs claims-based for others; document approval chain.
- Track credits by root-cause family and incident; feed quarterly reliability investment decisions.
- Target reducing annual credits from $1.3M to under $400K in first year.
18. Standardize blameless postmortems (depends on: 5, 6)
Make postmortems mandatory with fixed deadlines, a single format, and a blameless review forum, replacing 'various formats'.
- Mandatory: all SEV0/SEV1; SEV2 with customer impact, credits, >2h, repeat cause, or detected by customer; near-miss involving ledger.
- Draft within 3 business days, peer review within 5, publish within 10.
- One template: timeline, customer/financial impact, detection gap, response gap, contributing factors, what went well.
- Blameless rules: context-based, no individual blame, never used in performance reviews.
- Weekly Incident Review Board reviews all postmortems and challenges quality.
- Publish searchable postmortem library and quarterly recurring causes report.
19. Track postmortem actions with owner and due date (depends on: 18, 12)
Fix the 11/64 completion rate by giving every action the same status as customer commitments, with capacity and escalation.
- Every action gets named owner, priority, due date, and Jira ticket auto-created from postmortem.
- P0 actions prevent SEV0 recurrence, due 30 days; P1 60; P2 90.
- Reserve 15-20% team sprint capacity for incident actions.
- Escalation ladder: manager at 7 days overdue, director at 14, CTO dashboard at 30; P0 overdue blocks features.
- Monthly reporting of closure rate in engineering leadership.
- Target 90% closure of P0/P1 on time in two quarters.
20. Create runbooks and service readiness bar (depends on: 4, 8, 11)
Ensure every service is prepared for 3 a.m. response before it is allowed to page anyone.
- Readiness checklist for Tier 0/1: architecture diagram, dependencies, dashboards, rollback, feature flags, escalation contacts, data-loss statement.
- Write major-incident playbooks: shared PostgreSQL failure, cross-region failover, Kubernetes control plane loss, partner outage, settlement breach, security compromise.
- Prioritize ledger playbooks: documented failover, read-only mode, reconciliation, signed RPO/RTO.
- Runbooks must be tested twice yearly; stale runbooks marked in catalog.
- No paging alerts without readiness sign-off; gap reported to director.
21. Train, certify, and simulate incident response (depends on: 6, 13, 14, 15, 18, 20)
Build a certification path so 24x7 roles are staffed by people who have practiced, and validate the process through drills.
- Scribe course (2 hrs); Responder (half-day); Communications Lead (1 day); Incident Commander (2 days + shadowing/tabletop).
- Certify ICs and CLs; recertify annually.
- New responders shadow two shifts before primary; no primary in first 90 days.
- Monthly tabletop using real past incidents; quarterly game day including regional failover or ledger scenario.
- Twice-yearly unannounced paging drill; one regulator/legal exercise annually.
- Track drill metrics and action items.
22. Pilot on critical services and iterate (depends on: 12, 13, 15, 19, 21)
Prove the process on the highest-risk surface before full rollout. Run a six-week pilot with tight measurement and public results.
- Select 5-6 teams: payments core, ledger/database, platform/Kubernetes, API gateway, plus two existing on-call teams.
- Activate full stack: severity, roles, command rotation, paid on-call, single tooling, alert budget, status page, postmortems.
- Weekly pilot retro; fix defects within 48 hours.
- Exit criteria: MTTD <10 min, IC assigned within 5 min in 95% incidents, pages down 50%, postmortems on time, positive sentiment.
- Publish one-page pilot result as adoption argument.
23. Wave rollout to all teams and decommission legacy paths (depends on: 22, 20)
Roll out to all 28 teams in four waves by criticality, with explicit readiness gates and legacy tool shutdown.
- Wave 1 remaining Tier 0; Wave 2 Tier 1; Wave 3 Tier 2; Wave 4 Tier 3/internal.
- Per-team onboarding: catalog complete, alerts pruned, runbooks ready, rotation staffed with 6+ certified responders, one IC candidate, one drill passed.
- Coach per wave 3 weeks.
- Gate signed by director; failures rescheduled, not waived.
- After onboarding, freeze legacy alerting paths; no fallback to old tools.
- Publish adoption scoreboard.
24. Establish metrics, dashboards, and review cadence (depends on: 12, 19, 22)
Make process performance visible with a small set of metrics and fixed review meetings, so the program is owned by data.
- Response: MTTD, time to IC, MTTA, MTTM, MTTR, % customer-detected-first.
- Quality: page volume per person, alert actionability, postmortem on-time, action closure rate.
- Business: SLA credits, availability vs 99.95, repeat incidents.
- People: on-call load, page per engineer, sentiment, attrition.
- Cadence: weekly Incident Review Board, monthly Reliability Review, quarterly Executive/Board review.
- Dashboards self-serve with targets and named owners.
25. SOC 2 readiness and internal dry-run audit (depends on: 19, 23, 24)
Design evidence as a by-product and test it with an internal walkthrough before the external auditor arrives.
- Map process to SOC2 CC7.3/CC7.4, CC7.2, CC2.2/2.3, CC5, availability criteria.
- Publish versioned policy documents: Incident Response, Severity, On-Call, Communication, Postmortem.
- Automate evidence: incident records with timestamps, role assignments, paging logs, status history, postmortems, action board, training register, drill records.
- Ensure process operates at least 3 months before fieldwork.
- Run internal dry-run month 6; sample 15 incidents; fix gaps with 8 weeks to spare.
- Keep remediation log for process deviations.
26. Continuous improvement, culture, and sustainability (depends on: 23, 25)
Prevent the process from decaying after the audit by embedding review, feedback, and roadmap ownership.
- Quarterly process retrospective with IC pool and responders.
- Re-baseline metrics every six months; raise targets.
- Year-two roadmap: follow-the-sun coverage, self-healing top 3 causes, error budgets gating releases, blast-radius reduction for shared ledger.
- Annual policy review and certification renewal.
- Quarterly on-call sentiment survey with published actions.
- Board quarterly report on availability, credits, and incident trends.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Establish executive mandate, program office, and funding", "description": "Turn the CEO email into a **chartered company program** with one accountable owner and authority over all 28 teams.\n\n- Appoint the CTO as executive sponsor and a Head of Reliability as program owner with full-time authority.\n- Create a permanent program office: one program lead, one platform engineer, one analyst.\n- Form an eight-person steering group: Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours.\n- Lock the non-negotiables: one severity scale, one paging platform, one postmortem format, paid on-call, mandatory action tracking, named service ownership.\n- Approve budget anchored against the $1.3M in credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs.\n- Reserve 15% of engineering capacity for reliability and incident actions, protected by the sponsor.\n- Publish a one-page charter stating incident response is a company operating process, not a per-team choice.", "dependencies": []}, {"step_id": "S2", "title": "Install a seven-day interim command bridge", "description": "Do not leave the company unprotected while the permanent process is designed. Put a **crude but real** command structure in place within one week.\n\n- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.\n- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs. Compensate retroactively under the final policy.\n- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.\n- Direct Support to escalate credible customer reports immediately without waiting for engineering confirmation.\n- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.\n- Triage all 53 open historical actions: complete, re-plan, or formally risk-accept, prioritizing ledger integrity, duplicate payments, regional failover, and detection gaps.\n- Hold a 15-minute daily operations review until the permanent process is live.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Build the forensic baseline of incidents, alerts, and money lost", "description": "Rebuild the facts before designing anything. This becomes both the **design input** and the before picture for the executive and the auditor.\n\n- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.\n- For each of the 40% customer-first detections, name the specific missing signal. This list drives the detection backlog.\n- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.\n- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.\n- Freeze baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Run the listening tour and publish the on-call fairness deal", "description": "Engineer pushback is the **largest delivery risk**. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.\n\n- Interview all 28 teams plus Support, CS, and Sales in two weeks. Separate the real objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, fear of blame. Each has a different remedy.\n- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.\n- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.\n- Recruit 12–15 credible engineers as a design working group so the process is co-authored.\n- Run a baseline sentiment survey covering fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.", "dependencies": ["S1"]}, {"step_id": "S5", "title": "Map compliance, evidence, and audit requirements from day one", "description": "Design evidence as a **by-product of operations**, not a reconstruction before the auditor arrives. Confirm the SOC 2 observation window immediately.\n\n- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for communications; A1.2 for availability.\n- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure, training register, drill records, reportability decisions.\n- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.\n- Version and approve all policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard.\n- Record every control exception with an owner, compensating control, approval, and expiry date.\n- Start monthly evidence sampling immediately rather than reconstructing before the audit.", "dependencies": ["S1"]}, {"step_id": "S6", "title": "Build the service ownership catalog and criticality tiers", "description": "You cannot route a page correctly across 180 services until every service has a **named owner**. Build a machine-readable catalog as the single source of truth.\n\n- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.\n- Tier by business impact: Tier 0 (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).\n- Map every customer journey to its services, data stores, regions, and third parties including sponsor banks and processors.\n- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for each Tier 0/1 service.\n- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.\n- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.\n- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.", "dependencies": ["S3"]}, {"step_id": "S7", "title": "Adopt the severity scale, declaration rules, and lifecycle", "description": "Replace judgment calls with a **lookup table**. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter.\n\n- SEV1 (crisis): money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of all roles and executive; bridge in 5 minutes; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.\n- SEV2 (critical): material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.\n- SEV3 (major, contained): narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.\n- SEV4: no customer impact; ticket only, never pages.\n- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.\n- Anyone may declare and nobody is penalised for over-declaring. Only the commander may downgrade, with evidence recorded.\n- Lifecycle: Detected, Declared, Triaged, Mitigated (customer impact ends), Monitoring, Resolved (backlog processed and ledger reconciled), Reviewed.\n- Publish a decision tree with 12 worked examples from the real 31 incidents.", "dependencies": ["S3", "S6"]}, {"step_id": "S8", "title": "Define incident roles, authority, and handover discipline", "description": "Solve nobody-in-charge-for-an-hour by making command **explicit, single-holder, transferable, and logged**. Separate coordination from debugging.\n\n- Incident Commander: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.\n- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.\n- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates for SEV1/SEV2.\n- Subject-Matter Responders: engineers of the owning team; they mitigate, they do not run the room.\n- Executive Duty Officer (SEV1): removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.\n- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.\n- The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.", "dependencies": ["S7"]}, {"step_id": "S9", "title": "Design the three-layer 24x7 coverage model", "description": "Do not create 28 night rotations. **Centralise coordination** in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.\n\n- Layer A, Incident Command corps: approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of approximately 18 from Support, CS, and engineering management. A scribe pool used as the training entry point.\n- Layer B, critical-path domain rotations: consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.\n- Layer C, everyone else: business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.\n- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.\n- Evaluate in writing: a US-only paid night rotation now, a follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.\n- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.", "dependencies": ["S6", "S8"]}, {"step_id": "S10", "title": "Approve on-call compensation, labor compliance, and fatigue safeguards", "description": "Unpaid on-call in New York is both a **retention problem and a legal exposure**. Pay must be live in payroll before any mandatory night rotation starts.\n\n- Indicative scheme: approximately $1,000 per primary 24x7 week, $400 secondary, $250 for business-hours rotations, a separate $1,200 Duty Commander stipend, holiday premiums, approximately $150 per out-of-hours page plus hourly beyond the first hour.\n- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.\n- HR, Finance, and employment counsel publish amounts, eligibility, tax treatment, and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.\n- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.\n- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.\n- Non-cash elements: on-call counted as delivery load with a 15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.", "dependencies": ["S9", "S4"]}, {"step_id": "S11", "title": "Set the alert quality standard, page budget, and noise burn-down", "description": "3,400 alerts at 85% noise is why detection takes 22 minutes. Make alert quality a **condition of being allowed to page a human**.\n\n- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.\n- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.\n- Run new alerts in shadow mode for seven days; test both firing and recovery.\n- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.\n- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.\n- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly.\n- Target 3,400 to under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.", "dependencies": ["S3", "S6"]}, {"step_id": "S12", "title": "Build payment-outcome detection and ledger assurance", "description": "Stop customers telling you first. Detection must be driven by **payment outcomes and ledger truth**, not host metrics.\n\n- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.\n- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.\n- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.\n- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls. Five similar support tickets in ten minutes auto-creates a triage incident.\n- Route credible partner and processor notifications into the same declaration path within five minutes.\n- Record detection source on every incident. Customer detected first becomes a named defect class with a mandatory tracked action.\n- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.", "dependencies": ["S11", "S6"]}, {"step_id": "S13", "title": "Consolidate to one paging platform, one incident record, one status page", "description": "Collapse six alerting tools into a **single operational system of record** so there is one queue, one timeline, and one audit trail.\n\n- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.\n- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.\n- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.\n- Route every page through the service catalog: service label to owning domain schedule to escalation policy.\n- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.\n- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is unavailable. Test paging, escalation, status publication, and bridge access weekly.\n- Set a hard date after which pages outside this tool create no on-call obligation.", "dependencies": ["S7", "S9", "S11"]}, {"step_id": "S14", "title": "Codify the escalation ladder and five-minute command rule", "description": "Write one unskippable path from something looks wrong to **someone is in charge**. The default action is never waiting.\n\n- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.\n- SEV1/SEV2 ladder: owning primary immediately, secondary at 5 minutes, domain manager and Duty Commander at 10, Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.\n- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.\n- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.\n- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.\n- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension, and partner-bank notification, so the commander never waits for an executive.\n- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.", "dependencies": ["S8", "S9", "S13"]}, {"step_id": "S15", "title": "Write the live incident execution doctrine and major-incident playbooks", "description": "Give responders one short operating procedure from the first minutes through closure. Priority is **limiting customer and financial harm**, not proving a root cause.\n\n- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.\n- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.\n- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute before deep diagnosis.\n- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.\n- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.\n- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.\n- Treat the shared ledger cluster as the single largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.\n- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.\n- Closure requires a stability observation window, an explicit handback to the owning team, Support, and Customer Success, and reopening if impact recurs.\n- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.", "dependencies": ["S8", "S13", "S14"]}, {"step_id": "S16", "title": "Standardize internal communications protocol", "description": "Standardise the internal picture so executives, Support, and Sales are informed **without interrupting** the person running the incident.\n\n- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, and Sales.\n- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.\n- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.\n- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.\n- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.\n- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.", "dependencies": ["S8", "S13"]}, {"step_id": "S17", "title": "Build customer communications, status page, and account-manager outreach", "description": "Replace whoever is around with a **timed, owned, pre-approved process**. The Comms Lead is the single author and never writes from scratch under pressure.\n\n- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.\n- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.\n- Component-level status page mapped to customer journeys, with subscriptions, email, webhook, and RSS. All 2,100 customers subscribed by default.\n- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email.\n- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity, or blame.\n- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.", "dependencies": ["S7", "S16"]}, {"step_id": "S18", "title": "Create the regulatory, partner, and legal notification playbook", "description": "In payments, some incidents start a **legal clock at detection**. Build the assessment into the process so it is never remembered late.\n\n- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.\n- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded even when the answer is not reportable, with evidence, decision-maker, and timestamp.\n- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.\n- Pre-draft notification letters held under privilege review. Legal owns outbound regulatory text; the commander owns the facts.\n- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.\n- Rehearse the playbook once a quarter as part of the exercise programme.", "dependencies": ["S7", "S17"]}, {"step_id": "S19", "title": "Operationalize SLA credit and financial impact workflow", "description": "Tie incidents to money so severity, credits, and investment decisions **stay honest**, and Finance stops being surprised.\n\n- Agree with Legal and Finance the availability measurement method per contract and per component.\n- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.\n- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.\n- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.\n- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.\n- Target: credits from $1.3M to under $400k in year one. Use the delta as the standing business case for on-call pay and reliability capacity.", "dependencies": ["S7", "S17"]}, {"step_id": "S20", "title": "Establish the blameless postmortem standard and Incident Review Board", "description": "Replace some incidents, various formats with **one mandatory format, fixed deadlines, and a forum with teeth**.\n\n- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.\n- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.\n- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.\n- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.\n- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.\n- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.\n- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly top-five recurring causes analysis.", "dependencies": ["S7", "S8"]}, {"step_id": "S21", "title": "Enforce action ownership, capacity reservation, and tracking", "description": "Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of **customer commitments** and give them real capacity.\n\n- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.\n- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.\n- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor.\n- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.\n- Verify effectiveness before closing. A closed ticket without evidence does not close the action.\n- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.\n- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.", "dependencies": ["S20", "S13"]}, {"step_id": "S22", "title": "Build the training and certification academy", "description": "Command is a skill, not a title. **Certify before assigning duty**, and use paid working time for all of it.\n\n- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.\n- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.\n- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.\n- Incident Commander (2 days + shadowing): command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.\n- Comms Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.\n- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum; never primary in the first 90 days.\n- Certification valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.", "dependencies": ["S8", "S14", "S16", "S20"]}, {"step_id": "S23", "title": "Run the exercise programme: tabletops, game days, and unannounced drills", "description": "The process must meet a **simulated SEV1** before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply.\n\n- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.\n- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.\n- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.\n- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.\n- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.\n- Never inject uncontrolled change into the production ledger. Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.\n- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-correct-mitigation.", "dependencies": ["S22", "S13", "S15"]}, {"step_id": "S24", "title": "Publish Incident Management Policy v1", "description": "Collapse the design into a document people will **actually open mid-outage**, and make it official.\n\n- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.\n- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.\n- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.\n- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.\n- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.", "dependencies": ["S7", "S8", "S9", "S10", "S11", "S14", "S16", "S17", "S18", "S20", "S21"]}, {"step_id": "S25", "title": "Pilot on the payments critical path", "description": "Prove the process on the **highest-risk surface** with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.\n\n- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.\n- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.\n- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.\n- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.\n- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.\n- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment.\n- Publish a one-page result to the whole company.", "dependencies": ["S24", "S12", "S13", "S15", "S22", "S10"]}, {"step_id": "S26", "title": "Run the alert noise burn-down campaign", "description": "Run noise reduction as a **visible, quota-driven campaign** in parallel with rollout, not as a background hope.\n\n- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.\n- Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.\n- Publish a weekly noise leaderboard that names systems, never people.\n- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.\n- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness. Review missed detections monthly with the same seriousness as noise.\n- Milestones: 3,400 to 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.", "dependencies": ["S11", "S13", "S25"]}, {"step_id": "S27", "title": "Establish metrics, dashboards, review cadence, and anti-gaming", "description": "Instrument the process itself so improvement is visible and the auditor sees **evidence of monitoring and review**. Report median and 90th percentile, never averages alone.\n\n- Response: time to detect, declare, commander, acknowledgement, mitigate, resolve, split by severity, tier, journey, region, and detection source.\n- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.\n- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.\n- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.\n- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.\n- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.\n- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.\n- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.", "dependencies": ["S13", "S20", "S25"]}, {"step_id": "S28", "title": "Wave rollout to all 28 teams with readiness gates", "description": "Roll out in four waves of six to eight teams every two to three weeks, ordered by **customer risk**. Each wave passes an explicit gate rather than a date.\n\n- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.\n- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.\n- Each wave gets a named coach from the program office for three weeks.\n- Gates are signed by the director. Failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.\n- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.\n- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.\n- Publish a live adoption scoreboard.", "dependencies": ["S25", "S27"]}, {"step_id": "S29", "title": "Run change management, fairness, and pager culture programme", "description": "Run this from day one in parallel. Engineers judge the process on **fairness**; executives judge it on visible results. Both need constant, honest communication.\n\n- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.\n- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.\n- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.\n- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.\n- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.\n- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.", "dependencies": ["S4", "S10", "S25"]}, {"step_id": "S30", "title": "Build SOC 2 evidence by design, internal testing, and mock audit", "description": "Design the evidence as a **by-product of doing the work**. Never build a parallel audit process; the real process is the evidence.\n\n- Map the process with Compliance and the auditor to the Trust Services Criteria and confirm the observation window early.\n- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews.\n- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.\n- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.\n- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.\n- Freeze process wording for the remainder of the observation window once month 7 closes. Log any change as an exception.", "dependencies": ["S24", "S27", "S28"]}, {"step_id": "S31", "title": "Maintain the program risk register and contingencies", "description": "Name the ways this programme fails and **pre-commit the response**. Review it monthly with the sponsor.\n\n- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.\n- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.\n- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.\n- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.\n- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.\n- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.\n- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.", "dependencies": ["S1"]}, {"step_id": "S32", "title": "Ninety-day inspect-and-adapt, then year-two sustainability", "description": "Guard against the classic failure: the process **decays once the audit is signed**. Revise on data, then build the second-year plan before the first year ends.\n\n- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.\n- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.\n- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.\n- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.\n- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.\n- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.", "dependencies": ["S27", "S28", "S30"]}], "estimated_complexity": "high", "success_metrics": "- A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.\n- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.\n- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.\n- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.\n- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.\n- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.\n- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.\n- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.\n- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.\n- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.\n- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.\n- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.\n- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.\n- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.\n- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.\n- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.\n- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.\n- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.\n- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.\n- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.\n- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders."}Compressed to 26 steps while adding the round's neatest severity idea and the only staffing arithmetic. Gains: integrity/security flag (step 6), 260-engineer rotation math (step 8), week-1 audit clock (step 1), postmortems split from action tracking. Losses: the change-management step disappeared and the risk register is demoted to the end.
- Step 6's financial-integrity or security flag attachable to any severity delivers SEV0's forcing function — dual control, Legal, regulatory checkpoint — without inventing a fifth level.
- Step 8 justifies the coverage model numerically: 260 engineers can sustain 10–12 domain rotations plus one command corps, not 28 night rotas; no two-person 24x7 rotation, ever.
- Step 1 confirms the SOC 2 observation window with the auditor in week 1 so the seven-day operating floor already counts as evidence.
- Postmortem standard (17) and action tracking (18) separated, where round 1 fused them; reportability checkpoint tightened from two hours to one (step 15).
- Noise burn-down runs as a per-wave quota inside step 24 rather than a parallel campaign — one fewer concurrent workstream during rollout.
- SEV3 postmortem trigger tightened to customer-detected, over two hours, repeat, or credit-generating.
- The round-1 change-management step is gone: no office hours, no four-hour help-channel SLA, no awards, no non-retaliation clause, and no public closure of the two "nobody in charge" incidents — only the fairness contract (4) and a per-wave reminder (24) remain.
- The risk register is folded into step 26, which depends on steps 23–25, so contingencies only exist after rollout and mock audit, and drop from seven pre-committed responses to four.
- Step 15 is an 11-bullet mega-step spanning internal cadence, status page, account-manager outreach and NYDFS/PCI/FinCEN obligations — too much for one operating card.
- Step 24 still promises completion "at least five months before the audit report date" against a week-24 finish and a month-8 audit.
- Proposal 2 : Confirm the auditor's expected Type II observation period immediately.
- Proposal 2 : A distinct SEV0 tier for financial-integrity and security crises.
- Proposal 1 : Separate postmortem standard and action-ownership steps.
- Proposal 1 : One-hour reportability assessment for SEV1 and security SEV2, recorded even when not reportable.
- Proposal 1 : Pre-committed contingencies for volunteer shortfall, pay delay and a SEV1 mid-rollout.
- Proposal 1 : Three separate steps for internal, customer and regulatory communications.
- Proposal 1 : A standalone quota-driven noise burn-down campaign parallel to rollout.
- Proposal 2 : Do not merge small teams merely for schedule convenience.
+ Charter the program and start the audit clock+ Install a seven-day operating floor+ Rebuild the forensic baseline+ Map resistance and publish the fairness contract+ Lock severity levels and what each one triggers+ Staff 24x7 with a command corps, not 28 night rotations+ Pay on-call, meet New York labor rules, and cap fatigue+ Detect payment and ledger failures before customers do+ Write the five-minute path from signal to command+ Codify live execution, readiness, and ledger playbooks+ Run one communications clock for internals, customers, and regulators+ Tie incidents to SLA credits and true financial cost+ Track actions as risk commitments with reserved capacity+ Train and certify every role before independent duty+ Publish Policy v1 and the signed fairness contract+ Roll out in risk-ordered waves with readiness gates+ Inspect at 90 days and lock year-two ownershipExecutive mandate, single owner, funding and non-negotiablesSeven-day interim command bridgeForensic baseline of incidents, alerts and money lostListening tour, resistance map and the on-call dealSeverity scale, declaration rules and incident lifecycle24x7 coverage model: central command corps, local expertiseOn-call compensation, labour compliance and fatigue safeguardsDetection uplift on the money pathEscalation ladder and the five-minute command ruleLive incident execution doctrineInternal communications protocolCustomer communications and status page policyRegulatory, partner and legal notification playbookSLA credit and financial impact workflowAction ownership, capacity reservation and enforcementRunbooks, readiness bar and ledger blast-radius reductionTraining and certification academyPublish Incident Management Policy v1Wave rollout to all 28 teams with readiness gatesAlert noise burn-down campaignChange management, fairness and pager cultureRisk register and contingenciesNinety-day inspect-and-adapt, then year-two sustainability
The plan produced
1. Charter the program and start the audit clock new
Convert the CEO email into a company operating process with one owner, a budget, and an observation window that starts this week.
Incident response is no longer a per-team choice.
- Name the CTO as sponsor and a Head of Reliability as the single accountable owner, with a three-person program office.
- Form a small decision group: Engineering, SRE, Support, Customer Success, Security, Legal, Finance, and HR. The sponsor decides within 48 hours.
- Lock non-negotiables: one severity scale, one paging path, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, and reserved engineering capacity against the $1.3M in SLA credits.
- Timeline: operating floor in 7 days, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, mock audit in month 7, audit in month 8.
2. Install a seven-day operating floor (after 1)
Do not wait for new tools or the final policy. Put a crude but real process in place this week so the next outage has a named commander.
This week is the first audit evidence.
- Publish a one-page interim severity card and one declaration path: Slack command, phone number, and existing pagers, all reaching the same duty person.
- Staff interim primary and backup Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in the incident channel.
- Tell Support to declare from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline (after 1) new
Rebuild the facts before locking design. This is both the design input and the before picture for the CEO and the auditor.
Freeze the numbers so they cannot drift during design.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisy rules and every rule with no owner or runbook.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11 of 64 actions closed, on-call in 12 of 28 teams.
4. Map resistance and publish the fairness contract (after 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
Publish the deal in writing before any new mandatory pager is assigned.
- Interview all 28 teams plus Support, Customer Success, and Sales in two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the fairness contract: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and money-path tiers (after 3) from P1 step 5
You cannot page the right person across 180 services until every service has one owner. A wrong owner recreates the pager objection.
The catalog is the single source of truth for paging, impact, status-page components, and audit.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. Tier 0: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to services, data stores, regions, and third parties, including sponsor banks and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
6. Lock severity levels and what each one triggers (after 3, 5) from P5 step 5
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by who reported it or how hard the fix looks.
Anyone may declare. Nobody is punished for over-declaring. Only the Incident Commander may downgrade, with the evidence recorded.
- SEV1 crisis: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 major, contained: narrow or single-customer impact with a workaround. Owning team leads; business-hours comms; postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages.
- Attach a financial-integrity or security flag to any severity. The flag forces dual-control, Legal, and the regulatory checkpoint without inventing a fifth level.
- Auto-escalate: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer impact ends) → Monitoring → Resolved (backlog processed and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (after 6) from P2 step 6
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged.
Separate coordination from debugging so a commander never needs to know another team's code.
- Incident Commander: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. A human validates for SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1 and SEV2. Every handover is announced verbally and in writing with the exact time.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
8. Staff 24x7 with a command corps, not 28 night rotations (after 5, 7) new
Do not create 28 night rotations. That is what engineers are rejecting.
Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified volunteers from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to seven months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate 28 teams into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat Lisbon or APAC follow-the-sun as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
9. Pay on-call, meet New York labor rules, and cap fatigue (after 4, 8)
Unpaid on-call in New York is a retention problem and a legal exposure. Pay must be in payroll before any new mandatory night rotation starts.
Publish the numbers. Then ask people to sign up.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
10. Set the alert-quality bar and a hard page budget (after 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them.
Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers do (after 5, 10) from P2 step 12
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
12. Consolidate to one pager, one incident record, one status page (after 6, 8, 10) from P1 step 12
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail.
Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One Slack command creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is down. Test weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Write the five-minute path from signal to command (after 7, 8, 12) new
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting.
If nobody claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Codify live execution, readiness, and ledger playbooks (after 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving a root cause.
A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down. No command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers.
15. Run one communications clock for internals, customers, and regulators (after 6, 7, 12) new
Replace whoever is around with one timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is not reportable.
- Obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
16. Tie incidents to SLA credits and true financial cost (after 6, 15) new
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
Credit calculation is an output of the incident record, not a negotiation.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
17. Make blameless postmortems mandatory and consistent (after 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth.
The discipline lives in the deadlines and the review, not the template.
- Mandatory for every SEV1 and SEV2, any incident detected by a customer first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed.
18. Track actions as risk commitments with reserved capacity (after 12, 17) new
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
Closing a ticket without evidence of effectiveness does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Target 90% of high-priority actions closed by due date within two quarters.
19. Train and certify every role before independent duty (after 7, 13, 15, 17) new
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
The certification register is an audit artefact.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
20. Rehearse with tabletops, game days, and night drills (after 12, 14, 19) from P1 step 23
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
21. Publish Policy v1 and the signed fairness contract (after 6, 7, 8, 9, 10, 13, 15, 17, 18)
Collapse the design into a document people will actually open mid-outage, and make it official.
Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
22. Pilot on the payments critical path (after 9, 11, 12, 14, 19, 21) from P1 step 25
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse. Publish a one-page result to the whole company.
23. Instrument metrics, reviews, and anti-gaming (after 12, 17, 22) from P1 step 26
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review.
Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
24. Roll out in risk-ordered waves with readiness gates (after 22, 23) new
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave, not as a background hope. Give every team its noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
25. Produce SOC 2 evidence by operating, then mock-audit (after 21, 23, 24) from P1 step 30
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions. Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (after 23, 24, 25) new
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
Name the ways this programme fails and pre-commit the response.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
- A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 3efb9ff1-28a5-40aa-8f73-9853e91aa095, Agent: opus5_refine_1, LLM: anthropic/claude-opus-5):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (32):
1. Executive mandate, single owner, funding and non-negotiables
Convert the CEO email into a chartered program with one accountable owner and authority over all 28 teams. Incident response becomes a **company operating process**, not a per-team choice.
- Name an executive sponsor (CTO) and one accountable owner (Head of Reliability / Incident Management) with a small permanent office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group with Engineering, SRE/Platform, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours. It is not a 28-person committee.
- Fix the non-negotiables now: one severity scale, one paging tool, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.
- Set the clock deliberately earlier than the audit: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready week 28.
- Approve budget against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises.
2. Seven-day interim command bridge (depends on: 1)
Do not let design work leave the company unprotected for six weeks. Put a crude but real process in place within seven days and improve it later.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster from engineering managers and senior SREs, primary plus backup, 24x7. Pay it retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Tell Support to escalate credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items: complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This becomes both the design input and the "before" picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start / detect / declare / commander-assigned / mitigate / resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. This list drives the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome. Identify the top 50 rules producing most of the noise and every rule with no owner or runbook.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS and Sales in two weeks. Test what the objection actually is: unpaid work, lost sleep, unfamiliar code, missing runbooks, or fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey (fairness, trust in alerts, willingness, burnout) to re-measure at 60, 120 and 365 days.
5. Service catalog, ownership and money-path tiering (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that is the single source of truth for paging, impact and audit.
- One accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not by technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions and third parties, including sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Every orphan service gets an owner in 30 days or a decommission date approved by the sponsor. Tier 0 without an owner is an executive escalation.
6. Severity scale, declaration rules and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- **SEV3 (major, contained):** narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- **SEV4:** no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring; only the commander may downgrade, with the evidence recorded.
- Lifecycle: Detected → Declared → Triaged → **Mitigated** (customer impact ends) → Monitoring → **Resolved** (backlog processed and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Incident roles, authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, roles, cadence and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit spend. Retains command when a VP joins.
- **Communications Lead:** single voice for status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates for SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team; they mitigate, they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance and Vendor Management join on defined triggers.
- Rules: command claimed within 5 minutes, stated in channel ("I am IC"), distinct people for command, comms and technical lead at SEV1/SEV2, and every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation or privileged-access rules.
8. 24x7 coverage model: central command corps, local expertise (depends on: 5, 7)
Do not create 28 night rotations. Centralise coordination in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- **Layer A — Incident Command corps:** ~30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of ~18 from Support, CS and engineering management; a scribe pool used as the training entry point.
- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- **Layer C — everyone else:** business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing, with cost and timeline: a US-only paid night rotation now, a Lisbon or APAC follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
9. On-call compensation, labour compliance and fatigue safeguards (depends on: 8, 4)
Unpaid on-call in New York is both a retention problem and a legal exposure. Pay for it before asking anyone to sign up, and publish the numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 for business-hours rotations, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover rather than expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, tax treatment and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a ~15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
10. Alert quality standard, page budget and noise burn-down (depends on: 3, 5)
3,400 alerts at 85% noise is the reason detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human, and burn the backlog down deliberately rather than by mass silencing.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a **page budget** of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements; return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
- Target 3,400 → under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
11. Detection uplift on the money path (depends on: 10, 5)
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls; five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrating without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack or the identity provider is unavailable. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live incident execution doctrine (depends on: 7, 12, 13)
Give responders one short operating procedure for the first minutes through closure. Priority is limiting customer and financial harm, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; controlled backlog drain; **reconciliation completed before any payment or ledger incident is declared resolved**.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and Customer Success, and reopening if impact recurs.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook and RSS; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack; the long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity or blame; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded **even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review; Legal owns outbound regulatory text, the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
18. SLA credit and financial impact workflow (depends on: 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one, and use the delta as the standing business case for on-call pay and reliability capacity.
19. Blameless postmortem standard and Incident Review Board (depends on: 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first?** and **why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, capacity reservation and enforcement (depends on: 19, 12)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
21. Runbooks, readiness bar and ledger blast-radius reduction (depends on: 5, 8)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the single largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel workstream on blast-radius reduction — tenant or function partitioning, read replicas, and isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (depends on: 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days + shadowing):** command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence for their own domain before independent primary duty; two shadow shifts minimum, never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (depends on: 22, 12, 21)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 11, 12, 21, 22, 9)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect (from first impact), time to declare, time to commander, acknowledgement, mitigate, resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; team scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Wave rollout to all 28 teams with readiness gates (depends on: 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director; failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
28. Alert noise burn-down campaign (depends on: 10, 12, 25)
Run the noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
29. Change management, fairness and pager culture (depends on: 4, 9, 25)
Run this from day one in parallel. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (depends on: 24, 26, 27)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and contingencies (depends on: 1)
Name the ways this programme fails and pre-commit the response. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 26, 27, 30)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 2 (ID: 2e38c07e-bf01-4b70-ba95-0ca1ef39e2d8, Agent: gpt5.6-sol_refine_2, LLM: openai/gpt-5.6-sol):
Estimated Complexity: high
Success Metrics: - By day 7, every suspected SEV0–SEV2 uses one incident record, one coordination channel, and one named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named commander for more than 10 minutes; at least 95% are assigned within 5 minutes.
- By day 30, 100% of Tier 0 services have a named owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, every Tier 0 and Tier 1 customer journey has compensated 24×7 command and subject-matter coverage.
- By day 120, all 180 production services have an owner, response tier, tested escalation path, and appropriate coverage model.
- No mandatory night rotation starts before its compensation, training, access, and staffing controls are active.
- Every critical primary rotation has at least six qualified responders or a documented, expiring executive exception by day 120.
- No responder is routinely scheduled for primary duty more often than one week in six by day 120.
- At least 95% of critical pages are acknowledged within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by day 120 and under 60 minutes by month 8.
- At least 95% of SEV0 and SEV1 customer notices are issued within 15 minutes of declaration by month 3.
- At least 95% of customer-visible SEV2 notices are issued within 30 minutes by month 3.
- At least 95% of incidents meet their required update cadence by month 3.
- Monthly paging volume falls from 3,400 to no more than 1,500 by day 90 and no more than 700 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6 without loss of critical detection coverage.
- All new paging alerts satisfy the owner, impact, action, dashboard, runbook, deduplication, and escalation standard by day 60.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 currently open historical actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date by month 6, with effectiveness evidence.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Two end-to-end cross-company exercises, including regional failure and ledger recovery, are completed before the audit.
- Monthly availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- SLA credits decline by at least 50% on an annualized trailing basis within 12 months.
- Quarterly on-call surveys show improving fairness and sustainability, with at least 75% favorable responses by month 6.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (24):
1. Create the mandate, ownership, and funding
Launch incident management as a company operating program within 48 hours. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable program owner.
- Form a small steering group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Fund two to three implementation staff, paging and incident tooling, observability work, training, exercises, and on-call compensation.
- Reserve 10%–15% of engineering capacity for alert remediation, runbooks, and incident actions.
- Authorize incident commanders to freeze deployments, order rollback, disable features, shift traffic, and invoke continuity plans.
- Preserve financial controls. Incident commanders may coordinate ledger recovery but may not bypass dual approval, privileged-access controls, or reconciliation.
- Set the delivery target: critical controls operational within 60 days, enterprise rollout within 120 days, and a mock audit in month 7.
2. Install an interim process in seven days (depends on: 1)
Do not wait for new tools or the final policy. Put a minimum viable incident process into operation immediately and start collecting evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communications schedule.
- Establish one monitored declaration path through chat, telephone, and the existing paging tools.
- Create a standard incident channel, bridge, incident document, and naming convention.
- Staff temporary primary and backup incident commanders 24×7 from the existing on-call teams and engineering leadership.
- Compensate interim duties retroactively under the final compensation policy.
- Require a named incident commander within 10 minutes for every suspected major incident.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Hold a daily 15-minute control review until the permanent process is live.
3. Build the baseline, ownership catalog, and risk map (depends on: 1)
Establish the facts behind the current failures and assign every production component an owner. Use the resulting catalog as the source for routing, escalation, and audit evidence.
- Reconstruct all 31 customer-impacting incidents, including first impact, detection source, declaration, command assignment, mitigation, resolution, customer communications, and credits.
- Analyze the two incidents with no clear leader and every case detected first by customers.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, databases, queues, regional dependencies, payment processors, and banking partners.
- Assign one accountable team, engineering manager, product owner, primary escalation, secondary escalation, dashboard, and runbook to each service.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 based on financial integrity, customer impact, contractual exposure, and dependency centrality.
- Map critical customer journeys to their application, PostgreSQL, Kubernetes, regional, and third-party dependencies.
- Inventory the six alert sources and all 3,400 monthly alerts by owner, volume, actionability, and duplication.
- Interview representatives from all 28 teams and baseline on-call sentiment, fatigue, and objections.
- Give orphan services an owner or decommissioning decision within 30 days.
4. Design compliance and evidence controls from day one (depends on: 1)
Map the operating process to audit, legal, contractual, and record-retention requirements before finalizing it. Confirm the expected SOC 2 observation period with the auditor immediately.
- Map controls to applicable SOC 2 criteria for monitoring, incident identification, response, recovery, communications, corrective action, access, and availability.
- Define evidence required for declarations, pages, acknowledgements, role assignments, decisions, status updates, postmortems, actions, training, drills, and exceptions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve the Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, and Postmortem Standard.
- Record control exceptions with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing evidence before the audit.
5. Adopt one severity and lifecycle standard (depends on: 2, 3, 4)
Use four impact-based severity levels across operational, security, data, and third-party incidents. Start at the highest credible severity when facts are uncertain, then downgrade with recorded evidence.
- **SEV0 — financial or security crisis:** suspected ledger corruption, unauthorized or duplicated funds movement, material data compromise, both-region loss, or a decision to suspend payment processing. Page all roles immediately; engage executives, Security, Legal, Compliance, and Risk; assess regulatory duties within one hour; use distinct role holders; require a postmortem.
- **SEV1 — critical availability event:** a core payment journey is unavailable, payment failures exceed an initial 10% guardrail for five minutes, regional loss has impaired failover, a settlement deadline is at imminent risk, or rapid error-budget burn makes a material SLA breach likely. Assign all roles; notify internal stakeholders within 10 minutes; publish customer status within 15 minutes; update every 30 minutes; require a postmortem.
- **SEV2 — major bounded event:** approximately 1%–10% of payment attempts fail, a material customer subset or critical customer is down, degradation has a workaround, or contractual impact is likely. Assign an incident commander and responders; add communications and scribe roles for customer impact; publish status within 30 minutes; update every 60 minutes; require a postmortem for customer-visible events.
- **SEV3 — limited event:** localized impact, a safe workaround, and no financial-integrity, security, regulatory, or material contractual risk. The owning team leads; page only if immediate action is necessary; use a ticket otherwise.
- Treat the percentage thresholds as declaration guardrails, not reasons to under-classify integrity, settlement, security, or strategic-customer risk.
- Permit any employee to declare an incident. Only the incident commander may lower severity, with the rationale logged.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of active customer harm. Declare resolution only after stability, backlog recovery, and required ledger reconciliation.
6. Define roles, authority, and handoffs (depends on: 5)
Separate command, communication, recordkeeping, and technical repair. One named person must hold command at every moment of a major incident.
- **Incident Commander:** owns severity, priorities, role assignment, escalation, decision cadence, mitigation coordination, and closure. The commander does not act as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language. Legal or Compliance retains ownership of regulatory submissions.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, and status changes.
- **Subject-Matter Responders:** diagnose and mitigate services for which they have accepted ownership, access, training, and runbooks.
- **Executive Duty Officer:** removes organizational obstacles and approves exceptional business decisions without displacing the incident commander.
- Require separate commander, communications, scribe, and primary technical lead for SEV0 and SEV1.
- For customer-visible SEV2, keep the commander separate from the primary technical responder; communications and scribe may be combined if workload permits.
- Announce every role assignment and transfer in the incident channel. Require verbal and written handoff with current impact, decisions, risks, and next actions.
- Keep executives and account managers out of the technical command path; questions flow through the Executive Duty Officer or Communications Lead.
7. Create sustainable 24×7 coverage across the service estate (depends on: 3, 6)
Use central command coverage and service-domain responder coverage rather than creating 28 fragile night rotations. Engineers remain responsible only for services they own or have formally trained to support.
- Build a company incident-command pool of approximately 18–24 certified senior engineers and managers, with a primary and backup scheduled at all times.
- Build 12–16-person communications and scribe pools using Support, Customer Operations, Engineering Operations, and qualified engineering managers.
- Schedule an Engineering Director or equivalent as the 24×7 Executive Duty Officer.
- Group related service owners into approximately 8–12 coherent responder domains only where members have training, access, and explicit acceptance.
- Require Tier 0 and Tier 1 domains to provide 24×7 primary and secondary responders, normally with at least six trained people in each sustainable rotation.
- Give Tier 2 services business-hours coverage plus a maintained manager escalation path. Treat Tier 3 conditions as tickets unless their impact changes.
- Maintain distinct but coordinated coverage for the ledger application and the PostgreSQL platform.
- Route unknown-owner events to the duty commander and platform triage temporarily. Treat every such event as an ownership control defect.
- If a small team cannot staff a fair rotation, merge coverage only after training or provide headcount, service reassignment, or decommissioning.
8. Approve compensation and fatigue protections (depends on: 7)
End unpaid on-call before expanding mandatory coverage. Publish the policy through HR, Legal, Finance, and Payroll within 14 days.
- Pay fixed weekly stipends for primary and secondary service rotations.
- Pay separate stipends for duty commander, communications, and scribe assignments.
- Provide additional call-out compensation or equivalent paid recovery time for material after-hours work.
- Apply overtime and reporting rules correctly for non-exempt employees under federal and New York requirements.
- Pay higher rates for company holidays and provide a protected recovery day after qualifying overnight work, SEV0 events, or prolonged SEV1 response.
- Target no more than one primary week in six and prohibit simultaneous primary assignments.
- Avoid consecutive primary weeks and make all swaps visible in the paging system.
- Reduce sprint commitments for people carrying primary duty rather than expecting normal delivery capacity.
- Provide a documented accommodation path for health, disability, or caregiving constraints without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
9. Set the service readiness and runbook standard (depends on: 3, 5, 7)
A team cannot respond effectively at night without ownership, access, telemetry, and rehearsed recovery procedures. Apply a formal readiness gate to every Tier 0 and Tier 1 service.
- Require a current architecture diagram, dependency map, dashboards, SLOs, runbooks, rollback method, feature-control method, contacts, and tested escalation path.
- Record RTO, RPO, data-integrity requirements, regional mode, and customer-facing capabilities in the service catalog.
- Link every paging alert to the exact runbook step expected from the responder.
- Prevent new paging alerts for services that fail readiness review. Preserve existing critical detection through a documented exception until a safe replacement exists.
- Write priority playbooks for PostgreSQL failure, ledger-integrity investigation, regional failover, Kubernetes control-plane degradation, payment-processor failure, queue backlog, credential compromise, and payment suspension.
- For the ledger, document read-only or stop-processing modes, failover controls, replay and duplicate protections, backlog handling, and post-recovery reconciliation.
- Require runbook review after material incidents and at least twice per year through exercises.
10. Establish one paging and incident system of record (depends on: 4, 5, 6, 7)
Consolidate paging and incident coordination without requiring an unsafe big-bang replacement of every monitoring system. Monitoring sources may remain specialized, but all human pages must enter one controlled platform.
- Select one enterprise paging platform and one integrated incident record, with chat, telephone, SMS, conference, status-page, ticketing, and service-catalog integrations.
- Initially ingest events from all six current tools, then deduplicate, correlate, and route by service ownership.
- Automate incident channel and bridge creation, role paging, timeline capture, severity changes, communications reminders, and postmortem creation.
- Preserve immutable records of declarations, acknowledgements, role assignments, decisions, and messages.
- Use role-based access, multifactor authentication, break-glass controls, and periodic access reviews.
- Provide telephone and offline fallback procedures for loss of chat, identity, the paging vendor, or an AWS region.
- Test paging and fallback paths weekly.
- Retire a legacy paging route only after its signals have owners, quality review, successful end-to-end tests, and at least two weeks of verified operation in the new path.
11. Enforce alert quality and burn down noise safely (depends on: 3, 10)
Treat paging alerts as production products with owners and quality requirements. Do not reduce noise by silently disabling detection.
- Require every page to identify the service, owner, customer or SLO risk, urgency, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Define an actionable page as one requiring prompt human judgment or intervention that materially reduces customer, financial, security, or contractual risk.
- Route informational, capacity-planning, and non-urgent conditions to dashboards or ticket queues.
- Prefer symptom and error-budget burn alerts over raw CPU, memory, pod, or log-volume thresholds.
- Run new alerts in shadow mode for at least seven days unless an emergency exception is approved.
- Review alerts with less than 50% actionability, more than three firings in seven days, or repeated no-action acknowledgements within two business days.
- Require compensating detection and an owner before suppressing or removing an alert.
- Set a responder load target of no more than two after-hours pages per person per week. A breach creates a mandatory alert-remediation plan.
- Review alert actionability, duplication, missed detection, and page load monthly by domain.
- Prioritize the small number of rules producing most of the current 85% noise.
12. Detect payment and ledger failures before customers (depends on: 3, 11)
Shift detection from infrastructure symptoms to customer journeys and financial outcomes. Set internal objectives stricter than the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, webhooks, and reporting freshness.
- Run external synthetic transactions from outside the production boundary and through both regions at least every minute for critical paths.
- Monitor payment failure rates, latency, queue age, delayed value, settlement-window risk, regional asymmetry, and third-party response quality.
- Add continuous ledger controls for reconciliation breaks, unexpected balances, duplicate identifiers, replication lag, backup health, and failover readiness.
- Add tenant or cohort anomaly detection for high-value customers and common payment methods.
- Convert high-priority Support, account-manager, bank, and processor reports into incident candidates within five minutes.
- Review every customer-first incident as a missed-detection defect and create a corrective action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Codify escalation and live incident execution (depends on: 6, 10)
Create one time-bound path from signal to ownership and mitigation. Delivery of a notification does not count as acknowledgement.
- Page the owning Tier 0 or Tier 1 primary immediately; page the secondary after five unacknowledged minutes; page the domain manager at 10 minutes; escalate to the engineering director at 15 minutes.
- For SEV0–SEV2, page the duty commander immediately. Page the backup after five minutes and require the Engineering Director to assume command if no certified commander owns the event by 10 minutes.
- Automatically involve command for integrity concerns, security concerns, regional events, cross-team impact, critical customer journeys, or unresolved ownership.
- If impact remains materially unknown after 15 minutes, increase response posture rather than waiting for certainty.
- Open one incident channel, bridge, and system record. State severity, known impact, assigned roles, current objective, and next update time.
- Freeze unrelated changes during SEV0 and SEV1 unless the commander records an exception.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require controlled backlog processing and reconciliation before resolving payment or ledger incidents.
- Require a formal command handoff for incidents extending beyond four hours or when fatigue impairs a role holder.
14. Standardize internal and customer communications (depends on: 5, 6, 10)
Communicate known impact early without waiting for root cause. The Communications Lead uses approved facts and always states the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish customer status within 15 minutes when customer-visible and legally safe, then update at least every 30 minutes.
- For SEV1, brief internal stakeholders and publish status within 15 minutes, then update every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send notice within 30 minutes, then update every 60 minutes.
- For SEV3, communicate directly to affected customers only when impact or contract terms require it.
- Give account managers an approved statement and affected-customer list within 30 minutes for SEV0 or SEV1 and within 60 minutes for SEV2.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. Describe affected capabilities, symptoms, workarounds, and the next update.
- Do not speculate about root cause, blame, data integrity, security scope, or recovery time.
- Post a Monitoring update within 30 minutes of mitigation. Post Resolved only after stability and required reconciliation.
- Provide a customer-facing incident summary within five business days for qualifying incidents.
- Record and approve any delay or restriction of public details during an active security threat.
15. Operationalize regulatory, partner, contract, and credit decisions (depends on: 4, 5, 14)
Give Legal and Compliance a timed decision process while keeping technical command with the incident commander. Record both reportable and non-reportable determinations.
- Build a jurisdiction and obligation matrix covering applicable NYDFS requirements, state breach laws, GLBA or FTC requirements, PCI obligations, money-transmitter rules, sponsor banks, payment networks, cyber insurance, and customer contracts.
- Validate applicability and deadlines with counsel rather than assuming every operational incident is reportable.
- Begin reportability assessment immediately for every SEV0, security event, suspected ledger-integrity event, and relevant SEV1.
- Target a documented initial legal assessment within one hour for SEV0 and within two hours for other potentially reportable events.
- Record the decision, evidence, approver, legal deadline, submission owner, and confirmation of delivery.
- Maintain tested 24×7 contacts for regulators, banks, networks, insurers, outside counsel, and critical vendors.
- Encode customer-specific notification clocks and channels in the customer record used by the Communications Lead.
- Have Finance calculate affected minutes, likely credits, and contractual exposure from the incident record within five business days.
- Review whether proactive credits or claims-based handling applies by customer segment and contract.
16. Make postmortems mandatory, consistent, and blameless (depends on: 5, 6, 10)
Use one review standard to learn from incidents and test whether controls worked. Keep learning reviews separate from performance or misconduct processes.
- Require a postmortem for every SEV0 and SEV1.
- Require one for customer-visible SEV2, customer-first detection, incidents lasting more than two hours, contractual breaches, repeat failures, control gaps, and ledger-integrity near misses.
- Produce the factual draft within three business days, conduct the review within five, and publish the approved version within 10.
- Use one template covering summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, lessons, and actions.
- Analyze why detection was not earlier and why mitigation took as long as it did.
- Examine technical, organizational, process, testing, dependency, and incentive factors rather than forcing a single root cause.
- Use a trained facilitator and written blameless rules. Describe decisions in the context and information available at the time.
- Publish broadly useful findings internally while maintaining restricted versions for security, privacy, personnel, or privileged material.
- Review postmortem quality and recurring factors in a weekly Incident Review Board.
17. Enforce action ownership and effectiveness tracking (depends on: 16)
Treat corrective actions as risk commitments rather than suggestions. Closing a ticket is insufficient without evidence that the control or system behavior improved.
- Give every action one named individual owner, manager, priority, due date, expected risk reduction, verification method, and linked work item.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with intermediate milestones.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Place recurrence-prevention actions for SEV0 and SEV1 ahead of discretionary feature work unless an executive accepts the residual risk.
- Escalate overdue high-risk actions to the manager after seven days, director after 14 days, and CTO after 30 days.
- Require written residual-risk acceptance, compensating controls, and a new review date when high-risk work is deferred.
- Verify completed actions through tests, telemetry, drills, or production evidence.
- Triage the 53 currently open historical actions within 30 days. Complete, re-plan, or formally accept the risk, prioritizing ledger, regional, security, and detection items.
18. Measure outcomes, controls, business impact, and human load (depends on: 10, 11, 14, 17)
Use a balanced scorecard so teams are not rewarded for suppressing alerts or avoiding incident declarations. Report medians and 90th percentiles, not averages alone.
- Measure time from first impact to internal detection, declaration, acknowledgement, command assignment, mitigation, and resolution.
- Track customer-first detection, missed escalations, status-page timeliness, update-cadence compliance, and role conflicts.
- Track incident count, recurrence, availability, error-budget burn, affected payment value, delayed transactions, reconciliation breaks, and SLA credits.
- Track page volume, actionability, duplicates, after-hours pages, missed acknowledgements, and load per responder.
- Track postmortem timeliness, action completion, action age, verified effectiveness, and repeated contributing factors.
- Track rotation size, duty frequency, recovery days, swaps, attrition signals, and quarterly responder sentiment.
- Hold a weekly Incident Review Board for incidents, actions, missed controls, and noisy alerts.
- Hold a monthly executive reliability review for trends, investment decisions, contractual exposure, and accepted risks.
- Hold a quarterly resilience and controls review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile dashboard data against customer cases and sampled incident records monthly to identify missing incidents or metric gaming.
19. Train participants and address pager resistance (depends on: 6, 7, 8, 13, 14, 16)
Introduce the operating model as a fair exchange, not an audit mandate. The message is that people are paid, paged only for accepted services, supported by trained command, and given capacity to fix defects.
- Train all employees to recognize impact, declare an incident, and find the incident channel and customer status page.
- Give all 260 engineers role-based training in severity, escalation, evidence preservation, handoff, and financial-integrity precautions.
- Certify incident commanders through instruction, simulation, two shadowed events or exercises, and observed performance.
- Train Communications Leads in status writing, account-manager briefing, contractual clocks, and legal escalation.
- Train scribes in timeline quality, decision capture, fact-versus-hypothesis labeling, and evidence handling.
- Require responders to demonstrate access, dashboards, rollback, runbooks, and escalation competence before primary duty.
- Require two shadow shifts before independent primary on-call.
- Appoint one adoption champion in each team and hold weekly office hours during rollout.
- Publish compensation, fatigue protections, service boundaries, and the accommodation process before assigning shifts.
- Use paid working time for training, exercises, shadowing, runbook work, and certification.
- Survey engineers at baseline, day 60, day 120, and quarterly thereafter.
20. Pilot on the payment critical path (depends on: 8, 9, 11, 12, 13, 14, 17, 19)
Run a four-to-six-week pilot across the highest-risk customer journey. Use real incidents and exercises to correct the process before wider rollout.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, authentication or API edge, settlement, and Support intake.
- Activate paid primary and secondary rotations, central command, communications, the incident record, status templates, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new system authoritative.
- Review every pilot page within one business day for routing, actionability, responder load, and missing context.
- Have the program team coach incidents without silently taking command from assigned role holders.
- Hold a weekly pilot retrospective and correct critical process or tool defects within 48 hours.
- Require 95% command assignment within five minutes, 95% timely communications, no unpaid pages, complete postmortems, and at least a 50% page-noise reduction before expansion.
21. Roll out by customer journey and risk (depends on: 18, 20)
Expand in fixed waves rather than waiting for every team to become perfect. Apply explicit readiness gates and time-limited exceptions.
- Days 0–7: operate the interim declaration, command, and communications process.
- By day 30: complete Tier 0 ownership, activate the first certified command roster, approve compensation, and begin the pilot.
- By day 60: provide compensated 24×7 coverage for every Tier 0 and Tier 1 customer journey and route all critical pages through the new platform.
- By day 90: complete critical detection upgrades, customer communications, regulatory playbooks, and the first major alert-noise reduction.
- By day 120: assign every production service a response tier, owner, tested escalation, and appropriate coverage model.
- Roll out remaining teams in two-to-three-week waves ordered by customer and dependency risk.
- Require each wave to pass ownership, alert, runbook, access, training, compensation, and tabletop gates.
- Disable legacy paging paths after verified cutover rather than leaving ambiguous parallel obligations.
- Publish a weekly adoption dashboard by team and escalate failed gates as business risks.
- Never start mandatory night coverage before compensation, staffing, training, and access are ready.
22. Exercise command, regional resilience, and ledger recovery (depends on: 9, 13, 14, 15, 19)
Validate the process under realistic conditions before depending on it during a crisis. Use the same action-tracking rules for exercises and real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops covering payment failure, customer-first detection, third-party failure, and ambiguous ownership.
- Exercise loss of one AWS region, Kubernetes degradation, PostgreSQL failure, suspected duplicate payments, queue backlog, and settlement risk.
- Exercise simultaneous operational and security events to test command and disclosure boundaries.
- Exercise loss of chat, status-page, identity, or paging providers using telephone and offline fallbacks.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors.
- Validate backups, restoration, RPO, RTO, failover prerequisites, financial controls, and post-recovery reconciliation.
- Avoid uncontrolled production-ledger experiments; use staging, replicas, simulations, or tightly governed production tests.
- Complete at least two cross-company exercises before the audit, including one overnight or unannounced paging test.
23. Test SOC 2 operating effectiveness before the auditor (depends on: 4, 18, 21, 22)
Demonstrate that the controls operate consistently, not merely that policies exist. Correct failures through tracked remediation rather than rewriting historical records.
- Preserve policy approvals, service ownership, schedules, compensation activation, access reviews, training, certifications, incidents, communications, postmortems, actions, and exercises.
- Sample evidence monthly from initial signal through verified action closure.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Use exceptions to document missed acknowledgements, late communications, incomplete records, and compensating controls.
- Conduct a formal mock audit in month 7 using the populations and interview roles expected by the external auditor.
- Trace at least one SEV0 or SEV1, one SEV2, a customer-reported event, and an exercise end to end.
- Remediate evidence and operating gaps before external fieldwork.
- Brief commanders, responders, Support, and Compliance for auditor interviews without scripting inaccurate answers.
24. Institutionalize continuous improvement (depends on: 18, 21, 23)
Keep incident management active after the audit by assigning permanent owners, budgets, and review cycles. Use incident trends to drive architectural investment.
- Assign permanent owners for the policy, service catalog, paging platform, status page, training program, metrics, and evidence repository.
- Review severity thresholds, communications timing, staffing, and compensation annually and after material process failures.
- Recertify commanders and communications leads annually through observed exercises.
- Review recurring failure families quarterly and require executive action when remediation repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fatigue, fairness, tool friction, and psychological safety.
- Report severe incidents, SLA exposure, overdue risks, resilience investments, and customer-first detection to the board or risk committee quarterly.
- Prioritize reduction of the shared-ledger concentration risk, stronger regional independence, deployment safety, graceful degradation, and automated mitigation.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
Previous Proposal 3 (ID: 2437bc94-f488-46b5-866e-01e1d4afbf3b, Agent: qwen3.8-max_refine_3, LLM: alibaba/qwen3.8-max):
Estimated Complexity: high
Success Metrics: - Median time to detect customer-impacting incidents falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median mitigation time falls from 3h10 to under 90 minutes by day 120 and under 60 minutes by month 9.
- A named incident commander is assigned within 5 minutes in at least 95% of SEV1 and SEV2 incidents; zero incidents remain unowned for more than 10 minutes.
- First status-page update occurs within 15 minutes of SEV1 declaration and 30 minutes of SEV2 declaration in at least 95% of qualifying incidents.
- Monthly paging volume falls from 3,400 to under 500 actionable pages within six months, with alert actionability above 80%.
- Out-of-hours pages average no more than 2 per responder per week; sustained breaches trigger mandatory alert remediation.
- Six legacy alerting tools are consolidated into one paging and incident platform, with legacy paging paths disabled by week 16.
- 100% of Tier 0 and Tier 1 services have a named owning team, escalation path, dashboard, and runbook by day 60.
- All 28 teams are onboarded with readiness gates by week 22; every 24x7 critical-path rotation has at least six trained responders.
- 30 or more certified incident commanders and 20 or more certified communications leads provide continuous primary and secondary coverage.
- Paid on-call is approved by HR, legal, and finance and active in payroll before any new mandatory night rotation begins.
- 100% of required SEV1 and SEV2 postmortems are drafted within 3 business days and published within 10 business days in the standard template.
- Postmortem action closure rises from 17% to at least 90% of high-priority actions completed by their due date within two quarters.
- All 53 open historical actions are triaged within 30 days; high-risk unaccepted items are completed or formally risk-accepted within 90 days.
- Repeat incidents from a known unaddressed contributing factor decline by at least 50% within six months.
- Annualized SLA credits fall from $1.3M to under $400k within 12 months.
- Monthly availability meets or exceeds 99.95% by month 6, with exceptions reviewed at the executive reliability meeting.
- At least two cross-company exercises, including regional and ledger scenarios, are completed before the SOC 2 audit, with critical findings tracked.
- The month-6 internal dry-run audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete evidence.
- SOC 2 Type II incident-response controls pass with zero exceptions at the month-8 audit.
- On-call sentiment improves quarter over quarter; fewer than 10% of engineers report unwillingness to participate in their owned-service rotation by month 6.
- No increase in attrition among engineers on active rotations compared with baseline.
Steps (27):
1. Executive mandate, program funding, and governance
Convert the CEO email into a **written mandate within 48 hours**. Name one accountable program owner and a small decision group. Time-box design to five weeks so rollout starts well before the audit.
- Appoint the CTO as executive sponsor and a Head of Reliability or Incident Management as program owner with full-time authority.
- Create an 8–10 person design working group: SRE/platform lead, payments and ledger engineering managers, support lead, compliance, legal, HR, finance, and two rotating engineering managers.
- Approve budget lines: tooling consolidation, on-call compensation, training, coaching, and 2–3 dedicated program staff.
- State the non-negotiables: paid on-call, named service ownership, one severity scale, one paging platform, mandatory postmortems, and protected engineering capacity for reliability actions.
- Fix the master timeline: interim controls week 1, design weeks 1–5, pilot weeks 6–11, full rollout weeks 12–22, internal audit rehearsal month 6, audit-ready month 7.
- Publish a one-page charter to the company that says incident response is a company operating process, not an optional team practice.
2. Evidence baseline from incidents, alerts, and coverage gaps (depends on: 1)
Before changing anything, create an auditable **before picture** from the last 12 months. This baseline drives the severity design, staffing model, and executive reporting.
- Reconstruct all 31 customer-impacting incidents: detection source, first owner, severity, mitigation time, credits paid, and whether command was clear.
- Write a specific case review of the two incidents with unclear ownership for more than one hour.
- Inventory the six alert tools, alert volume by team, noise rate, rules with no owner, and rules with no runbook.
- Map the 16 teams without on-call and the 12 teams with unpaid on-call.
- Review the 53 open postmortem actions and triage the highest-risk items first.
- Freeze baseline metrics: 22-minute detection, 40% customer-first detection, 3h10 mitigation, $1.3M credits, 3,400 alerts, 85% noise, and 17% action closure.
3. Stakeholder listening, resistance mapping, and co-design group (depends on: 1)
Treat engineer pushback as **design input, not an attitude problem**. The objection to carrying a pager for other teams' code must shape ownership, routing, compensation, and command staffing.
- Interview leads from all 28 teams plus support, customer success, sales, compliance, and HR within two weeks.
- Separate the real causes of resistance: unpaid night work, unfamiliar systems, poor runbooks, unfair load, fear of blame, or unclear authority.
- Recruit 10–15 respected engineers and managers as co-designers so the model is built with the teams.
- Survey baseline on-call sentiment, trust in alerts, and psychological safety; repeat at 90, 180, and 365 days.
- Document the central promise: responders are paged for services they own, trained incident commanders coordinate, and all on-call work is compensated.
4. Service ownership catalog, criticality tiers, and dependency map (depends on: 2)
No alert can page the right team until every service has a **named owner**. Build the catalog as the routing foundation for on-call, severity impact mapping, status-page components, and audit evidence.
- Assign one accountable team to each of the 180 services, with an engineering manager, Slack channel, escalation policy, dashboard, and runbook link.
- Tier services: Tier 0 for money movement, ledger integrity, authentication, settlement, and shared PostgreSQL; Tier 1 for customer-facing degradable services; Tier 2 for internal or batch services; Tier 3 for non-critical services.
- Map customer journeys to services, databases, regions, third parties, and contractual SLA components.
- Mark orphan and shared services; require ownership, reassignment, or decommissioning within 30 days.
- Document cross-region dependencies, failover constraints, and services that defeat nominal regional redundancy.
- Treat missing ownership for Tier 0 or Tier 1 services as an executive escalation and a release-blocking risk.
5. Severity scale, declaration rights, and automatic triggers (depends on: 2, 4)
Adopt one severity scale so declaration is a lookup, not a debate. **Anyone may declare; only the incident commander may downgrade.** When uncertain, start higher.
- SEV1: money movement stopped or incorrect, ledger integrity in doubt, confirmed security or data event, both regions impaired, or broad SLA-credit exposure.
- SEV2: major degradation, settlement window at risk, one or more critical customers fully down, or likely contractual breach.
- SEV3: limited impact with a workaround, no financial-integrity or security risk.
- SEV4: no customer impact; handled by ticket during business hours.
- Define fixed triggers for each level: pages, command staffing, bridge call, status-page timing, account-manager outreach, executive notification, and postmortem obligation.
- Add automatic escalation: SEV3 open more than 2 hours becomes SEV2; any incident touching the shared ledger or unclear ownership becomes SEV2 unless the commander documents otherwise.
- Publish a decision tree and 10–12 worked examples from the actual 31 incidents.
6. Incident roles, authority, and handover rules (depends on: 5)
Solve the **nobody-was-in-charge failure** by making command explicit, trained, and transferable. Separate coordination from technical remediation so responders are not asked to debug unfamiliar code.
- Incident Commander: owns severity, priorities, escalation, mitigation strategy, role assignment, handoffs, and closure. Does not write code during the incident. May freeze deploys, invoke pre-approved failover, and pull any on-call responder.
- Communications Lead: owns status page, internal updates, account-manager briefs, and coordination with legal or compliance.
- Scribe: maintains the timestamped record of decisions, actions, role changes, and customer communications.
- Subject-matter responders: engineers from owning teams who diagnose and remediate within their accepted domain.
- Executive liaison: required for SEV1; shields the commander from executive questions and owns board or regulator escalation.
- Require the commander to claim the role within 5 minutes of SEV1 or SEV2 declaration, state it in the incident channel, and record every handover.
- Allow role combination only for SEV3 or SEV4; require separate people for command, communications, and primary technical work at SEV1.
7. Paid on-call, fatigue safeguards, and HR compliance (depends on: 3)
Unpaid on-call is a retention, fairness, and legal risk in New York. **Pay must be approved before any new mandatory rotation starts.** This is the fastest way to reduce pager resistance.
- Create a weekly stipend for primary and secondary on-call, differentiated by tier and night responsibility.
- Add per-page or per-incident pay for out-of-hours activation, plus guaranteed recovery time after night work or major incidents.
- Create a separate stipend for duty incident commanders and communications leads because command is a heavier burden.
- Verify FLSA, New York wage-hour, overtime, holiday, and payroll treatment with HR, legal, and finance.
- Cap rotation frequency: no more than one primary week in four or one week in six depending on staffing; no simultaneous primary assignments.
- Require at least six trained responders for any 24x7 rotation; fund hiring or service reassignment where teams are too small.
- Publish the compensation package and payroll start date before asking engineers to join rotations.
8. On-call staffing model and rotation rules (depends on: 4, 6, 7)
Do not force 28 identical night rotations. Use a **layered model**: central command coverage, critical-path team coverage, and business-hours coverage for lower-tier services.
- Company incident-command and communications rotation: 24x7 pool of 25–35 trained volunteers and designated senior staff, giving primary plus secondary coverage at all times.
- Critical-path teams: 24x7 primary and secondary on-call for Tier 0 and Tier 1 domains, including payments, ledger, authentication, API edge, Kubernetes platform, and shared PostgreSQL.
- Other teams: business-hours on-call with a documented night escalation list owned by the engineering manager.
- Group services into 8–12 coherent domains so rotations are sustainable; no rotation with fewer than six people is approved without an executive exception.
- Require handoff overlap, shadow shifts before first primary duty, and no on-call during approved PTO.
- Define cross-team pull rules: the commander may page another team's on-call with a 10-minute acknowledgement obligation; this is coordinated by command, not pushed onto responders.
- Publish schedules, swap rules, and load limits in the paging tool.
9. Alert quality standard, page budget, and noise controls (depends on: 2, 4)
With 3,400 alerts and 85% noise, detection fails because people stop trusting pages. Make alert quality a **condition of paging anyone**.
- Every page must have a named owning team, customer or SLO impact, severity, dashboard, runbook, expected action, and escalation policy.
- Page on customer-visible symptoms: payment success rate, latency, settlement deadlines, error-budget burn, ledger integrity, and replication health.
- Demote cause-based CPU, memory, or infrastructure-only alerts to dashboards or tickets unless they map to a customer journey.
- Set a page budget: maximum 2 out-of-hours pages per person per week; breach triggers a mandatory alert-tuning sprint.
- Auto-flag alerts that fire repeatedly without action or have high no-action acknowledgement rates.
- Require shadow mode for new alerts before they page humans, except for documented emergencies.
- Review alert quality monthly by team and publish a noise leaderboard.
- Target fewer than 500 actionable pages per month and above 80% actionability within six months.
10. Detection uplift across payments, ledger, and customer signals (depends on: 4, 9)
Customers detected 40% of incidents first. Detection must shift to **payment outcomes, ledger integrity, and inbound customer signals**, not host metrics alone.
- Define SLIs and SLOs for payment initiation, authorization, settlement, reconciliation, refunds, API availability, and reporting freshness.
- Run synthetic end-to-end payment tests from outside the platform in both AWS regions every 60 seconds.
- Add continuous ledger assurance: double-entry reconciliation, replication lag, failover readiness, disk pressure, and settlement-window countdown alerts.
- Monitor per-customer anomalies for top accounts so a single-tenant outage is detected before the account manager calls.
- Convert high-priority support tickets and account-manager reports into incident candidates within 5 minutes.
- Track detection source for every incident; make customer-first detection a reviewed defect with a corrective action.
- Add partner, banking, and card-network notification intake into the same declaration path.
11. Single incident platform and alert-tool consolidation (depends on: 5, 8, 9)
Collapse six tools into **one paging and incident-management system** with one queue, one timeline, and one audit trail. Do not create a parallel audit process.
- Select an integrated stack: paging and on-call scheduling, incident workflow, Slack or chat integration, bridge calling, status-page API, and ticketing integration.
- Implement one-command declaration that creates the incident record, channel, bridge, severity label, role prompts, and clock.
- Migrate alert sources by wave; retire a legacy paging path only after routing, ownership, and acknowledgement tests pass.
- Automate evidence capture: timestamps, acknowledgements, role assignments, severity changes, communications, and postmortem links.
- Integrate with the service catalog, on-call schedules, Jira or equivalent action tracker, customer-success tooling, and the status page.
- Verify the platform works during a single-region failure, including SMS, phone, and offline fallback paths.
- Set a hard date after which pages outside the chosen platform are not valid on-call obligations.
12. Escalation paths and five-minute command rule (depends on: 6, 8, 11)
Create one path from **signal to named commander in under five minutes**, any hour of the day. Make escalation automatic and time-bound.
- Accept declarations from automated alerts, engineers, support, account managers, partners, and customers through the same command.
- For SEV1 and SEV2, page the duty incident commander and owning-team primary immediately.
- Acknowledgement ladder: primary 5 minutes, secondary 10 minutes, manager 15 minutes, director or executive 20 minutes.
- If no commander claims the incident within 5 minutes, the platform assigns and announces one; the assignee may hand over but cannot leave the incident unowned.
- For SEV3, require team acknowledgement within 30 minutes; otherwise create a tracked work item.
- Give the commander pre-approved authority to invoke regional failover, ledger read-only mode, feature kill switches, and partner notifications without waiting for executive sign-off.
- Record every missed acknowledgement and escalation failure for weekly review.
13. Internal communications protocol (depends on: 6, 11)
Separate the **working incident channel from the audience channel** so responders can work and executives, support, and sales stay informed without interrupting the commander.
- Create one incident channel and one bridge per SEV1–SEV3 incident; use a read-only broadcast channel for executives and support.
- Set update cadence: every 15 minutes for SEV1, 30 minutes for SEV2, and at state changes for SEV3.
- Use a fixed update template: impact, customer-visible symptoms, current action, ETA or next update, commander, and communications lead.
- Brief support and customer success with a live affected-customer list and approved holding statements within 15 minutes of SEV1 or SEV2.
- Require executives to route questions through the executive liaison; the commander is not interrupted.
- Define fatigue rules and formal handover for incidents lasting more than 4 hours.
14. Customer status page, account-manager outreach, and SLA credit workflow (depends on: 5, 13)
Replace ad-hoc status updates with a **timed, owned, template-driven customer communication process**. Link incidents to SLA credits so finance and customer success are not surprised.
- Publish status-page updates within 15 minutes of SEV1 declaration and 30 minutes for SEV2; update every 30 or 60 minutes until resolution.
- Assign the communications lead as the single author; use pre-approved templates reviewed by legal and communications.
- Map status-page components to customer capabilities: payments, settlement, reporting, API, onboarding, and regional availability.
- For top accounts, require direct account-manager outreach within 30 minutes of SEV1 with an approved briefing pack.
- Send proactive email or webhook notifications to subscribed customers for SEV1 and SEV2.
- Publish a resolution notice within 30 minutes of mitigation and a customer-facing incident summary within 5 business days for SEV1.
- Automate SLA credit calculation from incident duration and affected capability; review with finance and legal within 5 business days.
- Track credits by incident and root-cause family to guide reliability investment.
15. Regulatory, legal, and partner notification playbook (depends on: 5, 14)
In payments, some incidents are reportable and the clock starts at detection. Make regulatory assessment a **mandatory step in the incident process**, not an afterthought.
- Map obligations: NYDFS cybersecurity-event rules, state breach laws, GLBA safeguards, PCI DSS if in scope, money-transmitter duties, card-network and sponsor-bank contracts, and cyber-insurer notice.
- Require a legal or compliance reportability assessment within 2 hours for every SEV1 and every security-related SEV2, even when the answer is not reportable.
- Maintain a 24x7 contact matrix for regulators, sponsor banks, card networks, outside counsel, insurers, and law enforcement.
- Pre-draft notification templates and preserve legal review.
- Encode customer-contract notification deadlines into account tiers and the communications workflow.
- Record every reportability decision, approver, deadline, and submission confirmation in the incident record.
16. Postmortem standard and blameless review (depends on: 5, 6)
Replace inconsistent postmortems with a **mandatory, blameless, single-format process**. The discipline comes from deadlines, facilitation, and action tracking.
- Require postmortems for all SEV1 and SEV2 incidents, customer-detected incidents, incidents over 2 hours, repeat failures, and ledger near-misses.
- Draft within 3 business days, peer review within 5, publish within 10 for SEV1 and SEV2.
- Use one template: summary, impact, timeline, detection analysis, response analysis, contributing factors, what worked, what failed, and action items.
- Make blamelessness explicit: focus on systems and decisions, not individual fault; never use postmortems in performance discipline.
- Hold a weekly incident review board to review postmortems, ratify severity, and challenge weak actions.
- Maintain a searchable postmortem library and quarterly recurring-cause analysis.
- Require a trained facilitator for major reviews; the incident commander attends but does not facilitate.
17. Action-item tracking, ownership, and delivery gates (depends on: 16, 11)
Only 11 of 64 actions were closed. Give postmortem actions the same status as **customer commitments**, with named owners and visible escalation.
- Create every action as a ticket with one named individual owner, priority, due date, and verification method.
- Use delivery classes: containment within 7 days, corrective work within 30 days, strategic work within 90 days.
- Reserve 15–20% of team sprint capacity for reliability and incident actions.
- Escalate overdue items: manager at 7 days, director at 14 days, CTO dashboard at 30 days.
- Block related feature releases when overdue P0 actions prevent recurrence of a severe incident.
- Require director approval and documented residual risk acceptance for overdue high-risk items.
- Verify effectiveness after completion; closing a ticket without evidence does not close the action.
- Target 90% of high-priority actions completed on time within two quarters.
18. Runbooks, critical-incident playbooks, and readiness bar (depends on: 4, 8)
Poor runbooks are a real cause of pager resistance and slow mitigation. Define a **minimum readiness bar** before a service is allowed to page anyone at night.
- Require for every Tier 0 and Tier 1 service: architecture summary, dependencies, dashboards, alert-to-runbook map, rollback procedure, feature flags, escalation contacts, and customer-impact statement.
- Write major playbooks for shared PostgreSQL ledger failure, regional failover, Kubernetes control-plane loss, payment-processor outage, settlement-window breach, duplicate-payment suspicion, and security compromise.
- Define ledger recovery rules: failover procedure, read-only degraded mode, reconciliation, RPO/RTO, and data-loss tolerance approved by executives.
- Test runbooks in drills at least twice a year; mark untested runbooks stale.
- Prevent paging alerts for services without readiness sign-off unless the engineering manager accepts the gap in writing.
- Keep runbooks linked from every alert and incident template.
19. Training, certification, and role readiness (depends on: 6, 12, 13, 16)
Command and communications are skills. Build a **tiered certification path** so rotations are staffed by people who have practiced, not by whoever is around.
- All employees: 1-hour module on declaring incidents, finding the incident channel, and reading the status page.
- All responders: half-day training on severity, acknowledgement, escalation, runbooks, and evidence hygiene.
- Incident commanders: 2-day course plus two shadowed incidents and one simulation before certification.
- Communications leads: training on status-page writing, customer language, account-manager briefs, and regulatory triggers.
- Scribes: training on timeline discipline and audit evidence.
- Certify for 12 months; renew through a simulation.
- Require shadow shifts before independent primary duty; no new hire holds primary within 90 days.
- Publish the certification register as an audit artifact.
20. Simulation program and game days (depends on: 19, 11, 18)
Rehearse the process before it meets a real SEV1. Simulations build commander confidence, expose runbook gaps, and produce audit evidence.
- Run monthly 60-minute tabletops using real incidents from the 31-incident baseline.
- Run quarterly game days covering regional failover, ledger replica promotion, dependency failure, partner outage, and security event.
- Run twice-yearly unannounced paging drills to measure night acknowledgement times.
- Include support, account managers, legal, compliance, and executives in at least one exercise per quarter.
- Produce tracked action items from every exercise using the same board as real incidents.
- Measure time to commander, time to first status update, and time to mitigation decision.
21. Critical-path pilot and gate review (depends on: 7, 10, 11, 14, 18, 19)
Prove the model on the highest-risk services with willing teams before full rollout. Run a **six-week pilot with daily feedback and public exit criteria**.
- Pilot with payments, ledger/database, platform/Kubernetes, API edge, authentication, and support intake.
- Activate severity scale, duty commanders, paid rotations, single paging platform, alert budget, status-page policy, and postmortem process.
- Hold a weekly pilot retrospective and fix process defects quickly.
- Validate night acknowledgement, cross-team pull response, severity clarity, and compensation payroll.
- Exit gate: commander assigned within 5 minutes in 95% of incidents, status page on time, page noise down at least 50%, postmortems on time, and positive on-call sentiment.
- Publish pilot results to the whole company as the main adoption argument.
22. Wave rollout across all 28 teams (depends on: 21)
Roll out by criticality and dependency, not by calendar alone. Use **readiness gates** so teams are not forced live without coverage.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1 teams. Wave 3: Tier 2 teams. Wave 4: Tier 3 and internal platforms.
- Gate per team: catalog entry complete, alerts migrated, runbooks ready, six-person rotation staffed, one commander candidate nominated, compensation in payroll, and one tabletop passed.
- Assign a program coach to each wave for three weeks.
- Freeze legacy paging paths for each team after successful onboarding.
- Publish a live adoption scoreboard by team.
- Complete all teams by week 22, leaving months of operating evidence before the audit.
23. Metrics, dashboards, and review cadence (depends on: 11, 16, 21)
Instrument the process itself. Leadership must see whether the program is working, and auditors must see **operating evidence, not retrospective paperwork**.
- Track response metrics: MTTD, time to declare, time to commander, acknowledgement time, mitigation time, resolution time, and customer-first detection rate.
- Track quality metrics: page volume, alert actionability, missed pages, postmortem timeliness, action closure rate, and status-page compliance.
- Track business metrics: availability against 99.95%, SLA credits, repeat incidents, and error-budget consumption.
- Track people metrics: on-call load, night pages per person, recovery time usage, sentiment, and attrition signals.
- Hold weekly incident review, monthly reliability review, quarterly executive review, and annual policy review.
- Publish dashboards internally and give every metric a target and owner.
24. Change management, incentives, and culture (depends on: 3, 7, 21)
The process will be judged on fairness. Communicate the deal repeatedly and make participation **recognized, compensated, and safe**.
- Core message: you are paid, paged only for what you own, supported by a trained commander, and given real capacity for actions.
- Run CTO all-hands, team roadshows, office hours, FAQ, and an internal incident-management hub.
- Add incident response and reliability work to promotion criteria and manager objectives.
- Recognize good postmortems, alert-noise reduction, and calm incident leadership.
- Provide a written path for engineers who cannot do nights; cover those shifts with paid volunteers or adjusted staffing.
- Prohibit retaliation for good-faith declaration or escalation.
- Publish sentiment survey results, including bad news, to maintain credibility.
25. SOC 2 evidence design and internal dry-run audit (depends on: 16, 17, 22, 23)
Design the process so audit evidence is a **by-product of normal operations**. Test it internally before the external auditor does.
- Map the process to SOC 2 criteria: incident identification, response, recovery, monitoring, communication, control activities, availability, and corrective action.
- Approve versioned policies: incident response policy, severity standard, on-call policy, communications policy, postmortem standard, and exception process.
- Retain incident records, paging logs, status-page history, postmortems, action tickets, training records, drill records, and access reviews for the audit period.
- Record regulatory reportability decisions even when no notification is required.
- Run an internal dry-run audit at month 6: sample at least 15 incidents and trace evidence end to end.
- Fix gaps with at least four weeks before the audit.
- Maintain an exception and remediation log instead of claiming perfection.
26. Program risk register and contingency planning (depends on: 1)
Name the likely failure modes now and pre-commit responses. Review the register monthly with the executive sponsor.
- Commander volunteer shortfall: roster command duty for senior engineers and managers until the pool reaches 25 trained people.
- Compensation delay: use immediate time-off-in-lieu plus phased stipend, but do not start mandatory night rotations without pay.
- Tool migration slip: preserve single paging and routing scope; defer advanced automation if needed.
- Alert pruning causing missed incidents: move noisy alerts to ticket first, observe 30 days, then delete.
- Burnout in experienced teams: monitor load weekly and cap pages per person.
- Major SEV1 during rollout: program owner shifts to incident support, wave schedule slips one wave, sponsor is informed same day.
- Shared ledger concentration: track blast-radius reduction and failover improvements as top strategic actions.
27. Continuous improvement, maturity roadmap, and post-audit sustainability (depends on: 23, 25)
Prevent the classic failure where the process decays after the audit. Build the second-year operating plan before the first year ends.
- Hold quarterly process retrospectives with commanders, responders, support, and communications leads.
- Re-baseline targets every six months; tighten goals once current targets are met.
- Move from lagging metrics to leading indicators: error-budget burn, near-miss rate, drill performance, and action ageing.
- Plan year-two improvements: follow-the-sun coverage, automated mitigation, ledger blast-radius reduction, error-budget release gates, and per-customer real-time impact reporting.
- Keep annual policy review, certification renewal, drills, and board reporting on a permanent calendar independent of the audit cycle.
- Report quarterly to the board or risk committee on availability, credits, severe incidents, overdue actions, and resilience investment.
Previous Proposal 4 (ID: 185429cf-5cd1-451d-ab7e-34d64ca403df, Agent: grok4.6_refine_4, LLM: xai/grok-4.6):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 10 minutes by month 3 and under 5 minutes by month 9.
- Customer-first detection falls from 40% to under 20% by month 3 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by month 4 and under 60 minutes by month 12.
- A named Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes after week 4.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases by month 3, with required update cadence met in 95% of cases.
- Monthly pages fall from 3,400 to under 500 within 6 months, with actionability above 75%; out-of-hours pages at or under 2 per person per week.
- All six legacy alerting tools route through one paging platform by week 16; legacy paging paths disabled per wave.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by due date within two quarters; all 53 currently open historical actions triaged within 30 days of S2.
- SLA credits fall from $1.3M to under $400k in the first 12 months; customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known unaddressed cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar; every Layer B rotation has 6+ certified responders or a time-limited executive exception.
- Paid on-call policy approved by HR, Legal, and Finance and in payroll before any mandatory night rotation starts.
- All 28 teams onboarded by week 24 with 24x7 command cover from 25+ certified ICs and 15+ certified Comms Leads.
- Internal dry-run at month 6 passes a 15-incident evidence walkthrough; month-7 mock audit finds no unowned high-risk control gap; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment improves quarter over quarter; no increase in attrition among engineers on rotation; pulse at 60 and 120 days shows at least 70% agree rotations are fair and limited to services they own.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked actions; two cross-company exercises completed before the audit.
Steps (24):
1. Program charter, mandate, and funding
Convert the CEO email into a named program with one owner, a budget, and a deadline earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, 28 teams will opt out.
- Appoint a Program Lead (Head of Reliability) with direct CTO and CEO sponsorship.
- Form a small steering group: CTO, VP Eng, Head of CS/Support, CISO, Legal, Finance, HR. Not a 28-team committee.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Timeline: operating floor in 7 days; design weeks 1–6; pilot weeks 7–12; rollout weeks 13–24; mock audit week 28; SOC 2 at month 8.
- Fund tooling ($150–250k/yr), on-call pay (~$600k–1M/yr), 2–3 program FTEs, and reserved engineering capacity. Anchor the ask against $1.3M in credits plus unmeasured incident cost.
- Freeze baselines now: 31 incidents, MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, $1.3M credits, 3,400 alerts/month at 85% noise, 11 of 64 actions closed.
2. Immediate 7-day operating floor (depends on: 1)
Do not wait for tooling, compensation, or the audit. Put a minimum process in place this week so the next outage has a named commander.
Start retaining artifacts on day one. This week becomes the first audit evidence.
- Publish a one-page interim severity guide and a single declaration path through Slack, phone, and pager.
- Staff interim primary and backup Incident Commander 24x7 from existing on-call veterans and engineering managers. Compensate this duty retroactively.
- Require a named IC within 10 minutes for every suspected major incident. If nobody claims it, the duty manager is IC.
- Use one channel, one bridge, one timeline doc, and one naming convention for every major incident.
- Direct Support to escalate credible customer reports immediately. They do not wait for engineering confirmation.
- Triage all 53 open historical actions. Close, re-plan, or formally risk-accept. Ledger integrity, payment duplication, regional resilience, security, and detection come first.
- Hold a daily 15-minute ops review until the permanent process is live.
3. Forensic baseline of incidents and alerts (depends on: 1)
Rebuild the facts before locking design. This is the before picture for the CEO and the auditor.
Every later design choice should trace to this evidence.
- Re-code each of the 31 incidents: trigger, service, detection source, timestamps, who led, credits paid, root-cause family.
- Quantify the 40% customer-first detections and name the missing signal in each case.
- Reconstruct the two nobody-in-charge incidents minute by minute. Use them as the burning-platform story.
- Audit the six alerting tools: volume per tool and team, top 50 noisy rules, rules with no owner or runbook.
- Freeze the baseline numbers. Do not let them drift during design.
4. Listening tour and the fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the main delivery risk. Treat it as a design constraint, not an attitude problem.
The answer that will actually stick is **the deal**: you are paid; you are paged only for services you own; a trained commander runs the room; your postmortem actions get real sprint capacity.
- Interview all 28 teams plus Support, CS, and Sales in two weeks.
- Separate the objections: unpaid work, nights, unfamiliar code, bad runbooks, fear of blame. Each needs a different fix.
- Collect informal practices from the 12 teams already on-call. They are the pilot candidates and the veteran pool.
- Recruit 10–15 credible engineers as a design working group so the process is co-authored.
- Baseline sentiment on alert trust, on-call willingness, and burnout. Re-measure at 6 and 12 months.
5. Service ownership catalog and criticality tiers (depends on: 3)
You cannot page the right person across 180 services until each one has a named owner. This is the foundation of fairness, routing, and audit evidence.
Build a machine-readable catalog as the single source of truth.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, the shared PostgreSQL cluster, queues, partner banks, and customer-facing endpoints.
- Assign one owning team, named engineering manager, Slack channel, escalation policy, and dependency list per service.
- Tier 0: money movement, ledger, auth, shared Postgres, regional control plane. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map Tier 0/1 services to customer capabilities: initiation, authorization, settlement, reporting, onboarding.
- Assign coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
6. Severity scale and declaration rules (depends on: 3, 5)
Replace in-the-moment debate with a lookup. Four payments-specific levels, with worked examples drawn from the 31 real incidents.
Anyone may declare. Only the Incident Commander may downgrade, with recorded rationale. When unsure, start high.
- **SEV1**: money movement stopped or incorrect; ledger integrity in doubt; material security or data exposure; both regions impaired; or more than 10% of customers impacted. Pages IC, comms, scribe, SMEs, and exec liaison. Bridge in 5 minutes. Status page in 15 minutes. Mandatory postmortem and regulator assessment.
- SEV2: severe degradation; settlement window at risk; a strategic customer fully down; SLA breach likely. IC and SMEs paged. Status page in 30 minutes. Mandatory postmortem.
- SEV3: partial impact with a workaround; no credit exposure. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, or a repeat.
- SEV4: internal or minor. Ticket only. No page.
- Auto-escalate: any SEV3 open more than 2 hours, or any incident touching the shared ledger, becomes SEV2. Unknown impact after 15 minutes is raised, not sat on.
- Nobody is punished for over-declaring. Publish that rule in writing and repeat it.
7. Roles, authority, and ledger dual-control (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, transferable, and logged.
The Incident Commander owns the incident, not the fix, and **never types** in a production terminal.
- Incident Commander: declares severity, pulls anyone, freezes deploys, invokes failover, authorises spend. One IC at a time. Assumed within 5 minutes and announced in channel.
- Communications Lead: single voice to customers, status page, account managers, and the exec summary.
- Scribe: timestamped timeline, decisions, and open questions. Feeds the postmortem and the audit trail. Required for SEV1 and SEV2.
- Subject-matter responders: diagnose and mitigate only services they own, with access and runbooks.
- Executive Liaison (SEV1): shields the IC from exec questions; owns regulator and board escalation.
- The IC may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access rules.
- Handover is verbal and written, with time and new owner recorded. Roles may combine below SEV2; never at SEV1. The IC stays in charge if a VP joins.
8. Three-layer 24x7 coverage model (depends on: 5, 7)
Do not put 28 teams on night rotation. That is what engineers are rejecting.
Staff command centrally. Page engineers only for code their team owns.
- Layer A, company command: Duty IC plus secondary, Duty Comms, and a scribe pool. Target 25–35 certified ICs and 15–20 comms leads. Roughly one week per person every 6–8 months. Five-minute ack SLA, then secondary, then on-call director.
- Layer B, Tier 0/1 domains: group 28 teams into 8–12 product and platform domains (ledger app, Postgres platform, payments orchestration, auth, API edge, Kubernetes/platform, settlement). Primary plus secondary 24x7. Minimum six trained people. Target one week in six, never worse than one in four.
- Layer C, Tier 2/3: business-hours on-call. After hours the IC pages the EM, who holds a written escalation list.
- No engineer joins another team's responder pool without training, access, runbooks, shadow shifts, and both teams' acceptance.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner.
- Evaluate follow-the-sun coverage as a 12-month option, not a year-1 dependency.
9. Paid on-call, NY labor compliance, and fatigue rules (depends on: 4, 8)
Unpaid on-call is a retention risk and a New York legal exposure. Pay must be live in payroll before any mandatory night rotation starts.
Treat interrupted nights and recovery time as compensable work.
- Weekly stipend by layer, higher for 24x7 Tier 0/1 and Duty IC, lower for business-hours, with holiday and weekend premiums.
- After-hours call-out pay or TOIL. A paid recovery day after prolonged overnight work, a SEV1, or a qualifying SEV2. Managers cover the next day.
- HR, Legal, and Finance publish dollar amounts, FLSA exempt/non-exempt treatment, NY wage-hour rules, tax treatment, and payroll timing within 14 days of charter.
- Load rules: no primary on two rotations; no consecutive primary weeks; no on-call the week after a SEV1 you commanded.
- A person may declare temporarily unfit after overnight work with no performance penalty.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- If a rotation cannot staff six people, merge domains or hire. Do not run two-person 24x7.
- Model annual cost against $1.3M in credits and get it as a CFO/board line item.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts a month at 85% noise is why detection takes 22 minutes. Make quality a condition of paging a human.
A page is a product with a quality bar, not a dump of host metrics.
- Every paging alert must have a named owner, customer or SLO impact, runbook, tested threshold, severity mapping, dashboard, expected action, and dedup key. Fail any of these and it becomes a ticket or is deleted.
- Page on customer symptoms: SLO burn, error budget, settlement-queue age. Cause-based CPU and memory alerts become dashboards or tickets.
- Set a **page budget** of at most two out-of-hours pages per person per week. A breach triggers a mandatory tuning sprint and blocks new paging alerts for that team.
- Auto-quarantine alerts that fire more than five times a month with no action, or that have more than 70% no-action acks. Never silently disable without compensating detection and a recorded decision.
- Run new alerts in shadow for seven days unless an emergency exception is approved.
- Monthly per-team kill, tune, or keep review. Target under 500 pages a month and actionability above 75% within six months.
11. Customer-journey detection and ledger assurance (depends on: 5, 10)
Stop customers being the monitoring system. Detect on money-path outcomes, not host metrics.
Every postmortem will ask why a customer saw it first.
- Define SLOs per Tier 0/1 capability: initiation success, auth latency, settlement timeliness, API availability, reporting freshness. Set an internal target stricter than 99.95%.
- Run synthetic full-lifecycle payments from outside the platform, in both regions, every 60 seconds. Include a small real-value canary where legally feasible.
- Add ledger assurance: continuous double-entry reconciliation, replication lag, failover-readiness, unexpected balances, duplicate identifiers, settlement-window countdown.
- Add per-customer anomaly detection for the top 100 accounts.
- Auto-create a triage incident within five minutes when support tickets or AM reports match impact keywords.
12. Single incident and paging platform (depends on: 6, 8, 10)
Collapse six alerting tools into one queue, one timeline, and one audit record.
The platform must still work if one AWS region is down. Verify SMS and phone fallback, plus an offline runbook.
- Select one paging product, one incident layer, and one hosted status page.
- One Slack command declares the incident, creates the channel and bridge, pages the Duty IC, sets severity, and starts the clock.
- Ingest the six existing tools first. Deduplicate and route. Retire a legacy path only after named owners and two weeks of verified operation.
- Auto-capture timestamps, acknowledgements, role assignments, severity changes, comms, mitigated time, and resolved time. Export for SOC 2.
- Integrate the service catalog, Jira, Salesforce or CS tooling for affected-customer lists, and Zoom or Slack Huddle.
- Test paging, escalation, status publication, and conference access every week.
- After a team's wave, old paging paths are disabled, not left as a fallback.
13. Detection-to-command escalation path (depends on: 7, 8, 12)
Write the single path from something looks wrong to someone is in charge. Target: a commander in under five minutes, any hour.
If nobody claims IC in five minutes, the platform assigns the Duty IC. The assignee can hand over, not decline.
- Converge every entry point on the same declare command: alert, engineer, support, account manager, partner bank, SEV hotline.
- SEV1: page the owning primary immediately; secondary at 5 minutes unacked; domain manager and Duty IC at 10; exec liaison at 15.
- SEV2: primary ack in 10 minutes; IC assigned in 15.
- Human acknowledgement is required. Delivery to a device does not count.
- The IC can page any team's on-call, with a 10-minute ack obligation. This reciprocity makes single-team ownership viable.
- Unowned alerts go to Layer A command, then the missing owner record is a control defect.
- Pre-authorise regional failover, ledger read-only mode, and partner-bank notice so the IC does not wait for an executive. Dual-control still applies to ledger writes.
14. Live execution and major-incident playbooks (depends on: 7, 13)
Limit customer and financial harm before proving root cause. One procedure from the first minute to handback.
A service cannot page at night until it meets the readiness bar.
- Open channel, bridge, record, and timeline immediately for SEV1 and SEV2. The IC states severity, known impact, hypothesis, objective, roles, and next update time.
- Freeze unrelated production changes during SEV1. Record exceptions the IC approves.
- Prefer reversible mitigation: rollback, feature flag, traffic isolation, rate limit, partner reroute.
- Write playbooks first for Postgres ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank outage, settlement-window breach, suspected security compromise, and suspected duplicate payments.
- Guard against split-brain, replay, and duplication during regional or database recovery. Reconcile and process backlog before calling the incident resolved.
- Mitigation means customer impact ended. Resolution means stable, backlog processed, and ledger reconciled.
- Formal IC handover after 4 hours. Reopen if impact recurs during the stability window.
- Readiness bar per Tier 0/1 service: diagram, dependencies, dashboard, runbook, rollback, kill switch, escalation contacts, RPO/RTO. Untested runbooks are marked stale.
15. Internal, customer, and regulatory communications (depends on: 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved process. The Communications Lead is the single author.
State impact and the next update time. Never speculate on cause.
- Internal: one working channel, one bridge, one read-only broadcast for execs, Support, and Sales. SEV1 update every 30 minutes even if unchanged; SEV2 every 60 minutes.
- Executives ask questions only of the Executive Liaison. Publish this as a signed exec behaviour rule.
- Status page: SEV1 within 15 minutes, SEV2 within 30 minutes; then 30/60-minute updates; resolve notice within 30 minutes of mitigation. Templates pre-approved with Legal.
- Top 100 accounts: named AM contact within 30 minutes of SEV1, with a briefing pack from Comms. Long tail gets the status page plus email or webhook.
- Customer-facing summary within 5 business days for SEV1.
- Mandatory regulatory checkpoint on every SEV1 and every security SEV2, within 2 hours, recorded even when not reportable. Cover NYDFS Part 500 72-hour clock, state breach laws, GLBA/FTC, PCI if in scope, sponsor-bank and card-network windows, FinCEN/OFAC if relevant.
- Legal owns outbound regulatory letters. The IC owns facts. Security or Legal may limit public detail during an active threat, with the reason recorded.
- Encode bespoke customer-contract notice SLAs into account tiering.
16. SLA credit and financial-impact workflow (depends on: 6, 15)
Link incidents to money so severity, credits, and investment stay consistent. Finance should not learn about outages from invoices.
Make credit calculation an output of the incident record, not a negotiation.
- Agree availability measurement per contract and component with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and capability telemetry. Produce a proposed credit schedule within 5 business days of resolution.
- Posture: proactive credits for top-tier accounts; claims-based for the rest. Document the approval chain.
- Track credits per incident and root-cause family. Quarterly report which reliability investments would have prevented which credits.
- Target: cut credits from $1.3M to under $400k in 12 months. Use that delta as the ongoing business case.
17. Blameless postmortems and action tracking (depends on: 6, 12)
Eleven of 64 actions closed is the clearest process failure. Make learning mandatory and actions as binding as customer commitments.
Closing a ticket without evidence of effectiveness does not close the action.
- Mandatory for every SEV1 and SEV2, every customer-first detection, every incident over 2 hours, every repeat of a known cause, and every ledger near-miss.
- Draft within 3 business days, review within 5, published internally within 10. The IC owns the draft. The owning EM is accountable.
- One template: timeline, customer and financial impact, why detection was late, why mitigation took that long, contributing factors, what went well, actions.
- Blameless in writing. No individual named as a cause. HR and management commit that postmortems are never used in performance reviews.
- Every action gets a named person, priority, due date, Jira ticket, and verification method. P0 (prevents SEV1 recurrence) due in 30 days and committed into the next sprint before roadmap work. P1 in 60 days. P2 in 90 days.
- Teams reserve 15–20% of sprint capacity for reliability and incident actions.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Weekly Incident Review Board with engineering directors. Target 90% of P0/P1 actions closed on time within two quarters.
18. Training, certification, and commander academy (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Nobody holds independent duty untrained. Training happens in paid working time.
The certification register is an audit artefact.
- All employees, 1 hour: recognise impact, declare, find the channel and status page.
- Responder, half day: severity, escalation, runbooks, comms hygiene. Mandatory before joining a rotation.
- Scribe, 2 hours: timeline discipline. The entry point.
- Incident Commander, two days plus shadowing: command presence, decisions under uncertainty, severity, handover, exec management. Two shadowed incidents and one simulated SEV1 before certification.
- Communications Lead, one day: status writing, customer tiering, legal boundaries, regulator triggers.
- Certification valid 12 months, renewed via simulation.
- New joiners shadow two shifts and do not hold primary in their first 90 days.
- Appoint a champion in each of the 28 teams.
19. Simulations and game days (depends on: 14, 15, 18)
Rehearse before the next real SEV1. Use a standing calendar, not a one-off exercise.
Do not inject uncontrolled changes into the production ledger.
- Monthly 60-minute tabletop per engineering group, using a real incident from the 31.
- Quarterly full-scale game day: regional failover, ledger replica promotion, dependency failure. Whole role structure, timed.
- Twice-yearly unannounced paging drill, including nights, to measure real acknowledgement times.
- One security incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Also exercise status-page failure and loss of the primary chat or pager.
- Use replicas, staging, or tightly governed tests for ledger scenarios.
- Every exercise produces actions in the same tracker as real incidents.
- Complete two cross-company exercises before the SOC 2 audit.
20. Critical-path pilot (depends on: 9, 11, 12, 18, 19)
Prove the process on the highest-risk surface with willing teams before asking 28 teams to adopt it.
Publish a one-page result company-wide. That result is the adoption argument.
- Six-week Wave 0: ledger app, Postgres platform, payments orchestration, Kubernetes/platform, API gateway, Support intake, plus two of the 12 teams already on-call.
- Activate the full stack: severity scale, Duty IC, one pager, page budget, status page, mandatory postmortems, paid on-call.
- Parallel-run old paths for one week, then cut over. The program lead coaches every SEV2+ and does not secretly take command.
- Weekly retro. Expect 20–40 process defects and fix them in the standard before rollout.
- Exit gate: MTTD under 10 minutes on pilot services; IC assigned within 5 minutes in 95% of incidents; page volume down 50%; all postmortems on time; no unpaid pages; sentiment not worse.
21. Metrics, reviews, and error budgets (depends on: 12, 17, 20)
Instrument the process itself. Do not reward hiding incidents or suppressing pages.
Every metric has a target and a named owner. Dashboards are public inside the company.
- Response: MTTD, time to declare, time to IC, MTTA, MTTM, MTTR, incidents by severity, percent customer-first. Report median and p90, by tier, journey, and region.
- Quality: pages per person per week, actionability, budget breaches, postmortem on-time rate, action closure and age, missed acknowledgements.
- Business: credits, 99.95% per capability, error-budget burn, failed-payment volume, reconciliation breaks, repeat-incident rate.
- People: rotation size, frequency, after-hours pages, recovery days, sentiment, attrition among on-call staff.
- Cadence: weekly Incident Review Board; monthly Reliability Review; quarterly exec and board review; annual policy review.
- Error budgets on Tier 0/1 SLOs. Burn too fast and the team pauses features to pay down reliability.
- Reconcile dashboards monthly against a sample of incident records and customer cases so missing incidents cannot hide.
22. Wave rollout with readiness gates (depends on: 20, 21)
Roll out in four waves by criticality, every three weeks. Gates keep the standard credible. A missed gate is rescheduled, not waived.
Finish all 28 teams by week 24 so roughly three months of operating evidence remain before audit fieldwork.
- Wave 1: remaining Tier 0. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team kit: catalog complete, alerts migrated and inside budget, runbooks at the readiness bar, rotation of 6+ certified responders, one IC nominee, one drill passed, paid rotation in payroll.
- Named coach for three weeks. The director signs the gate.
- Freeze legacy paging per wave.
- Publish a live adoption scoreboard.
- Any production service that cannot provide sustainable ownership gets an executive-reviewed deadline, compensating control, and expiry date. No indefinite verbal exceptions.
23. Change management, incentives, and the deal (depends on: 1, 4, 9)
Run this from day one in parallel with design. Engineers will judge fairness. Executives will judge visible results.
Repeat the deal until it is muscle memory: paid on-call, paged only for what you own, trained commander, real sprint capacity for actions.
- Launch via CTO all-hands, per-team roadshows, a one-page laptop card, an internal wiki, and a Slack help channel with a 4-hour answer SLA.
- Put incident-response contribution into promotion criteria. Award best postmortem and biggest noise cut quarterly. Thank people publicly after every SEV1.
- Put adoption, alert hygiene, action closure, and on-call load fairness into every engineering manager's quarterly objectives.
- Write an exception path for engineers who cannot do nights because of caring responsibilities or health, covered by stipended volunteers.
- Prohibit retaliation for good-faith declaration or escalation.
- Pulse-survey at 60 and 120 days. If fairness or load is red, pause expansion until fixed.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
24. SOC 2 evidence, mock audit, and sustainability (depends on: 17, 21, 22)
Evidence is a by-product of doing the work, not a reconstruction the week before the auditor. Protect the process after SOC 2 is signed.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to CC7.2–CC7.4, CC2.2/CC2.3, CC5, and availability A1.2. Confirm interpretation early.
- Maintain versioned, signed policies for incident response, severity, on-call, communications, and postmortems, reviewed annually.
- Automate evidence: incident records, paging and ack logs, status history, postmortem library, action closure, training register, drill records, and reportability decisions including not reportable.
- Internal dry-run at month 6 on a 15-incident evidence walkthrough. Mock audit at month 7. Fix with weeks to spare.
- Assign permanent owners for policy, pager, status page, catalog, training, and metrics.
- Year-two roadmap: follow-the-sun cell, self-healing for the top recurring causes, error-budget release gates, and blast-radius reduction on the shared ledger.
- Quarterly board summary of severe incidents, credits, overdue P0s, and resilience investment so attention does not die after the audit.
Previous Proposal 5 (ID: 0ab2b1c8-46af-4001-ab14-0c074d30a026, Agent: deepseek-v4-pro_refine_5, LLM: deepseek/deepseek-v4-pro):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 6 months of full rollout.
- Customer-first detection falls from 40% to below 10% within 6 months and below 5% within 12.
- Median time to mitigate SEV0/SEV1 falls from 3h10 to under 60 minutes within 12 months; SEV2 under 2 hours.
- Incident Commander assigned and announced within 5 minutes for 95% of SEV0/SEV1 incidents; zero incidents with unclear command beyond 10 minutes.
- Status page updated within policy time for 95% of SEV0/SEV1/SEV2 (10/15/30 minutes).
- Monthly pages fall from 3,400 to under 500 with actionability above 75%.
- All six legacy alerting tools decommissioned by week 16.
- 100% of SEV0/SEV1 incidents have a blameless postmortem published within 10 business days.
- Postmortem action closure rises from 17% to 90% for P0/P1 actions on time.
- SLA credits fall from $1.3M to under $400K in first 12 months.
- Customer-impacting incidents decline to under 15 per year; repeat root causes under 10%.
- 100% of 180 services have a named owning team and criticality tier.
- All 28 teams onboarded by week 24; every Tier 0/1 team has 24x7 primary+secondary coverage with 6+ certified responders.
- At least 30 certified Incident Commanders and 20 certified Communications Leads active.
- Paid on-call policy is approved and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among on-call engineers.
- Quarterly game days and twice-yearly unannounced paging drills executed on schedule, each with tracked action items.
Steps (26):
1. Executive mandate, program governance, and interim incident command
Convert the CEO's outage complaint into a company-level improvement program with a named owner, budget, and deadlines. In the first week, establish an interim command process so no incident remains unowned while the permanent process is designed.
- Appoint Head of Reliability as program owner and CTO/CEO as executive sponsor.
- Form steering group with Engineering, SRE, Support, CS, Legal, Compliance, HR, Finance, Security.
- Approve non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Fund tooling, on-call compensation, training, and 3 dedicated program FTEs; anchor to $1.3M credits.
- Set timeline: design weeks 1-4, tooling and pilot weeks 5-10, rollout weeks 11-20, audit evidence collection from week 8, dry-run month 6.
- Stand up interim 24x7 duty officer and single declaration path within 48 hours, compensated retroactively.
2. Baseline data and alert estate analysis (depends on: 1)
Re-open the last 31 incidents and profile the current alert estate so every later design decision is evidence-based.
- Re-code each incident: detection source, timestamps, owner, severity, credits, root-cause family.
- Quantify customer-first detections and missing signals.
- Analyze two 'nobody in charge' incidents minute-by-minute.
- Inventory six alerting tools: volume, noise, owner, runbook coverage, top 50 noisy rules.
- Freeze baseline metrics: MTTD 22m, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed.
3. Stakeholder listening and resistance mapping (depends on: 1, 2)
Treat engineer pushback as design input. Interview all 28 teams plus Support, CS, and Sales to understand real objections and find early adopters.
- Test objection: unpaid work, night load, unfamiliar code, poor runbooks, or blame.
- Document current informal practices from 12 on-call teams.
- Recruit 10-15 credible engineers as co-design working group.
- Survey baseline sentiment on on-call and alerts.
- Publish promise: paid on-call, page only for owned services, trained commander, reserved reliability capacity.
4. Service ownership catalog and criticality tiering (depends on: 2, 3)
Build a machine-readable service catalog that assigns each of the 180 services a single owning team, escalation policy, and criticality tier.
- Define fields: owner team, manager, Slack channel, escalation policy, dependencies, dashboards, runbooks.
- Tier 0: ledger, money movement, auth, shared PostgreSQL; Tier 1 customer-facing; Tier 2 internal; Tier 3 non-critical.
- Map customer-visible capabilities and dependencies across regions.
- Identify orphan and shared services; force ownership decision within 30 days or schedule decommissioning.
- Publish coverage gaps by tier; Tier 0 gaps become executive escalations.
5. Define severity levels and declaration triggers (depends on: 4)
Adopt a five-level severity model with objective payment-specific triggers, so declaration is a lookup rather than debate. Anyone may declare; only the Incident Commander may downgrade.
- SEV0: unauthorized, lost, duplicated, or corrupted money movement; ledger integrity loss; confirmed data breach; both regions failed. Page all roles, exec, legal, and consider payment pause.
- SEV1: widespread payment failure, no workaround, severe SLA breach, or single-region total loss. Full role activation and public status page.
- SEV2: significant degradation, multiple customers, or workaround available but SLA risk. IC and SMEs paged; status page if customer-visible.
- SEV3: limited impact with workaround; team-led, business-hours response.
- SEV4: internal/no customer impact; ticket only.
- Auto-escalation: unresolved SEV3 >2h becomes SEV2; unresolved SEV2 >1h becomes SEV1; ledger or security always at least SEV2.
- Provide decision tree and 12 worked examples from actual incidents.
6. Define incident roles, decision authority, and handover rules (depends on: 5)
Codify five roles with written responsibilities and explicit authority, so there is never confusion about who is in charge.
- Incident Commander: owns severity, priorities, cross-team pulls, mitigation decisions; does not code.
- Communications Lead: owns status page, internal updates, account manager briefs, regulator coordination.
- Scribe: maintains timeline, decisions, and evidence for postmortem/audit.
- Subject-Matter Responders: diagnose and remediate only their owned services.
- Executive Liaison (SEV0/SEV1): handles exec communications and external stakeholders.
- Rule: IC identified within 5 minutes and announced in channel; handover announced and logged; roles may combine below SEV2, never at SEV0/SEV1.
7. Design 24x7 incident command and comms staffing model (depends on: 6, 4)
Create a central, trained incident command rotation instead of making each of 28 teams field its own commander.
- Recruit 24-30 certified ICs and 12-16 Comms Leads from a cross-team volunteer pool with manager approval.
- Weekly rotations, primary and secondary; 5-minute acknowledgement SLA with auto-escalation.
- Scribe pool as entry-level rotation.
- Night coverage: paid command rotation in New York timezone initially; evaluate follow-the-sun coverage later.
- Eligibility: certification required; commanders can leave with 30 days notice.
- Ensure distinct persons for IC, CL, and primary SME on SEV0/SEV1.
8. Design team on-call rotations and cross-team escalation policies (depends on: 4, 6)
Define tiered team on-call obligations so engineers are only paged for services they own, and cross-team pages go through the IC.
- Tier 0/1 services: 24x7 primary+secondary, at least 6 trained responders per rotation, one-week shifts.
- Tier 2/3: business-hours on-call; after-hours escalation list to engineering manager.
- Platform, infrastructure, database: 24x7 due shared ledger and Kubernetes.
- Routing: every page resolves via service catalog to owning team's escalation policy; cross-team pages only by IC.
- Guardrails: max one primary week in four, no consecutive weeks, no on-call in week after leading SEV0/1, protected recovery time after night work.
9. Design paid on-call compensation and fatigue safeguards (depends on: 8, 3)
Make on-call paid and legally compliant before any new rotation starts, and convert unpaid pager culture into a fair employment condition.
- Weekly stipends: primary 24x7 $800-1,200, secondary 30-50%, business-hours $300-500; holidays premium.
- Out-of-hours incident pay: $150 per night page plus hourly beyond one hour; time-off-in-lieu after overnight work.
- Additional command rotation stipend and SEV0/1 bonus for active responders.
- Verify FLSA/NY wage rules with HR, Legal, Finance; document exemption treatment.
- Base annual budget against current $1.3M credits.
- Track load and trigger staffing review if >2 after-hours pages per person per week sustained.
10. Enforce alert quality standards and noise budget (depends on: 2, 4, 8)
Replace 3,400 monthly alerts at 85% noise with a contractual paging standard that makes on-call sustainable.
- Every paging alert must have owning team, customer impact statement, runbook link, severity mapping, and tested threshold.
- Page only on customer-visible symptoms or SLO burn; cause-based alerts become tickets or dashboards.
- Set page budget: max 2 out-of-hours pages per person per week; breach triggers mandatory alert-tuning sprint.
- Auto-quarantine alerts with >5 firings/month without action or >70% no-action acknowledgements.
- Target <500 actionable pages/month and >75% actionability within 6 months.
- Weekly per-team alert review, monthly cross-team review.
11. Build detection uplift: synthetics, SLOs, and support intake (depends on: 4, 10)
Shift detection from host metrics to customer outcomes so the company stops hearing about outages from clients first.
- Define SLOs per Tier 0/1 capability: payment initiation, auth, settlement timeliness, API availability, ledger consistency.
- Deploy external synthetic transactions from both regions every 60 seconds, covering full payment flow and ledger write.
- Add ledger assurance checks: replication lag, double-entry balance, settlement window countdown.
- Top-100 customer anomaly detection to catch single-tenant outages.
- Auto-create triage incident from support tickets or account manager keywords within 5 minutes.
- Track customer-detected-first as a defect and require a postmortem action.
12. Consolidate alerting and incident tooling (depends on: 5, 8, 10, 11)
Collapse six alerting tools into one integrated paging and incident management platform to create a single system of record for people and audit.
- Select paging/on-call platform and incident management layer (e.g., PagerDuty + incident.io/FireHydrant).
- Implement one-command Slack declaration that auto-creates channel, bridge, pages roles, sets severity, starts timeline.
- Migrate all monitoring sources into the one tool; decommission legacy paging only after two weeks verified.
- Integrate service catalog, status page, Jira action tracking, Salesforce/CS customer lists, and conference bridge.
- Ensure out-of-band paging and offline fallback if a region or chat tool is down.
- Automate evidence capture for SOC2: timestamps, role assignments, severity changes, comms sent.
13. Define acknowledgement and escalation paths (depends on: 12, 6, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts at first signal.
- For SEV0/SEV1: page primary on-call; after 5 minutes unacked page secondary; at 10 page manager and duty IC; at 15 page executive officer.
- For SEV2: primary ack within 10 minutes, IC assigned within 15; escalate on miss.
- For SEV3: ack within 30 minutes or create work item.
- Automatic duty IC page for ledger/security, cross-team, or unresolved ownership.
- If impact unknown after 15 minutes, raise severity.
- Cross-team responders summoned by IC have 10-minute acknowledgement obligation.
- Route alerts with no owner to command rotation, then treat missing ownership as control defect.
- Human acknowledgement required; delivery confirmation is not sufficient.
14. Standardize internal communications (depends on: 6, 5)
Separate the working war room from executive and stakeholder updates, with fixed cadence and pre-approved templates.
- Auto-create incident channel and one read-only broadcast channel for execs, support, sales.
- SEV0: internal update every 15 minutes; SEV1 every 30; SEV2 every 60; SEV3 on state change.
- Template: impact, what we know, what we are doing, ETA/next update, current IC and CL.
- Executives questions through Executive Liaison only; IC is not interrupted.
- Support/CS receives affected-customer list and holding statement within 15 minutes for SEV0/SEV1 and 30 for SEV2.
- Handover protocol for incidents lasting >4 hours: formal IC handover and fatigue check.
15. Standardize customer and status page communications (depends on: 5, 14)
Replace ad-hoc status page updates with a timed, role-owned, template-driven process, including account manager outreach.
- Status page timing: SEV0 initial post within 10 minutes, SEV1 within 15, SEV2 within 30; updates every 15/30/60 minutes until resolved.
- Resolution notice within 30 minutes of mitigation; customer-facing summary in 5 business days for SEV0/SEV1.
- Pre-approve 12-15 templates with Legal and Comms.
- Tiered outreach: top 100 accounts get direct call/email from AM within 30 minutes of SEV0/SEV1; long tail via subscription.
- Use factual language: state impact and next update; never speculate cause or blame vendor.
- Comms Lead is sole author for customer language.
16. Regulatory, legal, and account manager notification playbook (depends on: 5, 15)
Build a notification decision tree and contact matrix so legal/regulatory obligations are assessed early and never forgotten.
- Map obligations: NYDFS Part 500 72-hour cybersecurity event notification, state breach laws, GLBA/FTC, PCI, sponsor bank/card network contractual windows, FinCEN/OFAC if relevant.
- Add regulatory assessment checkpoint for every SEV0 and security SEV1 within 2 hours, even if not reportable.
- Maintain 24x7 contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel with named backups.
- Pre-draft notification templates and test quarterly.
- Encode customer-specific notification SLAs from enterprise contracts into customer tiering.
- Account managers receive legal-approved script and affected-customer list.
17. SLA credit workflow and financial impact measurement (depends on: 5, 15)
Link every incident to money automatically so severity, credits, and prioritization stay consistent and Finance is never surprised.
- Define availability measurement per contract and capability with Legal/Finance.
- Auto-compute affected minutes per customer from incident record and telemetry; generate credit proposal within 5 business days.
- Decide proactive credits for top tier vs claims-based for others; document approval chain.
- Track credits by root-cause family and incident; feed quarterly reliability investment decisions.
- Target reducing annual credits from $1.3M to under $400K in first year.
18. Standardize blameless postmortems (depends on: 5, 6)
Make postmortems mandatory with fixed deadlines, a single format, and a blameless review forum, replacing 'various formats'.
- Mandatory: all SEV0/SEV1; SEV2 with customer impact, credits, >2h, repeat cause, or detected by customer; near-miss involving ledger.
- Draft within 3 business days, peer review within 5, publish within 10.
- One template: timeline, customer/financial impact, detection gap, response gap, contributing factors, what went well.
- Blameless rules: context-based, no individual blame, never used in performance reviews.
- Weekly Incident Review Board reviews all postmortems and challenges quality.
- Publish searchable postmortem library and quarterly recurring causes report.
19. Track postmortem actions with owner and due date (depends on: 18, 12)
Fix the 11/64 completion rate by giving every action the same status as customer commitments, with capacity and escalation.
- Every action gets named owner, priority, due date, and Jira ticket auto-created from postmortem.
- P0 actions prevent SEV0 recurrence, due 30 days; P1 60; P2 90.
- Reserve 15-20% team sprint capacity for incident actions.
- Escalation ladder: manager at 7 days overdue, director at 14, CTO dashboard at 30; P0 overdue blocks features.
- Monthly reporting of closure rate in engineering leadership.
- Target 90% closure of P0/P1 on time in two quarters.
20. Create runbooks and service readiness bar (depends on: 4, 8, 11)
Ensure every service is prepared for 3 a.m. response before it is allowed to page anyone.
- Readiness checklist for Tier 0/1: architecture diagram, dependencies, dashboards, rollback, feature flags, escalation contacts, data-loss statement.
- Write major-incident playbooks: shared PostgreSQL failure, cross-region failover, Kubernetes control plane loss, partner outage, settlement breach, security compromise.
- Prioritize ledger playbooks: documented failover, read-only mode, reconciliation, signed RPO/RTO.
- Runbooks must be tested twice yearly; stale runbooks marked in catalog.
- No paging alerts without readiness sign-off; gap reported to director.
21. Train, certify, and simulate incident response (depends on: 6, 13, 14, 15, 18, 20)
Build a certification path so 24x7 roles are staffed by people who have practiced, and validate the process through drills.
- Scribe course (2 hrs); Responder (half-day); Communications Lead (1 day); Incident Commander (2 days + shadowing/tabletop).
- Certify ICs and CLs; recertify annually.
- New responders shadow two shifts before primary; no primary in first 90 days.
- Monthly tabletop using real past incidents; quarterly game day including regional failover or ledger scenario.
- Twice-yearly unannounced paging drill; one regulator/legal exercise annually.
- Track drill metrics and action items.
22. Pilot on critical services and iterate (depends on: 12, 13, 15, 19, 21)
Prove the process on the highest-risk surface before full rollout. Run a six-week pilot with tight measurement and public results.
- Select 5-6 teams: payments core, ledger/database, platform/Kubernetes, API gateway, plus two existing on-call teams.
- Activate full stack: severity, roles, command rotation, paid on-call, single tooling, alert budget, status page, postmortems.
- Weekly pilot retro; fix defects within 48 hours.
- Exit criteria: MTTD <10 min, IC assigned within 5 min in 95% incidents, pages down 50%, postmortems on time, positive sentiment.
- Publish one-page pilot result as adoption argument.
23. Wave rollout to all teams and decommission legacy paths (depends on: 22, 20)
Roll out to all 28 teams in four waves by criticality, with explicit readiness gates and legacy tool shutdown.
- Wave 1 remaining Tier 0; Wave 2 Tier 1; Wave 3 Tier 2; Wave 4 Tier 3/internal.
- Per-team onboarding: catalog complete, alerts pruned, runbooks ready, rotation staffed with 6+ certified responders, one IC candidate, one drill passed.
- Coach per wave 3 weeks.
- Gate signed by director; failures rescheduled, not waived.
- After onboarding, freeze legacy alerting paths; no fallback to old tools.
- Publish adoption scoreboard.
24. Establish metrics, dashboards, and review cadence (depends on: 12, 19, 22)
Make process performance visible with a small set of metrics and fixed review meetings, so the program is owned by data.
- Response: MTTD, time to IC, MTTA, MTTM, MTTR, % customer-detected-first.
- Quality: page volume per person, alert actionability, postmortem on-time, action closure rate.
- Business: SLA credits, availability vs 99.95, repeat incidents.
- People: on-call load, page per engineer, sentiment, attrition.
- Cadence: weekly Incident Review Board, monthly Reliability Review, quarterly Executive/Board review.
- Dashboards self-serve with targets and named owners.
25. SOC 2 readiness and internal dry-run audit (depends on: 19, 23, 24)
Design evidence as a by-product and test it with an internal walkthrough before the external auditor arrives.
- Map process to SOC2 CC7.3/CC7.4, CC7.2, CC2.2/2.3, CC5, availability criteria.
- Publish versioned policy documents: Incident Response, Severity, On-Call, Communication, Postmortem.
- Automate evidence: incident records with timestamps, role assignments, paging logs, status history, postmortems, action board, training register, drill records.
- Ensure process operates at least 3 months before fieldwork.
- Run internal dry-run month 6; sample 15 incidents; fix gaps with 8 weeks to spare.
- Keep remediation log for process deviations.
26. Continuous improvement, culture, and sustainability (depends on: 23, 25)
Prevent the process from decaying after the audit by embedding review, feedback, and roadmap ownership.
- Quarterly process retrospective with IC pool and responders.
- Re-baseline metrics every six months; raise targets.
- Year-two roadmap: follow-the-sun coverage, self-healing top 3 causes, error budgets gating releases, blast-radius reduction for shared ledger.
- Annual policy review and certification renewal.
- Quarterly on-call sentiment survey with published actions.
- Board quarterly report on availability, credits, and incident trends.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Charter the program and start the audit clock", "description": "Convert the CEO email into a company operating process with one owner, a budget, and an observation window that starts this week.\n\nIncident response is no longer a per-team choice.\n\n- Name the CTO as sponsor and a Head of Reliability as the single accountable owner, with a three-person program office.\n- Form a small decision group: Engineering, SRE, Support, Customer Success, Security, Legal, Finance, and HR. The sponsor decides within 48 hours.\n- Lock non-negotiables: one severity scale, one paging path, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.\n- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.\n- Fund tooling, compensation, training, and reserved engineering capacity against the $1.3M in SLA credits.\n- Timeline: operating floor in 7 days, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, mock audit in month 7, audit in month 8.", "dependencies": []}, {"step_id": "S2", "title": "Install a seven-day operating floor", "description": "Do not wait for new tools or the final policy. Put a crude but real process in place this week so the next outage has a named commander.\n\nThis week is the first audit evidence.\n\n- Publish a one-page interim severity card and one declaration path: Slack command, phone number, and existing pagers, all reaching the same duty person.\n- Staff interim primary and backup Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.\n- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in the incident channel.\n- Tell Support to declare from credible customer reports without waiting for engineering confirmation.\n- Triage the 53 open historical actions. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.\n- Hold a 15-minute daily operations review until the permanent process is live.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Rebuild the forensic baseline", "description": "Rebuild the facts before locking design. This is both the design input and the before picture for the CEO and the auditor.\n\nFreeze the numbers so they cannot drift during design.\n\n- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander assigned, mitigate, and resolve, who led, credits paid, and contributing factors.\n- For each of the 40% customer-first detections, name the specific signal that was missing. That list is the detection backlog.\n- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.\n- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisy rules and every rule with no owner or runbook.\n- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11 of 64 actions closed, on-call in 12 of 28 teams.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Map resistance and publish the fairness contract", "description": "Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.\n\nPublish the deal in writing before any new mandatory pager is assigned.\n\n- Interview all 28 teams plus Support, Customer Success, and Sales in two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, and fear of blame. Each needs a different remedy.\n- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders.\n- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.\n- Publish the **fairness contract**: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.\n- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.\n- No new mandatory night rotation starts until this contract, compensation, and training are live.", "dependencies": ["S1"]}, {"step_id": "S5", "title": "Build the service catalog, ownership, and money-path tiers", "description": "You cannot page the right person across 180 services until every service has one owner. A wrong owner recreates the pager objection.\n\nThe catalog is the single source of truth for paging, impact, status-page components, and audit.\n\n- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.\n- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.\n- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to services, data stores, regions, and third parties, including sponsor banks and processors.\n- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.\n- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform.\n- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.", "dependencies": ["S3"]}, {"step_id": "S6", "title": "Lock severity levels and what each one triggers", "description": "Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by who reported it or how hard the fix looks.\n\nAnyone may declare. Nobody is punished for over-declaring. Only the Incident Commander may downgrade, with the evidence recorded.\n\n- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.\n- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.\n- SEV3 major, contained: narrow or single-customer impact with a workaround. Owning team leads; business-hours comms; postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.\n- SEV4: no customer impact. Ticket only. Never pages.\n- Attach a financial-integrity or security flag to any severity. The flag forces dual-control, Legal, and the regulatory checkpoint without inventing a fifth level.\n- Auto-escalate: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.\n- Lifecycle: Detected → Declared → Triaged → Mitigated (customer impact ends) → Monitoring → Resolved (backlog processed and ledger reconciled) → Reviewed.\n- Publish a decision tree with 12 worked examples taken from the real 31 incidents.", "dependencies": ["S3", "S5"]}, {"step_id": "S7", "title": "Define roles, authority, dual-control, and handover", "description": "Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged.\n\nSeparate coordination from debugging so a commander never needs to know another team's code.\n\n- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.\n- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.\n- Scribe: timestamped record of observations, decisions, owners, and state changes. A human validates for SEV1 and SEV2.\n- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.\n- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.\n- Security, Legal, Compliance, and Vendor Management join on defined triggers.\n- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1 and SEV2. Every handover is announced verbally and in writing with the exact time.\n- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.", "dependencies": ["S6"]}, {"step_id": "S8", "title": "Staff 24x7 with a command corps, not 28 night rotations", "description": "Do not create 28 night rotations. That is what engineers are rejecting.\n\nCentralise coordination. Keep technical ownership local. Page people only for services they own.\n\n- Layer A, Incident Command corps: about 30 certified volunteers from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to seven months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.\n- Layer B, critical-path domains: consolidate 28 teams into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.\n- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.\n- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.\n- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.\n- Evaluate a US-only paid night rotation now. Treat Lisbon or APAC follow-the-sun as a 12-month option, not a year-one dependency.\n- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.", "dependencies": ["S5", "S7"]}, {"step_id": "S9", "title": "Pay on-call, meet New York labor rules, and cap fatigue", "description": "Unpaid on-call in New York is a retention problem and a legal exposure. Pay must be in payroll before any new mandatory night rotation starts.\n\nPublish the numbers. Then ask people to sign up.\n\n- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.\n- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.\n- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.\n- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.\n- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.\n- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.", "dependencies": ["S4", "S8"]}, {"step_id": "S10", "title": "Set the alert-quality bar and a hard page budget", "description": "3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them.\n\nMake alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.\n\n- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.\n- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.\n- Run new alerts in shadow mode for seven days. Test both firing and recovery.\n- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint.\n- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.\n- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.\n- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.", "dependencies": ["S3", "S5"]}, {"step_id": "S11", "title": "Detect payment and ledger failures before customers do", "description": "The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.\n\nEvery customer-first incident becomes a named defect with a tracked action.\n\n- Define SLOs and business SLIs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.\n- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.\n- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.\n- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.\n- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.\n- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.", "dependencies": ["S5", "S10"]}, {"step_id": "S12", "title": "Consolidate to one pager, one incident record, one status page", "description": "Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail.\n\nMigrate without creating a monitoring gap.\n\n- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.\n- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.\n- One Slack command creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.\n- Route every page through the service catalog: service label to owning domain schedule to escalation policy.\n- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.\n- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is down. Test weekly.\n- Set a hard date after which pages outside this tool create no on-call obligation.", "dependencies": ["S6", "S8", "S10"]}, {"step_id": "S13", "title": "Write the five-minute path from signal to command", "description": "Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting.\n\nIf nobody claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.\n\n- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.\n- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.\n- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.\n- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.\n- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.\n- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.", "dependencies": ["S7", "S8", "S12"]}, {"step_id": "S14", "title": "Codify live execution, readiness, and ledger playbooks", "description": "Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving a root cause.\n\nA service must earn the right to page a human at night.\n\n- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.\n- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.\n- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis.\n- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down. No command by direct message.\n- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.\n- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.\n- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date.\n- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.\n- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers.", "dependencies": ["S5", "S7", "S13"]}, {"step_id": "S15", "title": "Run one communications clock for internals, customers, and regulators", "description": "Replace whoever is around with one timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.\n\nState impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.\n\n- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.\n- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.\n- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.\n- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.\n- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.\n- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys. Subscribe all 2,100 customers by default.\n- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.\n- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is not reportable.\n- Obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, cyber-insurer notice.\n- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.", "dependencies": ["S6", "S7", "S12"]}, {"step_id": "S16", "title": "Tie incidents to SLA credits and true financial cost", "description": "Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.\n\nCredit calculation is an output of the incident record, not a negotiation.\n\n- Agree with Legal and Finance the availability measurement method per contract and per component.\n- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.\n- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.\n- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.\n- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.\n- Target credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.", "dependencies": ["S6", "S15"]}, {"step_id": "S17", "title": "Make blameless postmortems mandatory and consistent", "description": "Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth.\n\nThe discipline lives in the deadlines and the review, not the template.\n\n- Mandatory for every SEV1 and SEV2, any incident detected by a customer first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.\n- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.\n- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.\n- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.\n- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.\n- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.\n- Facilitators are trained and are never the commander of the incident being reviewed.", "dependencies": ["S6", "S7"]}, {"step_id": "S18", "title": "Track actions as risk commitments with reserved capacity", "description": "Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.\n\nClosing a ticket without evidence of effectiveness does not close the action.\n\n- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.\n- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.\n- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.\n- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.\n- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.\n- Triage all 53 currently open historical actions within 30 days. Target 90% of high-priority actions closed by due date within two quarters.", "dependencies": ["S12", "S17"]}, {"step_id": "S19", "title": "Train and certify every role before independent duty", "description": "Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.\n\nThe certification register is an audit artefact.\n\n- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.\n- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.\n- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.\n- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.\n- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.\n- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.\n- Certification is valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.", "dependencies": ["S7", "S13", "S15", "S17"]}, {"step_id": "S20", "title": "Rehearse with tabletops, game days, and night drills", "description": "The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.\n\nNever inject uncontrolled change into the production ledger.\n\n- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.\n- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.\n- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.\n- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.\n- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.\n- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.\n- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.\n- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.", "dependencies": ["S12", "S14", "S19"]}, {"step_id": "S21", "title": "Publish Policy v1 and the signed fairness contract", "description": "Collapse the design into a document people will actually open mid-outage, and make it official.\n\nTen pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings.\n\n- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.\n- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.\n- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.\n- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.\n- Link the policy directly from the incident tool.", "dependencies": ["S6", "S7", "S8", "S9", "S10", "S13", "S15", "S17", "S18"]}, {"step_id": "S22", "title": "Pilot on the payments critical path", "description": "Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.\n\nSix weeks, tightly measured, with a public verdict. That verdict is the adoption argument.\n\n- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.\n- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.\n- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.\n- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.\n- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.\n- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse. Publish a one-page result to the whole company.", "dependencies": ["S9", "S11", "S12", "S14", "S19", "S21"]}, {"step_id": "S23", "title": "Instrument metrics, reviews, and anti-gaming", "description": "Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review.\n\nReport median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.\n\n- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.\n- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.\n- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.\n- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.\n- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.\n- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.\n- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.\n- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.", "dependencies": ["S12", "S17", "S22"]}, {"step_id": "S24", "title": "Roll out in risk-ordered waves with readiness gates", "description": "Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.\n\nFinish all 28 teams at least five months before the audit report date so the observation window covers the whole company.\n\n- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.\n- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.\n- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.\n- Legacy alert paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.\n- Run noise burn-down as a quota inside each wave, not as a background hope. Give every team its noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.\n- Pair every suppression with a compensating-detection check. Publish a live adoption scoreboard.\n- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.", "dependencies": ["S22", "S23"]}, {"step_id": "S25", "title": "Produce SOC 2 evidence by operating, then mock-audit", "description": "Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.\n\nDocumented exceptions beat a claim of perfection. Never edit historical records to look clean.\n\n- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.\n- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.\n- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.\n- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.\n- Correct control failures through tracked actions. Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.", "dependencies": ["S21", "S23", "S24"]}, {"step_id": "S26", "title": "Inspect at 90 days and lock year-two ownership", "description": "Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.\n\nName the ways this programme fails and pre-commit the response.\n\n- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.\n- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.\n- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.\n- Contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.\n- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.\n- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.", "dependencies": ["S23", "S24", "S25"]}], "estimated_complexity": "high", "success_metrics": "- A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.\n- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.\n- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.\n- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.\n- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.\n- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.\n- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.\n- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.\n- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.\n- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.\n- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.\n- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.\n- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.\n- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.\n- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.\n- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.\n- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.\n- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.\n- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.\n- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.\n- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders."}Dropped SEV0 for a four-level scale, adopted P1's Layer A/B/C model, metric set and risk register, and added the readiness bar, simulations and a culture step. Still 26 steps, but two of them are now overloaded and the policy-publication step is missing.
- Step 6 replaces SEV0–SEV4 with SEV1–SEV4; the old SEV0 triggered nearly the same response as SEV1 and carried an unrealistic 10-minute status-page target.
- Step 8 states an explicit Layer A/B/C model with named domains, six-responder minimum and platform on-call as safety net rather than permanent owner.
- Adds what it lacked: readiness bar plus playbooks and the ledger blast-radius workstream (10), exercise programme (19), culture and fairness (23), risk register with pre-committed contingencies (25).
- Rotation cap tightened from one primary week in four to one in six (step 9), with the payroll-before-pager gate stated as a hard condition.
- Most honest schedule arithmetic in the round: week-24 completion leaves "~3 months of evidence before audit fieldwork" (step 22).
- Lost its distinctive time-based auto-escalation "unresolved SEV2 >1h becomes SEV1" — nobody else has it and nothing replaced it.
- No policy-publication step: Policy v1, laminated cards and an exception register survive only as evidence bullets in step 24, so the artefact people open mid-outage is never produced.
- Dropped its round-1 vendor candidates (PagerDuty + incident.io/FireHydrant), the only concrete tooling shortlist anyone offered.
- Steps 16 (regulatory + credits) and 17 (postmortems + actions + review board) are overloaded, and step 20's nine dependencies make the pilot a single choke point.
- Keeps "postmortem on request" for SEV3 and has no quota-driven noise burn-down campaign, only the standing alert contract in step 11.
- Proposal 1 : Three-layer coverage: ~30-person command corps, 10–12 domain rotations of six, business-hours Layer C with a manager escalation list.
- Proposal 1 : Service readiness bar, major-incident playbooks and a parallel ledger blast-radius reduction workstream.
- Proposal 1 : Risk register with pre-committed responses to volunteer shortfall, pay delay, tool slip, noise-induced blindness and a SEV1 mid-rollout.
- Proposal 1 : Change-management step with the repeated deal, public closure of the two historical incidents and a pulse-survey expansion gate.
- Proposal 1 : One primary week in six and compensation live in payroll before any mandatory rotation pages a human.
- Proposal 2 : A SEV0 tier for financial or security crises.
- Proposal 1 : Publishing Incident Management Policy v1 with one-page cards and an exception register as its own step.
- Proposal 1 : Quota-driven noise burn-down campaign with leaderboard and automatic downgrade rule.
+ Interim 7-day command+ Listening tour and fairness contract+ Service ownership catalog and criticality tiers+ 24x7 command and responder coverage model+ Service readiness bar and runbooks+ Escalation paths and acknowledgement SLAs+ Standardize internal and customer communications+ Simulations and game days+ Culture, fairness, and continuous feedbackSeven-day interim command bridgeListening tour, resistance map and the on-call dealService catalog, ownership and money-path tiering24x7 coverage model: central command corps, local expertiseEscalation ladder and the five-minute command ruleLive incident execution doctrineInternal communications protocolCustomer communications and status page policyRegulatory, partner and legal notification playbookAction ownership, capacity reservation and enforcementRunbooks, readiness bar and ledger blast-radius reductionExercise programme: tabletops, game days and unannounced drillsPublish Incident Management Policy v1Alert noise burn-down campaignChange management, fairness and pager culture
The plan produced
1. Executive mandate, budget, and governance from P3 step 1
Secure a written CTO/CEO charter making incident management a company operating process, not a team option. Name one accountable Head of Reliability and a small steering group with authority to decide within 48 hours.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking, and protected engineering capacity.
- Approve budget against the $1.3M annual credits: tooling $150-250k/yr, on-call compensation, training, and 2-3 program FTEs.
- Set timeline: interim process in week 1, design weeks 1-6, pilot weeks 7-12, rollout weeks 13-24, mock audit month 7, SOC 2 at month 8.
- Freeze current baselines: 31 incidents, 22-min MTTD, 40% customer-first detection, 3h10 MTTM, $1.3M credits, 3,400 alerts/month at 85% noise, 11/64 actions closed.
2. Interim 7-day command (after 1) new
Put a crude but real process in place immediately so no incident remains unowned while the permanent design proceeds.
- Publish a one-page interim severity card, one declaration path (phone, Slack command, pager), and one incident channel/bridge/timeline naming convention.
- Staff an interim 24x7 duty commander with primary and backup from engineering managers and senior SREs; compensate retroactively under final policy.
- Require a named incident commander within 10 minutes of any suspected major incident, announced in channel.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions in 30 days, prioritizing ledger integrity, duplicate payment, regional failover, security and detection gaps.
- Hold a daily 15-minute operations review until the permanent process is live.
3. Forensic baseline and alert estate analysis (after 1)
Reconstruct the true before picture from the last 12 months; it drives design and serves as audit baseline.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact start/detect/declare/commander/mitigate/resolve, who led, credits paid, contributing factors.
- For every customer-first detection, name the missing signal; this becomes the detection backlog.
- Reconstruct the two
nobody in chargeincidents minute by minute; use them as the burning-platform narrative and test case. - Profile the 3,400 monthly alerts by tool, team, rule, outcome; identify top 50 noisy rules and every rule with no owner or runbook.
- Freeze all baseline metrics in one signed document for the executive and the auditor.
4. Listening tour and fairness contract (after 1, 3) from P4 step 4
Treat engineer pushback as the main delivery risk and convert it into a written deal that removes the objection to carrying a pager for other teams' code.
- Interview all 28 teams plus Support, CS and Sales in two weeks to separate objections: unpaid work, nights, unfamiliar code, missing runbooks, or fear of blame.
- Harvest practices from the 12 teams already on-call; they supply pilot teams and first commanders.
- Publish the deal: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Recruit 10-15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey and repeat at 60, 120 and 365 days.
5. Service ownership catalog and criticality tiers (after 3)
Create a machine-readable catalog as the single source of truth for paging, routing, status page components, and audit evidence.
- Assign one accountable team, engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list per service.
- Tier by business impact: Tier 0 money movement/ledger/auth/shared PostgreSQL; Tier 1 customer-facing degradable; Tier 2 internal/batch; Tier 3 non-critical.
- Map customer journeys (initiate, authorize, settle, reconcile, report, onboard) to services, data stores, regions, third parties.
- Give orphan services an owner within 30 days or a decommission date approved by the sponsor; Tier 0 without an owner is an executive escalation.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
6. Severity scale, declaration rules, and automatic triggers (after 3, 5) from P3 step 5
Adopt one severity scale as a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade with recorded rationale. Start high when uncertain.
- SEV1: money moved wrongly/duplicated/lost, ledger integrity in doubt, confirmed security/data exposure, both regions impaired, payment processing halted. Full role page, bridge in 5 min, status page in 15 min, regulator assessment within 1 hour, mandatory postmortem.
- SEV2: material degradation of a payment journey, settlement window at risk, large/strategic customer fully down, SLA breach likely. Commander and SMEs paged, status page in 30 min, mandatory postmortem.
- SEV3: narrow or single-customer impact with workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any ledger cluster incident, any SEV3 open >2h, any unknown impact after 15 min, any cross-team incident become at least SEV2.
- Publish a decision tree with 12 worked examples from the 31 real incidents.
7. Incident roles, authority, and handoffs (after 6) from P2 step 6
Codify roles so coordination never depends on seniority or heroics. The Incident Commander coordinates; responders fix only services they own.
- IC: owns severity, priorities, roles, cadence, closure; pre-authorized to freeze deploys, rollback, disable features, shift traffic, invoke failover, put ledger read-only, commit spend; does not type in terminals; keeps command when VP joins.
- Communications Lead: sole author for status page, account managers, executives, and hand-off to Legal for regulators; speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, state changes; human validates for SEV1/SEV2.
- Subject-Matter Responders: diagnose and mitigate only their owned services.
- Executive Duty Officer (SEV1): removes obstacles, shields IC from exec questions, owns board/regulator escalation; does not take command unless formally transferred.
- Rule: IC claimed within 5 min and announced; distinct people for IC, Comms, technical lead at SEV1/SEV2; every handover announced verbally and in writing with time.
- Financial controls survive: IC coordinates ledger recovery but cannot bypass dual control, reconciliation or privileged access.
8. 24x7 command and responder coverage model (after 5, 7) new
Do not create 28 fragile night rotations. Centralize coordination in a trained command corps, keep technical ownership local.
- Layer A: incident command corps of ~30 certified volunteers with primary/secondary 24x7; paired Comms Lead pool ~18 and scribe pool as entry.
- Layer B: consolidate 28 teams into 10-12 critical-path domains (ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes, data/reporting, integrations) each with 24x7 primary+secondary and minimum six trained responders.
- Layer C: all other teams business-hours on-call with written after-hours escalation lists held by managers.
- Platform on-call is safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate follow-the-sun only as a 12-month option, not year-1 dependency.
- Acknowledgement discipline: 5 min at SEV1/SEV2 with automatic failover.
9. Paid on-call and fatigue safeguards (after 4, 8)
Unpaid on-call in New York is a retention and legal risk. Pay must be in payroll before any mandatory night rotation starts.
- Indicative: ~$1,000 per primary 24x7 week, secondary ~$400, business-hours ~$250, duty commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond one hour.
- Mandatory recovery: paid recovery day after >2h overnight work, SEV1, or qualifying SEV2; managers arrange coverage.
- HR/Legal/Finance publish amounts, tax treatment, FLSA/NY wage-hour status within 14 days.
- Load rules: no primary more often than 1 in 6, no consecutive primary/secondary weeks, no primary on two rotations, no on-call during PTO, annual cap.
- If a rotation averages >2 out-of-hours pages per person per week for 4 weeks, trigger staffing or alert-remediation review.
- Count on-call as ~15% delivery load; document exemption path for health/caring without career penalty.
10. Service readiness bar and runbooks (after 5, 8) from P2 step 9
A service must earn the right to page a human at 3 a.m. Define a minimum readiness bar and major incident playbooks before any Tier 0/1 service goes live with on-call.
- Require for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switch, escalation contacts, data-loss/latency impact.
- Write playbooks for top failures: shared PostgreSQL failure/corruption, cross-region failover, Kubernetes control-plane loss, processor/sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, duplicate payment.
- Treat ledger cluster as largest structural risk: documented failover, read-only degraded mode, reconciliation after recovery, executive-signed RPO/RTO.
- Start a parallel workstream on blast-radius reduction: tenant/function partitioning, read replicas, isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging at night unless manager accepts a dated exception in writing.
- Runbooks are peer-reviewed, version-controlled, marked stale if not exercised twice a year.
11. Alert quality standard and page budget (after 3, 5)
Make alert quality a condition of paging a human. Burn down noise deliberately rather than by mass silencing.
- Every paging alert must declare owning team, affected service, customer/SLO impact, expected action, dashboard, runbook, dedup key, severity mapping, escalation policy. Anything failing becomes a ticket.
- Page on symptoms of customer harm (SLO burn rate, payment success rate, queue age vs settlement deadlines), not raw CPU/memory.
- Run new alerts in shadow mode 7 days; test firing and recovery.
- Set page budget of 2 out-of-hours pages per person per week; breach blocks new alert creation and triggers tuning sprint.
- Auto-quarantine rules firing >5 times/month without action or >70% no-action acks; return to owner with correction deadline.
- Never silently disable: verify compensating detection, record decision, name owner, review missed detections monthly.
- Target 3,400 to <500 pages/month, actionability >75% within six months, no loss of Tier 0/1 detection.
12. Detection uplift on money path (after 5, 11) from P1 step 11
Stop customers telling you first by detecting on payment outcomes and ledger truth.
- Define SLOs and business SLIs per customer journey: initiation success, authorization latency, settlement timeliness, reconciliation break rate, API availability, reporting freshness; set internal targets stricter than 99.95%.
- Run external synthetic end-to-end payments every 60 seconds from both regions, covering all critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate detection, replication lag, failover-readiness, settlement-window countdown.
- Add per-customer anomaly detection for top 100 accounts; five similar support tickets in 10 minutes auto-creates triage incident.
- Route partner and processor notifications into declaration path within 5 minutes.
- Record detection source on every incident; treat
customer detected firstas a named defect class with mandatory action.
13. Consolidate to one paging and incident platform (after 6, 8, 11, 12)
Collapse six alerting tools into one operational system of record: one queue, one timeline, one audit trail, without creating a monitoring gap.
- Select one paging/scheduling platform, one incident record, one hosted status page; time-box selection to 2 weeks.
- Ingest all six sources first, deduplicate/correlate; retire legacy path only after named owners, successful end-to-end test, and 2 weeks verified operation.
- One-command declaration in Slack creates channel/bridge, pages commander, sets severity, opens timeline, starts clock.
- Route every page through service catalog: service label -> owning domain schedule -> escalation policy.
- Capture evidence automatically: declaration, acks, role assignments, severity changes, decisions, comms, mitigated/resolved times, postmortem link; retain 12 months.
- Test out-of-band SMS/phone paging, mobile fallback, offline runbook weekly; ensure works during one AWS region/chat/identity provider failure.
- Set hard date after which pages outside this tool create no on-call obligation.
14. Escalation paths and acknowledgement SLAs (after 7, 8, 13) new
Write one unskippable path from signal to named commander in under 5 minutes; default action is never waiting.
- Converge all entry points on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately -> secondary at 5 min -> domain manager + Duty Commander at 10 -> Executive Duty Officer at 15. SEV2 requires command assigned within 15 min.
- If no one claims command in 5 min, platform assigns and announces it; assignee may hand over but not decline.
- Commander may page any domain's on-call directly with 10-min ack obligation; this reciprocity makes single-team ownership viable.
- If ownership unclear after 10 min, commander keeps incident and names temporary owner; missing catalog entry logged as control defect.
- Pre-authorize regional failover, ledger read-only, payment suspension, partner-bank notification so IC never waits for an executive.
15. Standardize internal and customer communications (after 6, 7, 13) from P2 step 14
Replace whoever is around with a timed, owned, pre-approved process; the Communications Lead is single author and never writes from scratch under pressure.
- Internal: one working channel/bridge plus one read-only broadcast for executives/Support/Sales; cadence SEV1 15-30 min, SEV2 60 min even if nothing changed; fixed template with impact, action, next update, IC/Comms names.
- Executives ask questions only of Executive Duty Officer; publish this as signed behavior rule.
- Status page: SEV1 within 15 min, SEV2 within 30 min; updates every 30/60 min; resolution notice within 30 min of verified recovery; customer-facing summary within 5 business days for SEV1/qualifying SEV2.
- Pre-approve 12-15 templates with Legal for degradation, settlement delay, API errors, regional failure, data-integrity investigation, security event.
- Top 100 accounts get named account manager call/email within 30 min of SEV1 with briefing pack; all 2,100 subscribed by default.
- Language rules: state impact and next update; never speculate on cause, recovery time, data integrity or blame.
16. Regulatory, partner, and financial impact workflows (after 6, 15) from P1 step 18
In payments some incidents start a legal clock at detection. Build obligation assessment into the process and tie incidents to money.
- Legal/Compliance produce obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC, PCI, money-transmitter, sponsor-bank/card-network windows, cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and security-related SEV2 within 1 hour, recorded even when
not reportable, with decision-maker and evidence. - Maintain 24x7 contact matrix for regulators, sponsor banks, networks, outside counsel, insurer.
- Agree availability measurement method per contract; compute affected minutes per customer from incident record and journey telemetry; propose credit schedule within 5 business days.
- Attribute credits to root-cause families; track failed payment volume, delayed value, reconciliation breaks; target annual credits from $1.3M to under $400k.
17. Blameless postmortem standard and action tracking (after 6, 7) from P4 step 17
Replace various formats with one mandatory format, fixed deadlines, and enforceable actions. Closing a ticket without evidence does not close the action.
- Mandatory for every SEV1/SEV2, customer-first detection, incident >2h, repeat of known cause, credit-generating/contract breach, ledger near-miss.
- Draft within 3 business days, review within 5, publish within 10; IC owns delivery, owning manager accountable.
- One template: summary, customer/financial impact, detection source/gap, timeline, response analysis, contributing conditions, what worked, actions.
- Two mandatory questions: why did a customer see this first? why did mitigation take as long as it did?
- Blameless in writing: systems and context, no individual named as cause, never used in performance reviews; HR handles personnel/misconduct separately.
- Every action gets owner, priority, due date, verification method, ticket auto-created; classes containment 7 days, corrective 30, strategic 90; reserve 15% engineering capacity.
- Escalation: manager at +7 days, director at +14, CTO dashboard at +30; overdue P0 can block related releases.
- Weekly Incident Review Board ratifies severity, challenges quality, monitors actions; target 90% closure on time within two quarters.
18. Train and certify incident roles (after 7, 13, 15)
Command is a skill, not a title. Certify before duty; use paid working time.
- All employees: 30-min module on recognizing impact, declaring, finding channel/status page.
- Responder: half-day on severity, escalation, runbook, financial integrity; mandatory before joining rotation.
- Scribe: 2 hours timeline discipline; entry point.
- Incident Commander: 2 days + shadowing; command presence, delegation, severity calls, running room, handover; certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead: 1 day on status writing, customer tiering, legal boundaries, regulator triggers.
- Domain responders demonstrate dashboard, runbook, rollback, failover, access before primary; two shadow shifts; no primary in first 90 days.
- Certification valid 12 months, renewed by simulation; register is audit artifact.
19. Simulations and game days (after 10, 14, 15, 18) from P4 step 19
Rehearse process before real SEV1; exercise tooling failure and ledger scenarios.
- Monthly tabletop per group reusing an incident from baseline, rotating commander.
- Quarterly game day in staging or tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes degradation.
- Twice-yearly unannounced paging drill to measure real overnight ack times.
- Annually one combined operational/security exercise and one regulator-notification exercise with Legal/CISO.
- Exercise status page down, chat/paging provider unavailable.
- Never inject uncontrolled change into production ledger; validate backups/restore/RPO/RTO on replicas.
- Every exercise produces tracked actions; publish time-to-commander, time-to-first-update, time-to-mitigation.
20. Pilot on critical path (after 9, 11, 12, 13, 14, 15, 17, 18, 19) from P2 step 20
Prove full process on highest-risk surface with willing teams for six weeks before full rollout.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, Support intake, including two of the 12 already on-call teams.
- Activate severity scale, command corps, single paging tool, page budget, status page policy, mandatory postmortems, paid rotations.
- Parallel old paths for one week then cut over; real incidents use new process only.
- Program lead attends every SEV2+ as coach, never shadow commander; review every pilot page within one business day.
- Exit gate: commander named within 5 min in 95% cases, first status update on time, MTTD under 10 min for pilot services, pages halved, no unpaid page, all required postmortems on time, positive sentiment.
- Publish one-page result company-wide.
21. Metrics, dashboards, and review cadence (after 13, 17) from P3 step 23
Instrument the process so improvement is visible and audit sees monitoring/review evidence. Report median and p90, never averages alone.
- Response: time to detect/declare/commander/ack/mitigate/resolve split by severity, tier, journey, region, detection source.
- Quality: customer-first rate, status page timeliness, update cadence, missed escalations, role conflicts, alert actionability, out-of-hours pages per person.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat factors.
- Business: availability per journey, error budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, frequency, recovery days, sentiment, attrition.
- Weekly Incident Review Board; monthly reliability review per team; monthly executive review to CEO; quarterly control review with Security/Compliance/Risk; quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, status page history; team scorecards direct help, never individual penalties.
22. Wave rollout to all 28 teams with readiness gates (after 20, 21) from P1 step 27
Roll out in four waves by customer risk, each passing an explicit gate rather than a date. Complete all teams by week 24 to leave ~3 months of evidence before audit fieldwork.
- Wave 1 remaining Tier 0, Wave 2 Tier 1, Wave 3 Tier 2, Wave 4 Tier 3/internal.
- Per-team onboarding kit: catalog entry complete, alerts migrated within budget, runbooks at readiness bar, rotation staffed with 6 trained responders or time-limited exception, one IC candidate nominated, one tabletop passed, payroll active.
- Named coach per wave for three weeks; director signs gate.
- Legacy paging paths disabled per wave, not kept as comfort fallback.
- Publish live adoption scoreboard.
- Any team unable to staff fair rotation gets headcount, service reassignment, or explicit executive risk acceptance.
23. Culture, fairness, and continuous feedback (after 4, 9) new
Run this in parallel from day one. Engineers judge fairness; executives judge results.
- Repeat the deal in every forum: paid on-call, paged only for owned services, trained commander, real sprint capacity.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
- Weekly office hours first 8 weeks, Slack support channel 4-hour SLA, per-team champion, weekly newsletter with metrics and bad news.
- Incident response contribution in promotion criteria; quarterly awards for best postmortem and biggest noise reduction; public thanks after SEV1.
- Pulse survey at 60, 120, 365 days; if fairness or load red, pause expansion until fixed.
24. SOC 2 evidence design and internal dry-run (after 13, 17, 21, 22) from P3 step 25
Make evidence a by-product of operations; test before external auditor.
- Map with Compliance to Trust Services Criteria CC7.2-CC7.5, CC2.2/CC2.3, A1.2; confirm observation window early.
- Evidence set automated and indexed: versioned policies/exception, rotation schedules, compensation activation, incident records with timestamps/roles, paging/ack logs, status history, reportability decisions, postmortems, action closure with verification, training/certification register, drill records, access reviews.
- Internal Audit tests process in months 4 and 6, sampling incidents end-to-end.
- Run formal mock audit month 7 using same evidence populations; include interviews with IC, engineer, Support, Compliance.
- Correct failures through tracked actions, never by editing history.
- Freeze process wording after month 7; any change logged as exception.
25. Risk register and contingency planning (after 1) from P3 step 26
Name likely failure modes and pre-commit responses; review monthly with sponsor.
- Too few commander volunteers: make command rostered duty for managers and staff engineers until pool reaches 30.
- Compensation not approved in time: fallback to time-off-in-lieu plus phased stipend, never launch mandatory night on-call uncompensated.
- Tool migration slips: cut scope on incident record layer, never on paging consolidation.
- Noise pruning hides real failure: demote to ticket first, observe 30 days, keep recovery path, monthly missed-detection review.
- Burnout in experienced teams: weekly load monitoring with page caps.
- Major SEV1 mid-rollout: program lead becomes full-time responder, wave schedule slips one wave, sponsor told same day.
- Shared ledger concentration: if blast-radius workstream slips, escalate to board as accepted risk with dated plan.
26. Inspect and adapt; year-two sustainability (after 22, 24) from P1 step 32
Prevent decay after audit by revising on data and assigning permanent owners.
- At 90 days live, revise policy using measurements: severity calibration, Layer B/C membership from page data, uncovered shifts, commander burn, missed updates, action closure, survey results.
- Assign permanent owners for policy, paging platform, status page, service catalog, training, metrics.
- Re-baseline targets every six months; shift from lagging to leading metrics: error budget burn, near-miss rate, drill performance.
- Year-two candidates: follow-the-sun coverage, automated mitigation for top recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, completion of ledger blast-radius reduction.
- Keep annual policy review, certification renewal, exercise calendar, board reporting as permanent commitments.
- Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, no loss of Tier 0/1 detection.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% Tier 0/1 services have named owning team, tier, escalation policy, dashboard and runbook by day 30; all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10-12 domain rotations of at least six trained responders.
- 30 certified incident commanders and 18 certified communications leads active.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% required postmortems drafted in 3 business days, reviewed in 5, published in 10 from month 2.
- Postmortem action closure rises from 17% to 90% closed by due date within two quarters; all 53 legacy actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k annualised within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% SEV1/security SEV2 including not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two cross-company exercises completed before audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% sampled incidents; SOC 2 Type II incident response controls pass zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days; no increase in attrition among responders.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 3efb9ff1-28a5-40aa-8f73-9853e91aa095, Agent: opus5_refine_1, LLM: anthropic/claude-opus-5):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (32):
1. Executive mandate, single owner, funding and non-negotiables
Convert the CEO email into a chartered program with one accountable owner and authority over all 28 teams. Incident response becomes a **company operating process**, not a per-team choice.
- Name an executive sponsor (CTO) and one accountable owner (Head of Reliability / Incident Management) with a small permanent office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group with Engineering, SRE/Platform, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours. It is not a 28-person committee.
- Fix the non-negotiables now: one severity scale, one paging tool, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.
- Set the clock deliberately earlier than the audit: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready week 28.
- Approve budget against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises.
2. Seven-day interim command bridge (depends on: 1)
Do not let design work leave the company unprotected for six weeks. Put a crude but real process in place within seven days and improve it later.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster from engineering managers and senior SREs, primary plus backup, 24x7. Pay it retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Tell Support to escalate credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items: complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This becomes both the design input and the "before" picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start / detect / declare / commander-assigned / mitigate / resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. This list drives the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome. Identify the top 50 rules producing most of the noise and every rule with no owner or runbook.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS and Sales in two weeks. Test what the objection actually is: unpaid work, lost sleep, unfamiliar code, missing runbooks, or fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey (fairness, trust in alerts, willingness, burnout) to re-measure at 60, 120 and 365 days.
5. Service catalog, ownership and money-path tiering (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that is the single source of truth for paging, impact and audit.
- One accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not by technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions and third parties, including sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Every orphan service gets an owner in 30 days or a decommission date approved by the sponsor. Tier 0 without an owner is an executive escalation.
6. Severity scale, declaration rules and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- **SEV3 (major, contained):** narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- **SEV4:** no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring; only the commander may downgrade, with the evidence recorded.
- Lifecycle: Detected → Declared → Triaged → **Mitigated** (customer impact ends) → Monitoring → **Resolved** (backlog processed and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Incident roles, authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, roles, cadence and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit spend. Retains command when a VP joins.
- **Communications Lead:** single voice for status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates for SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team; they mitigate, they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance and Vendor Management join on defined triggers.
- Rules: command claimed within 5 minutes, stated in channel ("I am IC"), distinct people for command, comms and technical lead at SEV1/SEV2, and every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation or privileged-access rules.
8. 24x7 coverage model: central command corps, local expertise (depends on: 5, 7)
Do not create 28 night rotations. Centralise coordination in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- **Layer A — Incident Command corps:** ~30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of ~18 from Support, CS and engineering management; a scribe pool used as the training entry point.
- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- **Layer C — everyone else:** business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing, with cost and timeline: a US-only paid night rotation now, a Lisbon or APAC follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
9. On-call compensation, labour compliance and fatigue safeguards (depends on: 8, 4)
Unpaid on-call in New York is both a retention problem and a legal exposure. Pay for it before asking anyone to sign up, and publish the numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 for business-hours rotations, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover rather than expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, tax treatment and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a ~15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
10. Alert quality standard, page budget and noise burn-down (depends on: 3, 5)
3,400 alerts at 85% noise is the reason detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human, and burn the backlog down deliberately rather than by mass silencing.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a **page budget** of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements; return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
- Target 3,400 → under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
11. Detection uplift on the money path (depends on: 10, 5)
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls; five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrating without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack or the identity provider is unavailable. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live incident execution doctrine (depends on: 7, 12, 13)
Give responders one short operating procedure for the first minutes through closure. Priority is limiting customer and financial harm, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; controlled backlog drain; **reconciliation completed before any payment or ledger incident is declared resolved**.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and Customer Success, and reopening if impact recurs.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook and RSS; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack; the long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity or blame; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded **even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review; Legal owns outbound regulatory text, the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
18. SLA credit and financial impact workflow (depends on: 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one, and use the delta as the standing business case for on-call pay and reliability capacity.
19. Blameless postmortem standard and Incident Review Board (depends on: 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first?** and **why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, capacity reservation and enforcement (depends on: 19, 12)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
21. Runbooks, readiness bar and ledger blast-radius reduction (depends on: 5, 8)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the single largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel workstream on blast-radius reduction — tenant or function partitioning, read replicas, and isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (depends on: 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days + shadowing):** command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence for their own domain before independent primary duty; two shadow shifts minimum, never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (depends on: 22, 12, 21)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 11, 12, 21, 22, 9)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect (from first impact), time to declare, time to commander, acknowledgement, mitigate, resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; team scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Wave rollout to all 28 teams with readiness gates (depends on: 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director; failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
28. Alert noise burn-down campaign (depends on: 10, 12, 25)
Run the noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
29. Change management, fairness and pager culture (depends on: 4, 9, 25)
Run this from day one in parallel. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (depends on: 24, 26, 27)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and contingencies (depends on: 1)
Name the ways this programme fails and pre-commit the response. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 26, 27, 30)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 2 (ID: 2e38c07e-bf01-4b70-ba95-0ca1ef39e2d8, Agent: gpt5.6-sol_refine_2, LLM: openai/gpt-5.6-sol):
Estimated Complexity: high
Success Metrics: - By day 7, every suspected SEV0–SEV2 uses one incident record, one coordination channel, and one named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named commander for more than 10 minutes; at least 95% are assigned within 5 minutes.
- By day 30, 100% of Tier 0 services have a named owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, every Tier 0 and Tier 1 customer journey has compensated 24×7 command and subject-matter coverage.
- By day 120, all 180 production services have an owner, response tier, tested escalation path, and appropriate coverage model.
- No mandatory night rotation starts before its compensation, training, access, and staffing controls are active.
- Every critical primary rotation has at least six qualified responders or a documented, expiring executive exception by day 120.
- No responder is routinely scheduled for primary duty more often than one week in six by day 120.
- At least 95% of critical pages are acknowledged within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by day 120 and under 60 minutes by month 8.
- At least 95% of SEV0 and SEV1 customer notices are issued within 15 minutes of declaration by month 3.
- At least 95% of customer-visible SEV2 notices are issued within 30 minutes by month 3.
- At least 95% of incidents meet their required update cadence by month 3.
- Monthly paging volume falls from 3,400 to no more than 1,500 by day 90 and no more than 700 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6 without loss of critical detection coverage.
- All new paging alerts satisfy the owner, impact, action, dashboard, runbook, deduplication, and escalation standard by day 60.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 currently open historical actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date by month 6, with effectiveness evidence.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Two end-to-end cross-company exercises, including regional failure and ledger recovery, are completed before the audit.
- Monthly availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- SLA credits decline by at least 50% on an annualized trailing basis within 12 months.
- Quarterly on-call surveys show improving fairness and sustainability, with at least 75% favorable responses by month 6.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (24):
1. Create the mandate, ownership, and funding
Launch incident management as a company operating program within 48 hours. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable program owner.
- Form a small steering group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Fund two to three implementation staff, paging and incident tooling, observability work, training, exercises, and on-call compensation.
- Reserve 10%–15% of engineering capacity for alert remediation, runbooks, and incident actions.
- Authorize incident commanders to freeze deployments, order rollback, disable features, shift traffic, and invoke continuity plans.
- Preserve financial controls. Incident commanders may coordinate ledger recovery but may not bypass dual approval, privileged-access controls, or reconciliation.
- Set the delivery target: critical controls operational within 60 days, enterprise rollout within 120 days, and a mock audit in month 7.
2. Install an interim process in seven days (depends on: 1)
Do not wait for new tools or the final policy. Put a minimum viable incident process into operation immediately and start collecting evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communications schedule.
- Establish one monitored declaration path through chat, telephone, and the existing paging tools.
- Create a standard incident channel, bridge, incident document, and naming convention.
- Staff temporary primary and backup incident commanders 24×7 from the existing on-call teams and engineering leadership.
- Compensate interim duties retroactively under the final compensation policy.
- Require a named incident commander within 10 minutes for every suspected major incident.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Hold a daily 15-minute control review until the permanent process is live.
3. Build the baseline, ownership catalog, and risk map (depends on: 1)
Establish the facts behind the current failures and assign every production component an owner. Use the resulting catalog as the source for routing, escalation, and audit evidence.
- Reconstruct all 31 customer-impacting incidents, including first impact, detection source, declaration, command assignment, mitigation, resolution, customer communications, and credits.
- Analyze the two incidents with no clear leader and every case detected first by customers.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, databases, queues, regional dependencies, payment processors, and banking partners.
- Assign one accountable team, engineering manager, product owner, primary escalation, secondary escalation, dashboard, and runbook to each service.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 based on financial integrity, customer impact, contractual exposure, and dependency centrality.
- Map critical customer journeys to their application, PostgreSQL, Kubernetes, regional, and third-party dependencies.
- Inventory the six alert sources and all 3,400 monthly alerts by owner, volume, actionability, and duplication.
- Interview representatives from all 28 teams and baseline on-call sentiment, fatigue, and objections.
- Give orphan services an owner or decommissioning decision within 30 days.
4. Design compliance and evidence controls from day one (depends on: 1)
Map the operating process to audit, legal, contractual, and record-retention requirements before finalizing it. Confirm the expected SOC 2 observation period with the auditor immediately.
- Map controls to applicable SOC 2 criteria for monitoring, incident identification, response, recovery, communications, corrective action, access, and availability.
- Define evidence required for declarations, pages, acknowledgements, role assignments, decisions, status updates, postmortems, actions, training, drills, and exceptions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve the Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, and Postmortem Standard.
- Record control exceptions with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing evidence before the audit.
5. Adopt one severity and lifecycle standard (depends on: 2, 3, 4)
Use four impact-based severity levels across operational, security, data, and third-party incidents. Start at the highest credible severity when facts are uncertain, then downgrade with recorded evidence.
- **SEV0 — financial or security crisis:** suspected ledger corruption, unauthorized or duplicated funds movement, material data compromise, both-region loss, or a decision to suspend payment processing. Page all roles immediately; engage executives, Security, Legal, Compliance, and Risk; assess regulatory duties within one hour; use distinct role holders; require a postmortem.
- **SEV1 — critical availability event:** a core payment journey is unavailable, payment failures exceed an initial 10% guardrail for five minutes, regional loss has impaired failover, a settlement deadline is at imminent risk, or rapid error-budget burn makes a material SLA breach likely. Assign all roles; notify internal stakeholders within 10 minutes; publish customer status within 15 minutes; update every 30 minutes; require a postmortem.
- **SEV2 — major bounded event:** approximately 1%–10% of payment attempts fail, a material customer subset or critical customer is down, degradation has a workaround, or contractual impact is likely. Assign an incident commander and responders; add communications and scribe roles for customer impact; publish status within 30 minutes; update every 60 minutes; require a postmortem for customer-visible events.
- **SEV3 — limited event:** localized impact, a safe workaround, and no financial-integrity, security, regulatory, or material contractual risk. The owning team leads; page only if immediate action is necessary; use a ticket otherwise.
- Treat the percentage thresholds as declaration guardrails, not reasons to under-classify integrity, settlement, security, or strategic-customer risk.
- Permit any employee to declare an incident. Only the incident commander may lower severity, with the rationale logged.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of active customer harm. Declare resolution only after stability, backlog recovery, and required ledger reconciliation.
6. Define roles, authority, and handoffs (depends on: 5)
Separate command, communication, recordkeeping, and technical repair. One named person must hold command at every moment of a major incident.
- **Incident Commander:** owns severity, priorities, role assignment, escalation, decision cadence, mitigation coordination, and closure. The commander does not act as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language. Legal or Compliance retains ownership of regulatory submissions.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, and status changes.
- **Subject-Matter Responders:** diagnose and mitigate services for which they have accepted ownership, access, training, and runbooks.
- **Executive Duty Officer:** removes organizational obstacles and approves exceptional business decisions without displacing the incident commander.
- Require separate commander, communications, scribe, and primary technical lead for SEV0 and SEV1.
- For customer-visible SEV2, keep the commander separate from the primary technical responder; communications and scribe may be combined if workload permits.
- Announce every role assignment and transfer in the incident channel. Require verbal and written handoff with current impact, decisions, risks, and next actions.
- Keep executives and account managers out of the technical command path; questions flow through the Executive Duty Officer or Communications Lead.
7. Create sustainable 24×7 coverage across the service estate (depends on: 3, 6)
Use central command coverage and service-domain responder coverage rather than creating 28 fragile night rotations. Engineers remain responsible only for services they own or have formally trained to support.
- Build a company incident-command pool of approximately 18–24 certified senior engineers and managers, with a primary and backup scheduled at all times.
- Build 12–16-person communications and scribe pools using Support, Customer Operations, Engineering Operations, and qualified engineering managers.
- Schedule an Engineering Director or equivalent as the 24×7 Executive Duty Officer.
- Group related service owners into approximately 8–12 coherent responder domains only where members have training, access, and explicit acceptance.
- Require Tier 0 and Tier 1 domains to provide 24×7 primary and secondary responders, normally with at least six trained people in each sustainable rotation.
- Give Tier 2 services business-hours coverage plus a maintained manager escalation path. Treat Tier 3 conditions as tickets unless their impact changes.
- Maintain distinct but coordinated coverage for the ledger application and the PostgreSQL platform.
- Route unknown-owner events to the duty commander and platform triage temporarily. Treat every such event as an ownership control defect.
- If a small team cannot staff a fair rotation, merge coverage only after training or provide headcount, service reassignment, or decommissioning.
8. Approve compensation and fatigue protections (depends on: 7)
End unpaid on-call before expanding mandatory coverage. Publish the policy through HR, Legal, Finance, and Payroll within 14 days.
- Pay fixed weekly stipends for primary and secondary service rotations.
- Pay separate stipends for duty commander, communications, and scribe assignments.
- Provide additional call-out compensation or equivalent paid recovery time for material after-hours work.
- Apply overtime and reporting rules correctly for non-exempt employees under federal and New York requirements.
- Pay higher rates for company holidays and provide a protected recovery day after qualifying overnight work, SEV0 events, or prolonged SEV1 response.
- Target no more than one primary week in six and prohibit simultaneous primary assignments.
- Avoid consecutive primary weeks and make all swaps visible in the paging system.
- Reduce sprint commitments for people carrying primary duty rather than expecting normal delivery capacity.
- Provide a documented accommodation path for health, disability, or caregiving constraints without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
9. Set the service readiness and runbook standard (depends on: 3, 5, 7)
A team cannot respond effectively at night without ownership, access, telemetry, and rehearsed recovery procedures. Apply a formal readiness gate to every Tier 0 and Tier 1 service.
- Require a current architecture diagram, dependency map, dashboards, SLOs, runbooks, rollback method, feature-control method, contacts, and tested escalation path.
- Record RTO, RPO, data-integrity requirements, regional mode, and customer-facing capabilities in the service catalog.
- Link every paging alert to the exact runbook step expected from the responder.
- Prevent new paging alerts for services that fail readiness review. Preserve existing critical detection through a documented exception until a safe replacement exists.
- Write priority playbooks for PostgreSQL failure, ledger-integrity investigation, regional failover, Kubernetes control-plane degradation, payment-processor failure, queue backlog, credential compromise, and payment suspension.
- For the ledger, document read-only or stop-processing modes, failover controls, replay and duplicate protections, backlog handling, and post-recovery reconciliation.
- Require runbook review after material incidents and at least twice per year through exercises.
10. Establish one paging and incident system of record (depends on: 4, 5, 6, 7)
Consolidate paging and incident coordination without requiring an unsafe big-bang replacement of every monitoring system. Monitoring sources may remain specialized, but all human pages must enter one controlled platform.
- Select one enterprise paging platform and one integrated incident record, with chat, telephone, SMS, conference, status-page, ticketing, and service-catalog integrations.
- Initially ingest events from all six current tools, then deduplicate, correlate, and route by service ownership.
- Automate incident channel and bridge creation, role paging, timeline capture, severity changes, communications reminders, and postmortem creation.
- Preserve immutable records of declarations, acknowledgements, role assignments, decisions, and messages.
- Use role-based access, multifactor authentication, break-glass controls, and periodic access reviews.
- Provide telephone and offline fallback procedures for loss of chat, identity, the paging vendor, or an AWS region.
- Test paging and fallback paths weekly.
- Retire a legacy paging route only after its signals have owners, quality review, successful end-to-end tests, and at least two weeks of verified operation in the new path.
11. Enforce alert quality and burn down noise safely (depends on: 3, 10)
Treat paging alerts as production products with owners and quality requirements. Do not reduce noise by silently disabling detection.
- Require every page to identify the service, owner, customer or SLO risk, urgency, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Define an actionable page as one requiring prompt human judgment or intervention that materially reduces customer, financial, security, or contractual risk.
- Route informational, capacity-planning, and non-urgent conditions to dashboards or ticket queues.
- Prefer symptom and error-budget burn alerts over raw CPU, memory, pod, or log-volume thresholds.
- Run new alerts in shadow mode for at least seven days unless an emergency exception is approved.
- Review alerts with less than 50% actionability, more than three firings in seven days, or repeated no-action acknowledgements within two business days.
- Require compensating detection and an owner before suppressing or removing an alert.
- Set a responder load target of no more than two after-hours pages per person per week. A breach creates a mandatory alert-remediation plan.
- Review alert actionability, duplication, missed detection, and page load monthly by domain.
- Prioritize the small number of rules producing most of the current 85% noise.
12. Detect payment and ledger failures before customers (depends on: 3, 11)
Shift detection from infrastructure symptoms to customer journeys and financial outcomes. Set internal objectives stricter than the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, webhooks, and reporting freshness.
- Run external synthetic transactions from outside the production boundary and through both regions at least every minute for critical paths.
- Monitor payment failure rates, latency, queue age, delayed value, settlement-window risk, regional asymmetry, and third-party response quality.
- Add continuous ledger controls for reconciliation breaks, unexpected balances, duplicate identifiers, replication lag, backup health, and failover readiness.
- Add tenant or cohort anomaly detection for high-value customers and common payment methods.
- Convert high-priority Support, account-manager, bank, and processor reports into incident candidates within five minutes.
- Review every customer-first incident as a missed-detection defect and create a corrective action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Codify escalation and live incident execution (depends on: 6, 10)
Create one time-bound path from signal to ownership and mitigation. Delivery of a notification does not count as acknowledgement.
- Page the owning Tier 0 or Tier 1 primary immediately; page the secondary after five unacknowledged minutes; page the domain manager at 10 minutes; escalate to the engineering director at 15 minutes.
- For SEV0–SEV2, page the duty commander immediately. Page the backup after five minutes and require the Engineering Director to assume command if no certified commander owns the event by 10 minutes.
- Automatically involve command for integrity concerns, security concerns, regional events, cross-team impact, critical customer journeys, or unresolved ownership.
- If impact remains materially unknown after 15 minutes, increase response posture rather than waiting for certainty.
- Open one incident channel, bridge, and system record. State severity, known impact, assigned roles, current objective, and next update time.
- Freeze unrelated changes during SEV0 and SEV1 unless the commander records an exception.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require controlled backlog processing and reconciliation before resolving payment or ledger incidents.
- Require a formal command handoff for incidents extending beyond four hours or when fatigue impairs a role holder.
14. Standardize internal and customer communications (depends on: 5, 6, 10)
Communicate known impact early without waiting for root cause. The Communications Lead uses approved facts and always states the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish customer status within 15 minutes when customer-visible and legally safe, then update at least every 30 minutes.
- For SEV1, brief internal stakeholders and publish status within 15 minutes, then update every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send notice within 30 minutes, then update every 60 minutes.
- For SEV3, communicate directly to affected customers only when impact or contract terms require it.
- Give account managers an approved statement and affected-customer list within 30 minutes for SEV0 or SEV1 and within 60 minutes for SEV2.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. Describe affected capabilities, symptoms, workarounds, and the next update.
- Do not speculate about root cause, blame, data integrity, security scope, or recovery time.
- Post a Monitoring update within 30 minutes of mitigation. Post Resolved only after stability and required reconciliation.
- Provide a customer-facing incident summary within five business days for qualifying incidents.
- Record and approve any delay or restriction of public details during an active security threat.
15. Operationalize regulatory, partner, contract, and credit decisions (depends on: 4, 5, 14)
Give Legal and Compliance a timed decision process while keeping technical command with the incident commander. Record both reportable and non-reportable determinations.
- Build a jurisdiction and obligation matrix covering applicable NYDFS requirements, state breach laws, GLBA or FTC requirements, PCI obligations, money-transmitter rules, sponsor banks, payment networks, cyber insurance, and customer contracts.
- Validate applicability and deadlines with counsel rather than assuming every operational incident is reportable.
- Begin reportability assessment immediately for every SEV0, security event, suspected ledger-integrity event, and relevant SEV1.
- Target a documented initial legal assessment within one hour for SEV0 and within two hours for other potentially reportable events.
- Record the decision, evidence, approver, legal deadline, submission owner, and confirmation of delivery.
- Maintain tested 24×7 contacts for regulators, banks, networks, insurers, outside counsel, and critical vendors.
- Encode customer-specific notification clocks and channels in the customer record used by the Communications Lead.
- Have Finance calculate affected minutes, likely credits, and contractual exposure from the incident record within five business days.
- Review whether proactive credits or claims-based handling applies by customer segment and contract.
16. Make postmortems mandatory, consistent, and blameless (depends on: 5, 6, 10)
Use one review standard to learn from incidents and test whether controls worked. Keep learning reviews separate from performance or misconduct processes.
- Require a postmortem for every SEV0 and SEV1.
- Require one for customer-visible SEV2, customer-first detection, incidents lasting more than two hours, contractual breaches, repeat failures, control gaps, and ledger-integrity near misses.
- Produce the factual draft within three business days, conduct the review within five, and publish the approved version within 10.
- Use one template covering summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, lessons, and actions.
- Analyze why detection was not earlier and why mitigation took as long as it did.
- Examine technical, organizational, process, testing, dependency, and incentive factors rather than forcing a single root cause.
- Use a trained facilitator and written blameless rules. Describe decisions in the context and information available at the time.
- Publish broadly useful findings internally while maintaining restricted versions for security, privacy, personnel, or privileged material.
- Review postmortem quality and recurring factors in a weekly Incident Review Board.
17. Enforce action ownership and effectiveness tracking (depends on: 16)
Treat corrective actions as risk commitments rather than suggestions. Closing a ticket is insufficient without evidence that the control or system behavior improved.
- Give every action one named individual owner, manager, priority, due date, expected risk reduction, verification method, and linked work item.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with intermediate milestones.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Place recurrence-prevention actions for SEV0 and SEV1 ahead of discretionary feature work unless an executive accepts the residual risk.
- Escalate overdue high-risk actions to the manager after seven days, director after 14 days, and CTO after 30 days.
- Require written residual-risk acceptance, compensating controls, and a new review date when high-risk work is deferred.
- Verify completed actions through tests, telemetry, drills, or production evidence.
- Triage the 53 currently open historical actions within 30 days. Complete, re-plan, or formally accept the risk, prioritizing ledger, regional, security, and detection items.
18. Measure outcomes, controls, business impact, and human load (depends on: 10, 11, 14, 17)
Use a balanced scorecard so teams are not rewarded for suppressing alerts or avoiding incident declarations. Report medians and 90th percentiles, not averages alone.
- Measure time from first impact to internal detection, declaration, acknowledgement, command assignment, mitigation, and resolution.
- Track customer-first detection, missed escalations, status-page timeliness, update-cadence compliance, and role conflicts.
- Track incident count, recurrence, availability, error-budget burn, affected payment value, delayed transactions, reconciliation breaks, and SLA credits.
- Track page volume, actionability, duplicates, after-hours pages, missed acknowledgements, and load per responder.
- Track postmortem timeliness, action completion, action age, verified effectiveness, and repeated contributing factors.
- Track rotation size, duty frequency, recovery days, swaps, attrition signals, and quarterly responder sentiment.
- Hold a weekly Incident Review Board for incidents, actions, missed controls, and noisy alerts.
- Hold a monthly executive reliability review for trends, investment decisions, contractual exposure, and accepted risks.
- Hold a quarterly resilience and controls review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile dashboard data against customer cases and sampled incident records monthly to identify missing incidents or metric gaming.
19. Train participants and address pager resistance (depends on: 6, 7, 8, 13, 14, 16)
Introduce the operating model as a fair exchange, not an audit mandate. The message is that people are paid, paged only for accepted services, supported by trained command, and given capacity to fix defects.
- Train all employees to recognize impact, declare an incident, and find the incident channel and customer status page.
- Give all 260 engineers role-based training in severity, escalation, evidence preservation, handoff, and financial-integrity precautions.
- Certify incident commanders through instruction, simulation, two shadowed events or exercises, and observed performance.
- Train Communications Leads in status writing, account-manager briefing, contractual clocks, and legal escalation.
- Train scribes in timeline quality, decision capture, fact-versus-hypothesis labeling, and evidence handling.
- Require responders to demonstrate access, dashboards, rollback, runbooks, and escalation competence before primary duty.
- Require two shadow shifts before independent primary on-call.
- Appoint one adoption champion in each team and hold weekly office hours during rollout.
- Publish compensation, fatigue protections, service boundaries, and the accommodation process before assigning shifts.
- Use paid working time for training, exercises, shadowing, runbook work, and certification.
- Survey engineers at baseline, day 60, day 120, and quarterly thereafter.
20. Pilot on the payment critical path (depends on: 8, 9, 11, 12, 13, 14, 17, 19)
Run a four-to-six-week pilot across the highest-risk customer journey. Use real incidents and exercises to correct the process before wider rollout.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, authentication or API edge, settlement, and Support intake.
- Activate paid primary and secondary rotations, central command, communications, the incident record, status templates, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new system authoritative.
- Review every pilot page within one business day for routing, actionability, responder load, and missing context.
- Have the program team coach incidents without silently taking command from assigned role holders.
- Hold a weekly pilot retrospective and correct critical process or tool defects within 48 hours.
- Require 95% command assignment within five minutes, 95% timely communications, no unpaid pages, complete postmortems, and at least a 50% page-noise reduction before expansion.
21. Roll out by customer journey and risk (depends on: 18, 20)
Expand in fixed waves rather than waiting for every team to become perfect. Apply explicit readiness gates and time-limited exceptions.
- Days 0–7: operate the interim declaration, command, and communications process.
- By day 30: complete Tier 0 ownership, activate the first certified command roster, approve compensation, and begin the pilot.
- By day 60: provide compensated 24×7 coverage for every Tier 0 and Tier 1 customer journey and route all critical pages through the new platform.
- By day 90: complete critical detection upgrades, customer communications, regulatory playbooks, and the first major alert-noise reduction.
- By day 120: assign every production service a response tier, owner, tested escalation, and appropriate coverage model.
- Roll out remaining teams in two-to-three-week waves ordered by customer and dependency risk.
- Require each wave to pass ownership, alert, runbook, access, training, compensation, and tabletop gates.
- Disable legacy paging paths after verified cutover rather than leaving ambiguous parallel obligations.
- Publish a weekly adoption dashboard by team and escalate failed gates as business risks.
- Never start mandatory night coverage before compensation, staffing, training, and access are ready.
22. Exercise command, regional resilience, and ledger recovery (depends on: 9, 13, 14, 15, 19)
Validate the process under realistic conditions before depending on it during a crisis. Use the same action-tracking rules for exercises and real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops covering payment failure, customer-first detection, third-party failure, and ambiguous ownership.
- Exercise loss of one AWS region, Kubernetes degradation, PostgreSQL failure, suspected duplicate payments, queue backlog, and settlement risk.
- Exercise simultaneous operational and security events to test command and disclosure boundaries.
- Exercise loss of chat, status-page, identity, or paging providers using telephone and offline fallbacks.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors.
- Validate backups, restoration, RPO, RTO, failover prerequisites, financial controls, and post-recovery reconciliation.
- Avoid uncontrolled production-ledger experiments; use staging, replicas, simulations, or tightly governed production tests.
- Complete at least two cross-company exercises before the audit, including one overnight or unannounced paging test.
23. Test SOC 2 operating effectiveness before the auditor (depends on: 4, 18, 21, 22)
Demonstrate that the controls operate consistently, not merely that policies exist. Correct failures through tracked remediation rather than rewriting historical records.
- Preserve policy approvals, service ownership, schedules, compensation activation, access reviews, training, certifications, incidents, communications, postmortems, actions, and exercises.
- Sample evidence monthly from initial signal through verified action closure.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Use exceptions to document missed acknowledgements, late communications, incomplete records, and compensating controls.
- Conduct a formal mock audit in month 7 using the populations and interview roles expected by the external auditor.
- Trace at least one SEV0 or SEV1, one SEV2, a customer-reported event, and an exercise end to end.
- Remediate evidence and operating gaps before external fieldwork.
- Brief commanders, responders, Support, and Compliance for auditor interviews without scripting inaccurate answers.
24. Institutionalize continuous improvement (depends on: 18, 21, 23)
Keep incident management active after the audit by assigning permanent owners, budgets, and review cycles. Use incident trends to drive architectural investment.
- Assign permanent owners for the policy, service catalog, paging platform, status page, training program, metrics, and evidence repository.
- Review severity thresholds, communications timing, staffing, and compensation annually and after material process failures.
- Recertify commanders and communications leads annually through observed exercises.
- Review recurring failure families quarterly and require executive action when remediation repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fatigue, fairness, tool friction, and psychological safety.
- Report severe incidents, SLA exposure, overdue risks, resilience investments, and customer-first detection to the board or risk committee quarterly.
- Prioritize reduction of the shared-ledger concentration risk, stronger regional independence, deployment safety, graceful degradation, and automated mitigation.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
Previous Proposal 3 (ID: 2437bc94-f488-46b5-866e-01e1d4afbf3b, Agent: qwen3.8-max_refine_3, LLM: alibaba/qwen3.8-max):
Estimated Complexity: high
Success Metrics: - Median time to detect customer-impacting incidents falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median mitigation time falls from 3h10 to under 90 minutes by day 120 and under 60 minutes by month 9.
- A named incident commander is assigned within 5 minutes in at least 95% of SEV1 and SEV2 incidents; zero incidents remain unowned for more than 10 minutes.
- First status-page update occurs within 15 minutes of SEV1 declaration and 30 minutes of SEV2 declaration in at least 95% of qualifying incidents.
- Monthly paging volume falls from 3,400 to under 500 actionable pages within six months, with alert actionability above 80%.
- Out-of-hours pages average no more than 2 per responder per week; sustained breaches trigger mandatory alert remediation.
- Six legacy alerting tools are consolidated into one paging and incident platform, with legacy paging paths disabled by week 16.
- 100% of Tier 0 and Tier 1 services have a named owning team, escalation path, dashboard, and runbook by day 60.
- All 28 teams are onboarded with readiness gates by week 22; every 24x7 critical-path rotation has at least six trained responders.
- 30 or more certified incident commanders and 20 or more certified communications leads provide continuous primary and secondary coverage.
- Paid on-call is approved by HR, legal, and finance and active in payroll before any new mandatory night rotation begins.
- 100% of required SEV1 and SEV2 postmortems are drafted within 3 business days and published within 10 business days in the standard template.
- Postmortem action closure rises from 17% to at least 90% of high-priority actions completed by their due date within two quarters.
- All 53 open historical actions are triaged within 30 days; high-risk unaccepted items are completed or formally risk-accepted within 90 days.
- Repeat incidents from a known unaddressed contributing factor decline by at least 50% within six months.
- Annualized SLA credits fall from $1.3M to under $400k within 12 months.
- Monthly availability meets or exceeds 99.95% by month 6, with exceptions reviewed at the executive reliability meeting.
- At least two cross-company exercises, including regional and ledger scenarios, are completed before the SOC 2 audit, with critical findings tracked.
- The month-6 internal dry-run audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete evidence.
- SOC 2 Type II incident-response controls pass with zero exceptions at the month-8 audit.
- On-call sentiment improves quarter over quarter; fewer than 10% of engineers report unwillingness to participate in their owned-service rotation by month 6.
- No increase in attrition among engineers on active rotations compared with baseline.
Steps (27):
1. Executive mandate, program funding, and governance
Convert the CEO email into a **written mandate within 48 hours**. Name one accountable program owner and a small decision group. Time-box design to five weeks so rollout starts well before the audit.
- Appoint the CTO as executive sponsor and a Head of Reliability or Incident Management as program owner with full-time authority.
- Create an 8–10 person design working group: SRE/platform lead, payments and ledger engineering managers, support lead, compliance, legal, HR, finance, and two rotating engineering managers.
- Approve budget lines: tooling consolidation, on-call compensation, training, coaching, and 2–3 dedicated program staff.
- State the non-negotiables: paid on-call, named service ownership, one severity scale, one paging platform, mandatory postmortems, and protected engineering capacity for reliability actions.
- Fix the master timeline: interim controls week 1, design weeks 1–5, pilot weeks 6–11, full rollout weeks 12–22, internal audit rehearsal month 6, audit-ready month 7.
- Publish a one-page charter to the company that says incident response is a company operating process, not an optional team practice.
2. Evidence baseline from incidents, alerts, and coverage gaps (depends on: 1)
Before changing anything, create an auditable **before picture** from the last 12 months. This baseline drives the severity design, staffing model, and executive reporting.
- Reconstruct all 31 customer-impacting incidents: detection source, first owner, severity, mitigation time, credits paid, and whether command was clear.
- Write a specific case review of the two incidents with unclear ownership for more than one hour.
- Inventory the six alert tools, alert volume by team, noise rate, rules with no owner, and rules with no runbook.
- Map the 16 teams without on-call and the 12 teams with unpaid on-call.
- Review the 53 open postmortem actions and triage the highest-risk items first.
- Freeze baseline metrics: 22-minute detection, 40% customer-first detection, 3h10 mitigation, $1.3M credits, 3,400 alerts, 85% noise, and 17% action closure.
3. Stakeholder listening, resistance mapping, and co-design group (depends on: 1)
Treat engineer pushback as **design input, not an attitude problem**. The objection to carrying a pager for other teams' code must shape ownership, routing, compensation, and command staffing.
- Interview leads from all 28 teams plus support, customer success, sales, compliance, and HR within two weeks.
- Separate the real causes of resistance: unpaid night work, unfamiliar systems, poor runbooks, unfair load, fear of blame, or unclear authority.
- Recruit 10–15 respected engineers and managers as co-designers so the model is built with the teams.
- Survey baseline on-call sentiment, trust in alerts, and psychological safety; repeat at 90, 180, and 365 days.
- Document the central promise: responders are paged for services they own, trained incident commanders coordinate, and all on-call work is compensated.
4. Service ownership catalog, criticality tiers, and dependency map (depends on: 2)
No alert can page the right team until every service has a **named owner**. Build the catalog as the routing foundation for on-call, severity impact mapping, status-page components, and audit evidence.
- Assign one accountable team to each of the 180 services, with an engineering manager, Slack channel, escalation policy, dashboard, and runbook link.
- Tier services: Tier 0 for money movement, ledger integrity, authentication, settlement, and shared PostgreSQL; Tier 1 for customer-facing degradable services; Tier 2 for internal or batch services; Tier 3 for non-critical services.
- Map customer journeys to services, databases, regions, third parties, and contractual SLA components.
- Mark orphan and shared services; require ownership, reassignment, or decommissioning within 30 days.
- Document cross-region dependencies, failover constraints, and services that defeat nominal regional redundancy.
- Treat missing ownership for Tier 0 or Tier 1 services as an executive escalation and a release-blocking risk.
5. Severity scale, declaration rights, and automatic triggers (depends on: 2, 4)
Adopt one severity scale so declaration is a lookup, not a debate. **Anyone may declare; only the incident commander may downgrade.** When uncertain, start higher.
- SEV1: money movement stopped or incorrect, ledger integrity in doubt, confirmed security or data event, both regions impaired, or broad SLA-credit exposure.
- SEV2: major degradation, settlement window at risk, one or more critical customers fully down, or likely contractual breach.
- SEV3: limited impact with a workaround, no financial-integrity or security risk.
- SEV4: no customer impact; handled by ticket during business hours.
- Define fixed triggers for each level: pages, command staffing, bridge call, status-page timing, account-manager outreach, executive notification, and postmortem obligation.
- Add automatic escalation: SEV3 open more than 2 hours becomes SEV2; any incident touching the shared ledger or unclear ownership becomes SEV2 unless the commander documents otherwise.
- Publish a decision tree and 10–12 worked examples from the actual 31 incidents.
6. Incident roles, authority, and handover rules (depends on: 5)
Solve the **nobody-was-in-charge failure** by making command explicit, trained, and transferable. Separate coordination from technical remediation so responders are not asked to debug unfamiliar code.
- Incident Commander: owns severity, priorities, escalation, mitigation strategy, role assignment, handoffs, and closure. Does not write code during the incident. May freeze deploys, invoke pre-approved failover, and pull any on-call responder.
- Communications Lead: owns status page, internal updates, account-manager briefs, and coordination with legal or compliance.
- Scribe: maintains the timestamped record of decisions, actions, role changes, and customer communications.
- Subject-matter responders: engineers from owning teams who diagnose and remediate within their accepted domain.
- Executive liaison: required for SEV1; shields the commander from executive questions and owns board or regulator escalation.
- Require the commander to claim the role within 5 minutes of SEV1 or SEV2 declaration, state it in the incident channel, and record every handover.
- Allow role combination only for SEV3 or SEV4; require separate people for command, communications, and primary technical work at SEV1.
7. Paid on-call, fatigue safeguards, and HR compliance (depends on: 3)
Unpaid on-call is a retention, fairness, and legal risk in New York. **Pay must be approved before any new mandatory rotation starts.** This is the fastest way to reduce pager resistance.
- Create a weekly stipend for primary and secondary on-call, differentiated by tier and night responsibility.
- Add per-page or per-incident pay for out-of-hours activation, plus guaranteed recovery time after night work or major incidents.
- Create a separate stipend for duty incident commanders and communications leads because command is a heavier burden.
- Verify FLSA, New York wage-hour, overtime, holiday, and payroll treatment with HR, legal, and finance.
- Cap rotation frequency: no more than one primary week in four or one week in six depending on staffing; no simultaneous primary assignments.
- Require at least six trained responders for any 24x7 rotation; fund hiring or service reassignment where teams are too small.
- Publish the compensation package and payroll start date before asking engineers to join rotations.
8. On-call staffing model and rotation rules (depends on: 4, 6, 7)
Do not force 28 identical night rotations. Use a **layered model**: central command coverage, critical-path team coverage, and business-hours coverage for lower-tier services.
- Company incident-command and communications rotation: 24x7 pool of 25–35 trained volunteers and designated senior staff, giving primary plus secondary coverage at all times.
- Critical-path teams: 24x7 primary and secondary on-call for Tier 0 and Tier 1 domains, including payments, ledger, authentication, API edge, Kubernetes platform, and shared PostgreSQL.
- Other teams: business-hours on-call with a documented night escalation list owned by the engineering manager.
- Group services into 8–12 coherent domains so rotations are sustainable; no rotation with fewer than six people is approved without an executive exception.
- Require handoff overlap, shadow shifts before first primary duty, and no on-call during approved PTO.
- Define cross-team pull rules: the commander may page another team's on-call with a 10-minute acknowledgement obligation; this is coordinated by command, not pushed onto responders.
- Publish schedules, swap rules, and load limits in the paging tool.
9. Alert quality standard, page budget, and noise controls (depends on: 2, 4)
With 3,400 alerts and 85% noise, detection fails because people stop trusting pages. Make alert quality a **condition of paging anyone**.
- Every page must have a named owning team, customer or SLO impact, severity, dashboard, runbook, expected action, and escalation policy.
- Page on customer-visible symptoms: payment success rate, latency, settlement deadlines, error-budget burn, ledger integrity, and replication health.
- Demote cause-based CPU, memory, or infrastructure-only alerts to dashboards or tickets unless they map to a customer journey.
- Set a page budget: maximum 2 out-of-hours pages per person per week; breach triggers a mandatory alert-tuning sprint.
- Auto-flag alerts that fire repeatedly without action or have high no-action acknowledgement rates.
- Require shadow mode for new alerts before they page humans, except for documented emergencies.
- Review alert quality monthly by team and publish a noise leaderboard.
- Target fewer than 500 actionable pages per month and above 80% actionability within six months.
10. Detection uplift across payments, ledger, and customer signals (depends on: 4, 9)
Customers detected 40% of incidents first. Detection must shift to **payment outcomes, ledger integrity, and inbound customer signals**, not host metrics alone.
- Define SLIs and SLOs for payment initiation, authorization, settlement, reconciliation, refunds, API availability, and reporting freshness.
- Run synthetic end-to-end payment tests from outside the platform in both AWS regions every 60 seconds.
- Add continuous ledger assurance: double-entry reconciliation, replication lag, failover readiness, disk pressure, and settlement-window countdown alerts.
- Monitor per-customer anomalies for top accounts so a single-tenant outage is detected before the account manager calls.
- Convert high-priority support tickets and account-manager reports into incident candidates within 5 minutes.
- Track detection source for every incident; make customer-first detection a reviewed defect with a corrective action.
- Add partner, banking, and card-network notification intake into the same declaration path.
11. Single incident platform and alert-tool consolidation (depends on: 5, 8, 9)
Collapse six tools into **one paging and incident-management system** with one queue, one timeline, and one audit trail. Do not create a parallel audit process.
- Select an integrated stack: paging and on-call scheduling, incident workflow, Slack or chat integration, bridge calling, status-page API, and ticketing integration.
- Implement one-command declaration that creates the incident record, channel, bridge, severity label, role prompts, and clock.
- Migrate alert sources by wave; retire a legacy paging path only after routing, ownership, and acknowledgement tests pass.
- Automate evidence capture: timestamps, acknowledgements, role assignments, severity changes, communications, and postmortem links.
- Integrate with the service catalog, on-call schedules, Jira or equivalent action tracker, customer-success tooling, and the status page.
- Verify the platform works during a single-region failure, including SMS, phone, and offline fallback paths.
- Set a hard date after which pages outside the chosen platform are not valid on-call obligations.
12. Escalation paths and five-minute command rule (depends on: 6, 8, 11)
Create one path from **signal to named commander in under five minutes**, any hour of the day. Make escalation automatic and time-bound.
- Accept declarations from automated alerts, engineers, support, account managers, partners, and customers through the same command.
- For SEV1 and SEV2, page the duty incident commander and owning-team primary immediately.
- Acknowledgement ladder: primary 5 minutes, secondary 10 minutes, manager 15 minutes, director or executive 20 minutes.
- If no commander claims the incident within 5 minutes, the platform assigns and announces one; the assignee may hand over but cannot leave the incident unowned.
- For SEV3, require team acknowledgement within 30 minutes; otherwise create a tracked work item.
- Give the commander pre-approved authority to invoke regional failover, ledger read-only mode, feature kill switches, and partner notifications without waiting for executive sign-off.
- Record every missed acknowledgement and escalation failure for weekly review.
13. Internal communications protocol (depends on: 6, 11)
Separate the **working incident channel from the audience channel** so responders can work and executives, support, and sales stay informed without interrupting the commander.
- Create one incident channel and one bridge per SEV1–SEV3 incident; use a read-only broadcast channel for executives and support.
- Set update cadence: every 15 minutes for SEV1, 30 minutes for SEV2, and at state changes for SEV3.
- Use a fixed update template: impact, customer-visible symptoms, current action, ETA or next update, commander, and communications lead.
- Brief support and customer success with a live affected-customer list and approved holding statements within 15 minutes of SEV1 or SEV2.
- Require executives to route questions through the executive liaison; the commander is not interrupted.
- Define fatigue rules and formal handover for incidents lasting more than 4 hours.
14. Customer status page, account-manager outreach, and SLA credit workflow (depends on: 5, 13)
Replace ad-hoc status updates with a **timed, owned, template-driven customer communication process**. Link incidents to SLA credits so finance and customer success are not surprised.
- Publish status-page updates within 15 minutes of SEV1 declaration and 30 minutes for SEV2; update every 30 or 60 minutes until resolution.
- Assign the communications lead as the single author; use pre-approved templates reviewed by legal and communications.
- Map status-page components to customer capabilities: payments, settlement, reporting, API, onboarding, and regional availability.
- For top accounts, require direct account-manager outreach within 30 minutes of SEV1 with an approved briefing pack.
- Send proactive email or webhook notifications to subscribed customers for SEV1 and SEV2.
- Publish a resolution notice within 30 minutes of mitigation and a customer-facing incident summary within 5 business days for SEV1.
- Automate SLA credit calculation from incident duration and affected capability; review with finance and legal within 5 business days.
- Track credits by incident and root-cause family to guide reliability investment.
15. Regulatory, legal, and partner notification playbook (depends on: 5, 14)
In payments, some incidents are reportable and the clock starts at detection. Make regulatory assessment a **mandatory step in the incident process**, not an afterthought.
- Map obligations: NYDFS cybersecurity-event rules, state breach laws, GLBA safeguards, PCI DSS if in scope, money-transmitter duties, card-network and sponsor-bank contracts, and cyber-insurer notice.
- Require a legal or compliance reportability assessment within 2 hours for every SEV1 and every security-related SEV2, even when the answer is not reportable.
- Maintain a 24x7 contact matrix for regulators, sponsor banks, card networks, outside counsel, insurers, and law enforcement.
- Pre-draft notification templates and preserve legal review.
- Encode customer-contract notification deadlines into account tiers and the communications workflow.
- Record every reportability decision, approver, deadline, and submission confirmation in the incident record.
16. Postmortem standard and blameless review (depends on: 5, 6)
Replace inconsistent postmortems with a **mandatory, blameless, single-format process**. The discipline comes from deadlines, facilitation, and action tracking.
- Require postmortems for all SEV1 and SEV2 incidents, customer-detected incidents, incidents over 2 hours, repeat failures, and ledger near-misses.
- Draft within 3 business days, peer review within 5, publish within 10 for SEV1 and SEV2.
- Use one template: summary, impact, timeline, detection analysis, response analysis, contributing factors, what worked, what failed, and action items.
- Make blamelessness explicit: focus on systems and decisions, not individual fault; never use postmortems in performance discipline.
- Hold a weekly incident review board to review postmortems, ratify severity, and challenge weak actions.
- Maintain a searchable postmortem library and quarterly recurring-cause analysis.
- Require a trained facilitator for major reviews; the incident commander attends but does not facilitate.
17. Action-item tracking, ownership, and delivery gates (depends on: 16, 11)
Only 11 of 64 actions were closed. Give postmortem actions the same status as **customer commitments**, with named owners and visible escalation.
- Create every action as a ticket with one named individual owner, priority, due date, and verification method.
- Use delivery classes: containment within 7 days, corrective work within 30 days, strategic work within 90 days.
- Reserve 15–20% of team sprint capacity for reliability and incident actions.
- Escalate overdue items: manager at 7 days, director at 14 days, CTO dashboard at 30 days.
- Block related feature releases when overdue P0 actions prevent recurrence of a severe incident.
- Require director approval and documented residual risk acceptance for overdue high-risk items.
- Verify effectiveness after completion; closing a ticket without evidence does not close the action.
- Target 90% of high-priority actions completed on time within two quarters.
18. Runbooks, critical-incident playbooks, and readiness bar (depends on: 4, 8)
Poor runbooks are a real cause of pager resistance and slow mitigation. Define a **minimum readiness bar** before a service is allowed to page anyone at night.
- Require for every Tier 0 and Tier 1 service: architecture summary, dependencies, dashboards, alert-to-runbook map, rollback procedure, feature flags, escalation contacts, and customer-impact statement.
- Write major playbooks for shared PostgreSQL ledger failure, regional failover, Kubernetes control-plane loss, payment-processor outage, settlement-window breach, duplicate-payment suspicion, and security compromise.
- Define ledger recovery rules: failover procedure, read-only degraded mode, reconciliation, RPO/RTO, and data-loss tolerance approved by executives.
- Test runbooks in drills at least twice a year; mark untested runbooks stale.
- Prevent paging alerts for services without readiness sign-off unless the engineering manager accepts the gap in writing.
- Keep runbooks linked from every alert and incident template.
19. Training, certification, and role readiness (depends on: 6, 12, 13, 16)
Command and communications are skills. Build a **tiered certification path** so rotations are staffed by people who have practiced, not by whoever is around.
- All employees: 1-hour module on declaring incidents, finding the incident channel, and reading the status page.
- All responders: half-day training on severity, acknowledgement, escalation, runbooks, and evidence hygiene.
- Incident commanders: 2-day course plus two shadowed incidents and one simulation before certification.
- Communications leads: training on status-page writing, customer language, account-manager briefs, and regulatory triggers.
- Scribes: training on timeline discipline and audit evidence.
- Certify for 12 months; renew through a simulation.
- Require shadow shifts before independent primary duty; no new hire holds primary within 90 days.
- Publish the certification register as an audit artifact.
20. Simulation program and game days (depends on: 19, 11, 18)
Rehearse the process before it meets a real SEV1. Simulations build commander confidence, expose runbook gaps, and produce audit evidence.
- Run monthly 60-minute tabletops using real incidents from the 31-incident baseline.
- Run quarterly game days covering regional failover, ledger replica promotion, dependency failure, partner outage, and security event.
- Run twice-yearly unannounced paging drills to measure night acknowledgement times.
- Include support, account managers, legal, compliance, and executives in at least one exercise per quarter.
- Produce tracked action items from every exercise using the same board as real incidents.
- Measure time to commander, time to first status update, and time to mitigation decision.
21. Critical-path pilot and gate review (depends on: 7, 10, 11, 14, 18, 19)
Prove the model on the highest-risk services with willing teams before full rollout. Run a **six-week pilot with daily feedback and public exit criteria**.
- Pilot with payments, ledger/database, platform/Kubernetes, API edge, authentication, and support intake.
- Activate severity scale, duty commanders, paid rotations, single paging platform, alert budget, status-page policy, and postmortem process.
- Hold a weekly pilot retrospective and fix process defects quickly.
- Validate night acknowledgement, cross-team pull response, severity clarity, and compensation payroll.
- Exit gate: commander assigned within 5 minutes in 95% of incidents, status page on time, page noise down at least 50%, postmortems on time, and positive on-call sentiment.
- Publish pilot results to the whole company as the main adoption argument.
22. Wave rollout across all 28 teams (depends on: 21)
Roll out by criticality and dependency, not by calendar alone. Use **readiness gates** so teams are not forced live without coverage.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1 teams. Wave 3: Tier 2 teams. Wave 4: Tier 3 and internal platforms.
- Gate per team: catalog entry complete, alerts migrated, runbooks ready, six-person rotation staffed, one commander candidate nominated, compensation in payroll, and one tabletop passed.
- Assign a program coach to each wave for three weeks.
- Freeze legacy paging paths for each team after successful onboarding.
- Publish a live adoption scoreboard by team.
- Complete all teams by week 22, leaving months of operating evidence before the audit.
23. Metrics, dashboards, and review cadence (depends on: 11, 16, 21)
Instrument the process itself. Leadership must see whether the program is working, and auditors must see **operating evidence, not retrospective paperwork**.
- Track response metrics: MTTD, time to declare, time to commander, acknowledgement time, mitigation time, resolution time, and customer-first detection rate.
- Track quality metrics: page volume, alert actionability, missed pages, postmortem timeliness, action closure rate, and status-page compliance.
- Track business metrics: availability against 99.95%, SLA credits, repeat incidents, and error-budget consumption.
- Track people metrics: on-call load, night pages per person, recovery time usage, sentiment, and attrition signals.
- Hold weekly incident review, monthly reliability review, quarterly executive review, and annual policy review.
- Publish dashboards internally and give every metric a target and owner.
24. Change management, incentives, and culture (depends on: 3, 7, 21)
The process will be judged on fairness. Communicate the deal repeatedly and make participation **recognized, compensated, and safe**.
- Core message: you are paid, paged only for what you own, supported by a trained commander, and given real capacity for actions.
- Run CTO all-hands, team roadshows, office hours, FAQ, and an internal incident-management hub.
- Add incident response and reliability work to promotion criteria and manager objectives.
- Recognize good postmortems, alert-noise reduction, and calm incident leadership.
- Provide a written path for engineers who cannot do nights; cover those shifts with paid volunteers or adjusted staffing.
- Prohibit retaliation for good-faith declaration or escalation.
- Publish sentiment survey results, including bad news, to maintain credibility.
25. SOC 2 evidence design and internal dry-run audit (depends on: 16, 17, 22, 23)
Design the process so audit evidence is a **by-product of normal operations**. Test it internally before the external auditor does.
- Map the process to SOC 2 criteria: incident identification, response, recovery, monitoring, communication, control activities, availability, and corrective action.
- Approve versioned policies: incident response policy, severity standard, on-call policy, communications policy, postmortem standard, and exception process.
- Retain incident records, paging logs, status-page history, postmortems, action tickets, training records, drill records, and access reviews for the audit period.
- Record regulatory reportability decisions even when no notification is required.
- Run an internal dry-run audit at month 6: sample at least 15 incidents and trace evidence end to end.
- Fix gaps with at least four weeks before the audit.
- Maintain an exception and remediation log instead of claiming perfection.
26. Program risk register and contingency planning (depends on: 1)
Name the likely failure modes now and pre-commit responses. Review the register monthly with the executive sponsor.
- Commander volunteer shortfall: roster command duty for senior engineers and managers until the pool reaches 25 trained people.
- Compensation delay: use immediate time-off-in-lieu plus phased stipend, but do not start mandatory night rotations without pay.
- Tool migration slip: preserve single paging and routing scope; defer advanced automation if needed.
- Alert pruning causing missed incidents: move noisy alerts to ticket first, observe 30 days, then delete.
- Burnout in experienced teams: monitor load weekly and cap pages per person.
- Major SEV1 during rollout: program owner shifts to incident support, wave schedule slips one wave, sponsor is informed same day.
- Shared ledger concentration: track blast-radius reduction and failover improvements as top strategic actions.
27. Continuous improvement, maturity roadmap, and post-audit sustainability (depends on: 23, 25)
Prevent the classic failure where the process decays after the audit. Build the second-year operating plan before the first year ends.
- Hold quarterly process retrospectives with commanders, responders, support, and communications leads.
- Re-baseline targets every six months; tighten goals once current targets are met.
- Move from lagging metrics to leading indicators: error-budget burn, near-miss rate, drill performance, and action ageing.
- Plan year-two improvements: follow-the-sun coverage, automated mitigation, ledger blast-radius reduction, error-budget release gates, and per-customer real-time impact reporting.
- Keep annual policy review, certification renewal, drills, and board reporting on a permanent calendar independent of the audit cycle.
- Report quarterly to the board or risk committee on availability, credits, severe incidents, overdue actions, and resilience investment.
Previous Proposal 4 (ID: 185429cf-5cd1-451d-ab7e-34d64ca403df, Agent: grok4.6_refine_4, LLM: xai/grok-4.6):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 10 minutes by month 3 and under 5 minutes by month 9.
- Customer-first detection falls from 40% to under 20% by month 3 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by month 4 and under 60 minutes by month 12.
- A named Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes after week 4.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases by month 3, with required update cadence met in 95% of cases.
- Monthly pages fall from 3,400 to under 500 within 6 months, with actionability above 75%; out-of-hours pages at or under 2 per person per week.
- All six legacy alerting tools route through one paging platform by week 16; legacy paging paths disabled per wave.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by due date within two quarters; all 53 currently open historical actions triaged within 30 days of S2.
- SLA credits fall from $1.3M to under $400k in the first 12 months; customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known unaddressed cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar; every Layer B rotation has 6+ certified responders or a time-limited executive exception.
- Paid on-call policy approved by HR, Legal, and Finance and in payroll before any mandatory night rotation starts.
- All 28 teams onboarded by week 24 with 24x7 command cover from 25+ certified ICs and 15+ certified Comms Leads.
- Internal dry-run at month 6 passes a 15-incident evidence walkthrough; month-7 mock audit finds no unowned high-risk control gap; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment improves quarter over quarter; no increase in attrition among engineers on rotation; pulse at 60 and 120 days shows at least 70% agree rotations are fair and limited to services they own.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked actions; two cross-company exercises completed before the audit.
Steps (24):
1. Program charter, mandate, and funding
Convert the CEO email into a named program with one owner, a budget, and a deadline earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, 28 teams will opt out.
- Appoint a Program Lead (Head of Reliability) with direct CTO and CEO sponsorship.
- Form a small steering group: CTO, VP Eng, Head of CS/Support, CISO, Legal, Finance, HR. Not a 28-team committee.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Timeline: operating floor in 7 days; design weeks 1–6; pilot weeks 7–12; rollout weeks 13–24; mock audit week 28; SOC 2 at month 8.
- Fund tooling ($150–250k/yr), on-call pay (~$600k–1M/yr), 2–3 program FTEs, and reserved engineering capacity. Anchor the ask against $1.3M in credits plus unmeasured incident cost.
- Freeze baselines now: 31 incidents, MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, $1.3M credits, 3,400 alerts/month at 85% noise, 11 of 64 actions closed.
2. Immediate 7-day operating floor (depends on: 1)
Do not wait for tooling, compensation, or the audit. Put a minimum process in place this week so the next outage has a named commander.
Start retaining artifacts on day one. This week becomes the first audit evidence.
- Publish a one-page interim severity guide and a single declaration path through Slack, phone, and pager.
- Staff interim primary and backup Incident Commander 24x7 from existing on-call veterans and engineering managers. Compensate this duty retroactively.
- Require a named IC within 10 minutes for every suspected major incident. If nobody claims it, the duty manager is IC.
- Use one channel, one bridge, one timeline doc, and one naming convention for every major incident.
- Direct Support to escalate credible customer reports immediately. They do not wait for engineering confirmation.
- Triage all 53 open historical actions. Close, re-plan, or formally risk-accept. Ledger integrity, payment duplication, regional resilience, security, and detection come first.
- Hold a daily 15-minute ops review until the permanent process is live.
3. Forensic baseline of incidents and alerts (depends on: 1)
Rebuild the facts before locking design. This is the before picture for the CEO and the auditor.
Every later design choice should trace to this evidence.
- Re-code each of the 31 incidents: trigger, service, detection source, timestamps, who led, credits paid, root-cause family.
- Quantify the 40% customer-first detections and name the missing signal in each case.
- Reconstruct the two nobody-in-charge incidents minute by minute. Use them as the burning-platform story.
- Audit the six alerting tools: volume per tool and team, top 50 noisy rules, rules with no owner or runbook.
- Freeze the baseline numbers. Do not let them drift during design.
4. Listening tour and the fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the main delivery risk. Treat it as a design constraint, not an attitude problem.
The answer that will actually stick is **the deal**: you are paid; you are paged only for services you own; a trained commander runs the room; your postmortem actions get real sprint capacity.
- Interview all 28 teams plus Support, CS, and Sales in two weeks.
- Separate the objections: unpaid work, nights, unfamiliar code, bad runbooks, fear of blame. Each needs a different fix.
- Collect informal practices from the 12 teams already on-call. They are the pilot candidates and the veteran pool.
- Recruit 10–15 credible engineers as a design working group so the process is co-authored.
- Baseline sentiment on alert trust, on-call willingness, and burnout. Re-measure at 6 and 12 months.
5. Service ownership catalog and criticality tiers (depends on: 3)
You cannot page the right person across 180 services until each one has a named owner. This is the foundation of fairness, routing, and audit evidence.
Build a machine-readable catalog as the single source of truth.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, the shared PostgreSQL cluster, queues, partner banks, and customer-facing endpoints.
- Assign one owning team, named engineering manager, Slack channel, escalation policy, and dependency list per service.
- Tier 0: money movement, ledger, auth, shared Postgres, regional control plane. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map Tier 0/1 services to customer capabilities: initiation, authorization, settlement, reporting, onboarding.
- Assign coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
6. Severity scale and declaration rules (depends on: 3, 5)
Replace in-the-moment debate with a lookup. Four payments-specific levels, with worked examples drawn from the 31 real incidents.
Anyone may declare. Only the Incident Commander may downgrade, with recorded rationale. When unsure, start high.
- **SEV1**: money movement stopped or incorrect; ledger integrity in doubt; material security or data exposure; both regions impaired; or more than 10% of customers impacted. Pages IC, comms, scribe, SMEs, and exec liaison. Bridge in 5 minutes. Status page in 15 minutes. Mandatory postmortem and regulator assessment.
- SEV2: severe degradation; settlement window at risk; a strategic customer fully down; SLA breach likely. IC and SMEs paged. Status page in 30 minutes. Mandatory postmortem.
- SEV3: partial impact with a workaround; no credit exposure. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, or a repeat.
- SEV4: internal or minor. Ticket only. No page.
- Auto-escalate: any SEV3 open more than 2 hours, or any incident touching the shared ledger, becomes SEV2. Unknown impact after 15 minutes is raised, not sat on.
- Nobody is punished for over-declaring. Publish that rule in writing and repeat it.
7. Roles, authority, and ledger dual-control (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, transferable, and logged.
The Incident Commander owns the incident, not the fix, and **never types** in a production terminal.
- Incident Commander: declares severity, pulls anyone, freezes deploys, invokes failover, authorises spend. One IC at a time. Assumed within 5 minutes and announced in channel.
- Communications Lead: single voice to customers, status page, account managers, and the exec summary.
- Scribe: timestamped timeline, decisions, and open questions. Feeds the postmortem and the audit trail. Required for SEV1 and SEV2.
- Subject-matter responders: diagnose and mitigate only services they own, with access and runbooks.
- Executive Liaison (SEV1): shields the IC from exec questions; owns regulator and board escalation.
- The IC may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access rules.
- Handover is verbal and written, with time and new owner recorded. Roles may combine below SEV2; never at SEV1. The IC stays in charge if a VP joins.
8. Three-layer 24x7 coverage model (depends on: 5, 7)
Do not put 28 teams on night rotation. That is what engineers are rejecting.
Staff command centrally. Page engineers only for code their team owns.
- Layer A, company command: Duty IC plus secondary, Duty Comms, and a scribe pool. Target 25–35 certified ICs and 15–20 comms leads. Roughly one week per person every 6–8 months. Five-minute ack SLA, then secondary, then on-call director.
- Layer B, Tier 0/1 domains: group 28 teams into 8–12 product and platform domains (ledger app, Postgres platform, payments orchestration, auth, API edge, Kubernetes/platform, settlement). Primary plus secondary 24x7. Minimum six trained people. Target one week in six, never worse than one in four.
- Layer C, Tier 2/3: business-hours on-call. After hours the IC pages the EM, who holds a written escalation list.
- No engineer joins another team's responder pool without training, access, runbooks, shadow shifts, and both teams' acceptance.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner.
- Evaluate follow-the-sun coverage as a 12-month option, not a year-1 dependency.
9. Paid on-call, NY labor compliance, and fatigue rules (depends on: 4, 8)
Unpaid on-call is a retention risk and a New York legal exposure. Pay must be live in payroll before any mandatory night rotation starts.
Treat interrupted nights and recovery time as compensable work.
- Weekly stipend by layer, higher for 24x7 Tier 0/1 and Duty IC, lower for business-hours, with holiday and weekend premiums.
- After-hours call-out pay or TOIL. A paid recovery day after prolonged overnight work, a SEV1, or a qualifying SEV2. Managers cover the next day.
- HR, Legal, and Finance publish dollar amounts, FLSA exempt/non-exempt treatment, NY wage-hour rules, tax treatment, and payroll timing within 14 days of charter.
- Load rules: no primary on two rotations; no consecutive primary weeks; no on-call the week after a SEV1 you commanded.
- A person may declare temporarily unfit after overnight work with no performance penalty.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- If a rotation cannot staff six people, merge domains or hire. Do not run two-person 24x7.
- Model annual cost against $1.3M in credits and get it as a CFO/board line item.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts a month at 85% noise is why detection takes 22 minutes. Make quality a condition of paging a human.
A page is a product with a quality bar, not a dump of host metrics.
- Every paging alert must have a named owner, customer or SLO impact, runbook, tested threshold, severity mapping, dashboard, expected action, and dedup key. Fail any of these and it becomes a ticket or is deleted.
- Page on customer symptoms: SLO burn, error budget, settlement-queue age. Cause-based CPU and memory alerts become dashboards or tickets.
- Set a **page budget** of at most two out-of-hours pages per person per week. A breach triggers a mandatory tuning sprint and blocks new paging alerts for that team.
- Auto-quarantine alerts that fire more than five times a month with no action, or that have more than 70% no-action acks. Never silently disable without compensating detection and a recorded decision.
- Run new alerts in shadow for seven days unless an emergency exception is approved.
- Monthly per-team kill, tune, or keep review. Target under 500 pages a month and actionability above 75% within six months.
11. Customer-journey detection and ledger assurance (depends on: 5, 10)
Stop customers being the monitoring system. Detect on money-path outcomes, not host metrics.
Every postmortem will ask why a customer saw it first.
- Define SLOs per Tier 0/1 capability: initiation success, auth latency, settlement timeliness, API availability, reporting freshness. Set an internal target stricter than 99.95%.
- Run synthetic full-lifecycle payments from outside the platform, in both regions, every 60 seconds. Include a small real-value canary where legally feasible.
- Add ledger assurance: continuous double-entry reconciliation, replication lag, failover-readiness, unexpected balances, duplicate identifiers, settlement-window countdown.
- Add per-customer anomaly detection for the top 100 accounts.
- Auto-create a triage incident within five minutes when support tickets or AM reports match impact keywords.
12. Single incident and paging platform (depends on: 6, 8, 10)
Collapse six alerting tools into one queue, one timeline, and one audit record.
The platform must still work if one AWS region is down. Verify SMS and phone fallback, plus an offline runbook.
- Select one paging product, one incident layer, and one hosted status page.
- One Slack command declares the incident, creates the channel and bridge, pages the Duty IC, sets severity, and starts the clock.
- Ingest the six existing tools first. Deduplicate and route. Retire a legacy path only after named owners and two weeks of verified operation.
- Auto-capture timestamps, acknowledgements, role assignments, severity changes, comms, mitigated time, and resolved time. Export for SOC 2.
- Integrate the service catalog, Jira, Salesforce or CS tooling for affected-customer lists, and Zoom or Slack Huddle.
- Test paging, escalation, status publication, and conference access every week.
- After a team's wave, old paging paths are disabled, not left as a fallback.
13. Detection-to-command escalation path (depends on: 7, 8, 12)
Write the single path from something looks wrong to someone is in charge. Target: a commander in under five minutes, any hour.
If nobody claims IC in five minutes, the platform assigns the Duty IC. The assignee can hand over, not decline.
- Converge every entry point on the same declare command: alert, engineer, support, account manager, partner bank, SEV hotline.
- SEV1: page the owning primary immediately; secondary at 5 minutes unacked; domain manager and Duty IC at 10; exec liaison at 15.
- SEV2: primary ack in 10 minutes; IC assigned in 15.
- Human acknowledgement is required. Delivery to a device does not count.
- The IC can page any team's on-call, with a 10-minute ack obligation. This reciprocity makes single-team ownership viable.
- Unowned alerts go to Layer A command, then the missing owner record is a control defect.
- Pre-authorise regional failover, ledger read-only mode, and partner-bank notice so the IC does not wait for an executive. Dual-control still applies to ledger writes.
14. Live execution and major-incident playbooks (depends on: 7, 13)
Limit customer and financial harm before proving root cause. One procedure from the first minute to handback.
A service cannot page at night until it meets the readiness bar.
- Open channel, bridge, record, and timeline immediately for SEV1 and SEV2. The IC states severity, known impact, hypothesis, objective, roles, and next update time.
- Freeze unrelated production changes during SEV1. Record exceptions the IC approves.
- Prefer reversible mitigation: rollback, feature flag, traffic isolation, rate limit, partner reroute.
- Write playbooks first for Postgres ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank outage, settlement-window breach, suspected security compromise, and suspected duplicate payments.
- Guard against split-brain, replay, and duplication during regional or database recovery. Reconcile and process backlog before calling the incident resolved.
- Mitigation means customer impact ended. Resolution means stable, backlog processed, and ledger reconciled.
- Formal IC handover after 4 hours. Reopen if impact recurs during the stability window.
- Readiness bar per Tier 0/1 service: diagram, dependencies, dashboard, runbook, rollback, kill switch, escalation contacts, RPO/RTO. Untested runbooks are marked stale.
15. Internal, customer, and regulatory communications (depends on: 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved process. The Communications Lead is the single author.
State impact and the next update time. Never speculate on cause.
- Internal: one working channel, one bridge, one read-only broadcast for execs, Support, and Sales. SEV1 update every 30 minutes even if unchanged; SEV2 every 60 minutes.
- Executives ask questions only of the Executive Liaison. Publish this as a signed exec behaviour rule.
- Status page: SEV1 within 15 minutes, SEV2 within 30 minutes; then 30/60-minute updates; resolve notice within 30 minutes of mitigation. Templates pre-approved with Legal.
- Top 100 accounts: named AM contact within 30 minutes of SEV1, with a briefing pack from Comms. Long tail gets the status page plus email or webhook.
- Customer-facing summary within 5 business days for SEV1.
- Mandatory regulatory checkpoint on every SEV1 and every security SEV2, within 2 hours, recorded even when not reportable. Cover NYDFS Part 500 72-hour clock, state breach laws, GLBA/FTC, PCI if in scope, sponsor-bank and card-network windows, FinCEN/OFAC if relevant.
- Legal owns outbound regulatory letters. The IC owns facts. Security or Legal may limit public detail during an active threat, with the reason recorded.
- Encode bespoke customer-contract notice SLAs into account tiering.
16. SLA credit and financial-impact workflow (depends on: 6, 15)
Link incidents to money so severity, credits, and investment stay consistent. Finance should not learn about outages from invoices.
Make credit calculation an output of the incident record, not a negotiation.
- Agree availability measurement per contract and component with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and capability telemetry. Produce a proposed credit schedule within 5 business days of resolution.
- Posture: proactive credits for top-tier accounts; claims-based for the rest. Document the approval chain.
- Track credits per incident and root-cause family. Quarterly report which reliability investments would have prevented which credits.
- Target: cut credits from $1.3M to under $400k in 12 months. Use that delta as the ongoing business case.
17. Blameless postmortems and action tracking (depends on: 6, 12)
Eleven of 64 actions closed is the clearest process failure. Make learning mandatory and actions as binding as customer commitments.
Closing a ticket without evidence of effectiveness does not close the action.
- Mandatory for every SEV1 and SEV2, every customer-first detection, every incident over 2 hours, every repeat of a known cause, and every ledger near-miss.
- Draft within 3 business days, review within 5, published internally within 10. The IC owns the draft. The owning EM is accountable.
- One template: timeline, customer and financial impact, why detection was late, why mitigation took that long, contributing factors, what went well, actions.
- Blameless in writing. No individual named as a cause. HR and management commit that postmortems are never used in performance reviews.
- Every action gets a named person, priority, due date, Jira ticket, and verification method. P0 (prevents SEV1 recurrence) due in 30 days and committed into the next sprint before roadmap work. P1 in 60 days. P2 in 90 days.
- Teams reserve 15–20% of sprint capacity for reliability and incident actions.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Weekly Incident Review Board with engineering directors. Target 90% of P0/P1 actions closed on time within two quarters.
18. Training, certification, and commander academy (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Nobody holds independent duty untrained. Training happens in paid working time.
The certification register is an audit artefact.
- All employees, 1 hour: recognise impact, declare, find the channel and status page.
- Responder, half day: severity, escalation, runbooks, comms hygiene. Mandatory before joining a rotation.
- Scribe, 2 hours: timeline discipline. The entry point.
- Incident Commander, two days plus shadowing: command presence, decisions under uncertainty, severity, handover, exec management. Two shadowed incidents and one simulated SEV1 before certification.
- Communications Lead, one day: status writing, customer tiering, legal boundaries, regulator triggers.
- Certification valid 12 months, renewed via simulation.
- New joiners shadow two shifts and do not hold primary in their first 90 days.
- Appoint a champion in each of the 28 teams.
19. Simulations and game days (depends on: 14, 15, 18)
Rehearse before the next real SEV1. Use a standing calendar, not a one-off exercise.
Do not inject uncontrolled changes into the production ledger.
- Monthly 60-minute tabletop per engineering group, using a real incident from the 31.
- Quarterly full-scale game day: regional failover, ledger replica promotion, dependency failure. Whole role structure, timed.
- Twice-yearly unannounced paging drill, including nights, to measure real acknowledgement times.
- One security incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Also exercise status-page failure and loss of the primary chat or pager.
- Use replicas, staging, or tightly governed tests for ledger scenarios.
- Every exercise produces actions in the same tracker as real incidents.
- Complete two cross-company exercises before the SOC 2 audit.
20. Critical-path pilot (depends on: 9, 11, 12, 18, 19)
Prove the process on the highest-risk surface with willing teams before asking 28 teams to adopt it.
Publish a one-page result company-wide. That result is the adoption argument.
- Six-week Wave 0: ledger app, Postgres platform, payments orchestration, Kubernetes/platform, API gateway, Support intake, plus two of the 12 teams already on-call.
- Activate the full stack: severity scale, Duty IC, one pager, page budget, status page, mandatory postmortems, paid on-call.
- Parallel-run old paths for one week, then cut over. The program lead coaches every SEV2+ and does not secretly take command.
- Weekly retro. Expect 20–40 process defects and fix them in the standard before rollout.
- Exit gate: MTTD under 10 minutes on pilot services; IC assigned within 5 minutes in 95% of incidents; page volume down 50%; all postmortems on time; no unpaid pages; sentiment not worse.
21. Metrics, reviews, and error budgets (depends on: 12, 17, 20)
Instrument the process itself. Do not reward hiding incidents or suppressing pages.
Every metric has a target and a named owner. Dashboards are public inside the company.
- Response: MTTD, time to declare, time to IC, MTTA, MTTM, MTTR, incidents by severity, percent customer-first. Report median and p90, by tier, journey, and region.
- Quality: pages per person per week, actionability, budget breaches, postmortem on-time rate, action closure and age, missed acknowledgements.
- Business: credits, 99.95% per capability, error-budget burn, failed-payment volume, reconciliation breaks, repeat-incident rate.
- People: rotation size, frequency, after-hours pages, recovery days, sentiment, attrition among on-call staff.
- Cadence: weekly Incident Review Board; monthly Reliability Review; quarterly exec and board review; annual policy review.
- Error budgets on Tier 0/1 SLOs. Burn too fast and the team pauses features to pay down reliability.
- Reconcile dashboards monthly against a sample of incident records and customer cases so missing incidents cannot hide.
22. Wave rollout with readiness gates (depends on: 20, 21)
Roll out in four waves by criticality, every three weeks. Gates keep the standard credible. A missed gate is rescheduled, not waived.
Finish all 28 teams by week 24 so roughly three months of operating evidence remain before audit fieldwork.
- Wave 1: remaining Tier 0. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team kit: catalog complete, alerts migrated and inside budget, runbooks at the readiness bar, rotation of 6+ certified responders, one IC nominee, one drill passed, paid rotation in payroll.
- Named coach for three weeks. The director signs the gate.
- Freeze legacy paging per wave.
- Publish a live adoption scoreboard.
- Any production service that cannot provide sustainable ownership gets an executive-reviewed deadline, compensating control, and expiry date. No indefinite verbal exceptions.
23. Change management, incentives, and the deal (depends on: 1, 4, 9)
Run this from day one in parallel with design. Engineers will judge fairness. Executives will judge visible results.
Repeat the deal until it is muscle memory: paid on-call, paged only for what you own, trained commander, real sprint capacity for actions.
- Launch via CTO all-hands, per-team roadshows, a one-page laptop card, an internal wiki, and a Slack help channel with a 4-hour answer SLA.
- Put incident-response contribution into promotion criteria. Award best postmortem and biggest noise cut quarterly. Thank people publicly after every SEV1.
- Put adoption, alert hygiene, action closure, and on-call load fairness into every engineering manager's quarterly objectives.
- Write an exception path for engineers who cannot do nights because of caring responsibilities or health, covered by stipended volunteers.
- Prohibit retaliation for good-faith declaration or escalation.
- Pulse-survey at 60 and 120 days. If fairness or load is red, pause expansion until fixed.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
24. SOC 2 evidence, mock audit, and sustainability (depends on: 17, 21, 22)
Evidence is a by-product of doing the work, not a reconstruction the week before the auditor. Protect the process after SOC 2 is signed.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to CC7.2–CC7.4, CC2.2/CC2.3, CC5, and availability A1.2. Confirm interpretation early.
- Maintain versioned, signed policies for incident response, severity, on-call, communications, and postmortems, reviewed annually.
- Automate evidence: incident records, paging and ack logs, status history, postmortem library, action closure, training register, drill records, and reportability decisions including not reportable.
- Internal dry-run at month 6 on a 15-incident evidence walkthrough. Mock audit at month 7. Fix with weeks to spare.
- Assign permanent owners for policy, pager, status page, catalog, training, and metrics.
- Year-two roadmap: follow-the-sun cell, self-healing for the top recurring causes, error-budget release gates, and blast-radius reduction on the shared ledger.
- Quarterly board summary of severe incidents, credits, overdue P0s, and resilience investment so attention does not die after the audit.
Previous Proposal 5 (ID: 0ab2b1c8-46af-4001-ab14-0c074d30a026, Agent: deepseek-v4-pro_refine_5, LLM: deepseek/deepseek-v4-pro):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 6 months of full rollout.
- Customer-first detection falls from 40% to below 10% within 6 months and below 5% within 12.
- Median time to mitigate SEV0/SEV1 falls from 3h10 to under 60 minutes within 12 months; SEV2 under 2 hours.
- Incident Commander assigned and announced within 5 minutes for 95% of SEV0/SEV1 incidents; zero incidents with unclear command beyond 10 minutes.
- Status page updated within policy time for 95% of SEV0/SEV1/SEV2 (10/15/30 minutes).
- Monthly pages fall from 3,400 to under 500 with actionability above 75%.
- All six legacy alerting tools decommissioned by week 16.
- 100% of SEV0/SEV1 incidents have a blameless postmortem published within 10 business days.
- Postmortem action closure rises from 17% to 90% for P0/P1 actions on time.
- SLA credits fall from $1.3M to under $400K in first 12 months.
- Customer-impacting incidents decline to under 15 per year; repeat root causes under 10%.
- 100% of 180 services have a named owning team and criticality tier.
- All 28 teams onboarded by week 24; every Tier 0/1 team has 24x7 primary+secondary coverage with 6+ certified responders.
- At least 30 certified Incident Commanders and 20 certified Communications Leads active.
- Paid on-call policy is approved and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among on-call engineers.
- Quarterly game days and twice-yearly unannounced paging drills executed on schedule, each with tracked action items.
Steps (26):
1. Executive mandate, program governance, and interim incident command
Convert the CEO's outage complaint into a company-level improvement program with a named owner, budget, and deadlines. In the first week, establish an interim command process so no incident remains unowned while the permanent process is designed.
- Appoint Head of Reliability as program owner and CTO/CEO as executive sponsor.
- Form steering group with Engineering, SRE, Support, CS, Legal, Compliance, HR, Finance, Security.
- Approve non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Fund tooling, on-call compensation, training, and 3 dedicated program FTEs; anchor to $1.3M credits.
- Set timeline: design weeks 1-4, tooling and pilot weeks 5-10, rollout weeks 11-20, audit evidence collection from week 8, dry-run month 6.
- Stand up interim 24x7 duty officer and single declaration path within 48 hours, compensated retroactively.
2. Baseline data and alert estate analysis (depends on: 1)
Re-open the last 31 incidents and profile the current alert estate so every later design decision is evidence-based.
- Re-code each incident: detection source, timestamps, owner, severity, credits, root-cause family.
- Quantify customer-first detections and missing signals.
- Analyze two 'nobody in charge' incidents minute-by-minute.
- Inventory six alerting tools: volume, noise, owner, runbook coverage, top 50 noisy rules.
- Freeze baseline metrics: MTTD 22m, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed.
3. Stakeholder listening and resistance mapping (depends on: 1, 2)
Treat engineer pushback as design input. Interview all 28 teams plus Support, CS, and Sales to understand real objections and find early adopters.
- Test objection: unpaid work, night load, unfamiliar code, poor runbooks, or blame.
- Document current informal practices from 12 on-call teams.
- Recruit 10-15 credible engineers as co-design working group.
- Survey baseline sentiment on on-call and alerts.
- Publish promise: paid on-call, page only for owned services, trained commander, reserved reliability capacity.
4. Service ownership catalog and criticality tiering (depends on: 2, 3)
Build a machine-readable service catalog that assigns each of the 180 services a single owning team, escalation policy, and criticality tier.
- Define fields: owner team, manager, Slack channel, escalation policy, dependencies, dashboards, runbooks.
- Tier 0: ledger, money movement, auth, shared PostgreSQL; Tier 1 customer-facing; Tier 2 internal; Tier 3 non-critical.
- Map customer-visible capabilities and dependencies across regions.
- Identify orphan and shared services; force ownership decision within 30 days or schedule decommissioning.
- Publish coverage gaps by tier; Tier 0 gaps become executive escalations.
5. Define severity levels and declaration triggers (depends on: 4)
Adopt a five-level severity model with objective payment-specific triggers, so declaration is a lookup rather than debate. Anyone may declare; only the Incident Commander may downgrade.
- SEV0: unauthorized, lost, duplicated, or corrupted money movement; ledger integrity loss; confirmed data breach; both regions failed. Page all roles, exec, legal, and consider payment pause.
- SEV1: widespread payment failure, no workaround, severe SLA breach, or single-region total loss. Full role activation and public status page.
- SEV2: significant degradation, multiple customers, or workaround available but SLA risk. IC and SMEs paged; status page if customer-visible.
- SEV3: limited impact with workaround; team-led, business-hours response.
- SEV4: internal/no customer impact; ticket only.
- Auto-escalation: unresolved SEV3 >2h becomes SEV2; unresolved SEV2 >1h becomes SEV1; ledger or security always at least SEV2.
- Provide decision tree and 12 worked examples from actual incidents.
6. Define incident roles, decision authority, and handover rules (depends on: 5)
Codify five roles with written responsibilities and explicit authority, so there is never confusion about who is in charge.
- Incident Commander: owns severity, priorities, cross-team pulls, mitigation decisions; does not code.
- Communications Lead: owns status page, internal updates, account manager briefs, regulator coordination.
- Scribe: maintains timeline, decisions, and evidence for postmortem/audit.
- Subject-Matter Responders: diagnose and remediate only their owned services.
- Executive Liaison (SEV0/SEV1): handles exec communications and external stakeholders.
- Rule: IC identified within 5 minutes and announced in channel; handover announced and logged; roles may combine below SEV2, never at SEV0/SEV1.
7. Design 24x7 incident command and comms staffing model (depends on: 6, 4)
Create a central, trained incident command rotation instead of making each of 28 teams field its own commander.
- Recruit 24-30 certified ICs and 12-16 Comms Leads from a cross-team volunteer pool with manager approval.
- Weekly rotations, primary and secondary; 5-minute acknowledgement SLA with auto-escalation.
- Scribe pool as entry-level rotation.
- Night coverage: paid command rotation in New York timezone initially; evaluate follow-the-sun coverage later.
- Eligibility: certification required; commanders can leave with 30 days notice.
- Ensure distinct persons for IC, CL, and primary SME on SEV0/SEV1.
8. Design team on-call rotations and cross-team escalation policies (depends on: 4, 6)
Define tiered team on-call obligations so engineers are only paged for services they own, and cross-team pages go through the IC.
- Tier 0/1 services: 24x7 primary+secondary, at least 6 trained responders per rotation, one-week shifts.
- Tier 2/3: business-hours on-call; after-hours escalation list to engineering manager.
- Platform, infrastructure, database: 24x7 due shared ledger and Kubernetes.
- Routing: every page resolves via service catalog to owning team's escalation policy; cross-team pages only by IC.
- Guardrails: max one primary week in four, no consecutive weeks, no on-call in week after leading SEV0/1, protected recovery time after night work.
9. Design paid on-call compensation and fatigue safeguards (depends on: 8, 3)
Make on-call paid and legally compliant before any new rotation starts, and convert unpaid pager culture into a fair employment condition.
- Weekly stipends: primary 24x7 $800-1,200, secondary 30-50%, business-hours $300-500; holidays premium.
- Out-of-hours incident pay: $150 per night page plus hourly beyond one hour; time-off-in-lieu after overnight work.
- Additional command rotation stipend and SEV0/1 bonus for active responders.
- Verify FLSA/NY wage rules with HR, Legal, Finance; document exemption treatment.
- Base annual budget against current $1.3M credits.
- Track load and trigger staffing review if >2 after-hours pages per person per week sustained.
10. Enforce alert quality standards and noise budget (depends on: 2, 4, 8)
Replace 3,400 monthly alerts at 85% noise with a contractual paging standard that makes on-call sustainable.
- Every paging alert must have owning team, customer impact statement, runbook link, severity mapping, and tested threshold.
- Page only on customer-visible symptoms or SLO burn; cause-based alerts become tickets or dashboards.
- Set page budget: max 2 out-of-hours pages per person per week; breach triggers mandatory alert-tuning sprint.
- Auto-quarantine alerts with >5 firings/month without action or >70% no-action acknowledgements.
- Target <500 actionable pages/month and >75% actionability within 6 months.
- Weekly per-team alert review, monthly cross-team review.
11. Build detection uplift: synthetics, SLOs, and support intake (depends on: 4, 10)
Shift detection from host metrics to customer outcomes so the company stops hearing about outages from clients first.
- Define SLOs per Tier 0/1 capability: payment initiation, auth, settlement timeliness, API availability, ledger consistency.
- Deploy external synthetic transactions from both regions every 60 seconds, covering full payment flow and ledger write.
- Add ledger assurance checks: replication lag, double-entry balance, settlement window countdown.
- Top-100 customer anomaly detection to catch single-tenant outages.
- Auto-create triage incident from support tickets or account manager keywords within 5 minutes.
- Track customer-detected-first as a defect and require a postmortem action.
12. Consolidate alerting and incident tooling (depends on: 5, 8, 10, 11)
Collapse six alerting tools into one integrated paging and incident management platform to create a single system of record for people and audit.
- Select paging/on-call platform and incident management layer (e.g., PagerDuty + incident.io/FireHydrant).
- Implement one-command Slack declaration that auto-creates channel, bridge, pages roles, sets severity, starts timeline.
- Migrate all monitoring sources into the one tool; decommission legacy paging only after two weeks verified.
- Integrate service catalog, status page, Jira action tracking, Salesforce/CS customer lists, and conference bridge.
- Ensure out-of-band paging and offline fallback if a region or chat tool is down.
- Automate evidence capture for SOC2: timestamps, role assignments, severity changes, comms sent.
13. Define acknowledgement and escalation paths (depends on: 12, 6, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts at first signal.
- For SEV0/SEV1: page primary on-call; after 5 minutes unacked page secondary; at 10 page manager and duty IC; at 15 page executive officer.
- For SEV2: primary ack within 10 minutes, IC assigned within 15; escalate on miss.
- For SEV3: ack within 30 minutes or create work item.
- Automatic duty IC page for ledger/security, cross-team, or unresolved ownership.
- If impact unknown after 15 minutes, raise severity.
- Cross-team responders summoned by IC have 10-minute acknowledgement obligation.
- Route alerts with no owner to command rotation, then treat missing ownership as control defect.
- Human acknowledgement required; delivery confirmation is not sufficient.
14. Standardize internal communications (depends on: 6, 5)
Separate the working war room from executive and stakeholder updates, with fixed cadence and pre-approved templates.
- Auto-create incident channel and one read-only broadcast channel for execs, support, sales.
- SEV0: internal update every 15 minutes; SEV1 every 30; SEV2 every 60; SEV3 on state change.
- Template: impact, what we know, what we are doing, ETA/next update, current IC and CL.
- Executives questions through Executive Liaison only; IC is not interrupted.
- Support/CS receives affected-customer list and holding statement within 15 minutes for SEV0/SEV1 and 30 for SEV2.
- Handover protocol for incidents lasting >4 hours: formal IC handover and fatigue check.
15. Standardize customer and status page communications (depends on: 5, 14)
Replace ad-hoc status page updates with a timed, role-owned, template-driven process, including account manager outreach.
- Status page timing: SEV0 initial post within 10 minutes, SEV1 within 15, SEV2 within 30; updates every 15/30/60 minutes until resolved.
- Resolution notice within 30 minutes of mitigation; customer-facing summary in 5 business days for SEV0/SEV1.
- Pre-approve 12-15 templates with Legal and Comms.
- Tiered outreach: top 100 accounts get direct call/email from AM within 30 minutes of SEV0/SEV1; long tail via subscription.
- Use factual language: state impact and next update; never speculate cause or blame vendor.
- Comms Lead is sole author for customer language.
16. Regulatory, legal, and account manager notification playbook (depends on: 5, 15)
Build a notification decision tree and contact matrix so legal/regulatory obligations are assessed early and never forgotten.
- Map obligations: NYDFS Part 500 72-hour cybersecurity event notification, state breach laws, GLBA/FTC, PCI, sponsor bank/card network contractual windows, FinCEN/OFAC if relevant.
- Add regulatory assessment checkpoint for every SEV0 and security SEV1 within 2 hours, even if not reportable.
- Maintain 24x7 contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel with named backups.
- Pre-draft notification templates and test quarterly.
- Encode customer-specific notification SLAs from enterprise contracts into customer tiering.
- Account managers receive legal-approved script and affected-customer list.
17. SLA credit workflow and financial impact measurement (depends on: 5, 15)
Link every incident to money automatically so severity, credits, and prioritization stay consistent and Finance is never surprised.
- Define availability measurement per contract and capability with Legal/Finance.
- Auto-compute affected minutes per customer from incident record and telemetry; generate credit proposal within 5 business days.
- Decide proactive credits for top tier vs claims-based for others; document approval chain.
- Track credits by root-cause family and incident; feed quarterly reliability investment decisions.
- Target reducing annual credits from $1.3M to under $400K in first year.
18. Standardize blameless postmortems (depends on: 5, 6)
Make postmortems mandatory with fixed deadlines, a single format, and a blameless review forum, replacing 'various formats'.
- Mandatory: all SEV0/SEV1; SEV2 with customer impact, credits, >2h, repeat cause, or detected by customer; near-miss involving ledger.
- Draft within 3 business days, peer review within 5, publish within 10.
- One template: timeline, customer/financial impact, detection gap, response gap, contributing factors, what went well.
- Blameless rules: context-based, no individual blame, never used in performance reviews.
- Weekly Incident Review Board reviews all postmortems and challenges quality.
- Publish searchable postmortem library and quarterly recurring causes report.
19. Track postmortem actions with owner and due date (depends on: 18, 12)
Fix the 11/64 completion rate by giving every action the same status as customer commitments, with capacity and escalation.
- Every action gets named owner, priority, due date, and Jira ticket auto-created from postmortem.
- P0 actions prevent SEV0 recurrence, due 30 days; P1 60; P2 90.
- Reserve 15-20% team sprint capacity for incident actions.
- Escalation ladder: manager at 7 days overdue, director at 14, CTO dashboard at 30; P0 overdue blocks features.
- Monthly reporting of closure rate in engineering leadership.
- Target 90% closure of P0/P1 on time in two quarters.
20. Create runbooks and service readiness bar (depends on: 4, 8, 11)
Ensure every service is prepared for 3 a.m. response before it is allowed to page anyone.
- Readiness checklist for Tier 0/1: architecture diagram, dependencies, dashboards, rollback, feature flags, escalation contacts, data-loss statement.
- Write major-incident playbooks: shared PostgreSQL failure, cross-region failover, Kubernetes control plane loss, partner outage, settlement breach, security compromise.
- Prioritize ledger playbooks: documented failover, read-only mode, reconciliation, signed RPO/RTO.
- Runbooks must be tested twice yearly; stale runbooks marked in catalog.
- No paging alerts without readiness sign-off; gap reported to director.
21. Train, certify, and simulate incident response (depends on: 6, 13, 14, 15, 18, 20)
Build a certification path so 24x7 roles are staffed by people who have practiced, and validate the process through drills.
- Scribe course (2 hrs); Responder (half-day); Communications Lead (1 day); Incident Commander (2 days + shadowing/tabletop).
- Certify ICs and CLs; recertify annually.
- New responders shadow two shifts before primary; no primary in first 90 days.
- Monthly tabletop using real past incidents; quarterly game day including regional failover or ledger scenario.
- Twice-yearly unannounced paging drill; one regulator/legal exercise annually.
- Track drill metrics and action items.
22. Pilot on critical services and iterate (depends on: 12, 13, 15, 19, 21)
Prove the process on the highest-risk surface before full rollout. Run a six-week pilot with tight measurement and public results.
- Select 5-6 teams: payments core, ledger/database, platform/Kubernetes, API gateway, plus two existing on-call teams.
- Activate full stack: severity, roles, command rotation, paid on-call, single tooling, alert budget, status page, postmortems.
- Weekly pilot retro; fix defects within 48 hours.
- Exit criteria: MTTD <10 min, IC assigned within 5 min in 95% incidents, pages down 50%, postmortems on time, positive sentiment.
- Publish one-page pilot result as adoption argument.
23. Wave rollout to all teams and decommission legacy paths (depends on: 22, 20)
Roll out to all 28 teams in four waves by criticality, with explicit readiness gates and legacy tool shutdown.
- Wave 1 remaining Tier 0; Wave 2 Tier 1; Wave 3 Tier 2; Wave 4 Tier 3/internal.
- Per-team onboarding: catalog complete, alerts pruned, runbooks ready, rotation staffed with 6+ certified responders, one IC candidate, one drill passed.
- Coach per wave 3 weeks.
- Gate signed by director; failures rescheduled, not waived.
- After onboarding, freeze legacy alerting paths; no fallback to old tools.
- Publish adoption scoreboard.
24. Establish metrics, dashboards, and review cadence (depends on: 12, 19, 22)
Make process performance visible with a small set of metrics and fixed review meetings, so the program is owned by data.
- Response: MTTD, time to IC, MTTA, MTTM, MTTR, % customer-detected-first.
- Quality: page volume per person, alert actionability, postmortem on-time, action closure rate.
- Business: SLA credits, availability vs 99.95, repeat incidents.
- People: on-call load, page per engineer, sentiment, attrition.
- Cadence: weekly Incident Review Board, monthly Reliability Review, quarterly Executive/Board review.
- Dashboards self-serve with targets and named owners.
25. SOC 2 readiness and internal dry-run audit (depends on: 19, 23, 24)
Design evidence as a by-product and test it with an internal walkthrough before the external auditor arrives.
- Map process to SOC2 CC7.3/CC7.4, CC7.2, CC2.2/2.3, CC5, availability criteria.
- Publish versioned policy documents: Incident Response, Severity, On-Call, Communication, Postmortem.
- Automate evidence: incident records with timestamps, role assignments, paging logs, status history, postmortems, action board, training register, drill records.
- Ensure process operates at least 3 months before fieldwork.
- Run internal dry-run month 6; sample 15 incidents; fix gaps with 8 weeks to spare.
- Keep remediation log for process deviations.
26. Continuous improvement, culture, and sustainability (depends on: 23, 25)
Prevent the process from decaying after the audit by embedding review, feedback, and roadmap ownership.
- Quarterly process retrospective with IC pool and responders.
- Re-baseline metrics every six months; raise targets.
- Year-two roadmap: follow-the-sun coverage, self-healing top 3 causes, error budgets gating releases, blast-radius reduction for shared ledger.
- Annual policy review and certification renewal.
- Quarterly on-call sentiment survey with published actions.
- Board quarterly report on availability, credits, and incident trends.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
Answer with a single JSON object, and nothing else, that conforms to this JSON schema:
{"description": "The structure every proposing agent must answer with, in every round.", "properties": {"steps": {"items": {"description": "One step of a proposed plan. Dependencies are step_ids of the same proposal.", "properties": {"step_id": {"title": "Step Id", "type": "string"}, "title": {"title": "Title", "type": "string"}, "description": {"title": "Description", "type": "string"}, "dependencies": {"items": {"type": "string"}, "title": "Dependencies", "type": "array"}}, "required": ["step_id", "title", "description", "dependencies"], "title": "Step", "type": "object", "additionalProperties": false}, "title": "Steps", "type": "array"}, "estimated_complexity": {"title": "Estimated Complexity", "type": "string"}, "success_metrics": {"title": "Success Metrics", "type": "string"}}, "required": ["steps", "estimated_complexity", "success_metrics"], "title": "ProposalContent", "type": "object", "additionalProperties": false}
{"steps": [{"step_id": "S1", "title": "Executive mandate, budget, and governance", "description": "Secure a written CTO/CEO charter making incident management a company operating process, not a team option. Name one accountable Head of Reliability and a small steering group with authority to decide within 48 hours.\n\n- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking, and protected engineering capacity.\n- Approve budget against the $1.3M annual credits: tooling $150-250k/yr, on-call compensation, training, and 2-3 program FTEs.\n- Set timeline: interim process in week 1, design weeks 1-6, pilot weeks 7-12, rollout weeks 13-24, mock audit month 7, SOC 2 at month 8.\n- Freeze current baselines: 31 incidents, 22-min MTTD, 40% customer-first detection, 3h10 MTTM, $1.3M credits, 3,400 alerts/month at 85% noise, 11/64 actions closed.", "dependencies": []}, {"step_id": "S2", "title": "Interim 7-day command", "description": "Put a crude but real process in place immediately so no incident remains unowned while the permanent design proceeds.\n\n- Publish a one-page interim severity card, one declaration path (phone, Slack command, pager), and one incident channel/bridge/timeline naming convention.\n- Staff an interim 24x7 duty commander with primary and backup from engineering managers and senior SREs; compensate retroactively under final policy.\n- Require a named incident commander within 10 minutes of any suspected major incident, announced in channel.\n- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.\n- Triage the 53 open historical actions in 30 days, prioritizing ledger integrity, duplicate payment, regional failover, security and detection gaps.\n- Hold a daily 15-minute operations review until the permanent process is live.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Forensic baseline and alert estate analysis", "description": "Reconstruct the true before picture from the last 12 months; it drives design and serves as audit baseline.\n\n- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact start/detect/declare/commander/mitigate/resolve, who led, credits paid, contributing factors.\n- For every customer-first detection, name the missing signal; this becomes the detection backlog.\n- Reconstruct the two `nobody in charge` incidents minute by minute; use them as the burning-platform narrative and test case.\n- Profile the 3,400 monthly alerts by tool, team, rule, outcome; identify top 50 noisy rules and every rule with no owner or runbook.\n- Freeze all baseline metrics in one signed document for the executive and the auditor.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Listening tour and fairness contract", "description": "Treat engineer pushback as the main delivery risk and convert it into a written deal that removes the objection to carrying a pager for other teams' code.\n\n- Interview all 28 teams plus Support, CS and Sales in two weeks to separate objections: unpaid work, nights, unfamiliar code, missing runbooks, or fear of blame.\n- Harvest practices from the 12 teams already on-call; they supply pilot teams and first commanders.\n- Publish the deal: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.\n- Recruit 10-15 credible engineers as a design working group so the process is co-authored.\n- Run a baseline sentiment survey and repeat at 60, 120 and 365 days.", "dependencies": ["S1", "S3"]}, {"step_id": "S5", "title": "Service ownership catalog and criticality tiers", "description": "Create a machine-readable catalog as the single source of truth for paging, routing, status page components, and audit evidence.\n\n- Assign one accountable team, engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list per service.\n- Tier by business impact: Tier 0 money movement/ledger/auth/shared PostgreSQL; Tier 1 customer-facing degradable; Tier 2 internal/batch; Tier 3 non-critical.\n- Map customer journeys (initiate, authorize, settle, reconcile, report, onboard) to services, data stores, regions, third parties.\n- Give orphan services an owner within 30 days or a decommission date approved by the sponsor; Tier 0 without an owner is an executive escalation.\n- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.", "dependencies": ["S3"]}, {"step_id": "S6", "title": "Severity scale, declaration rules, and automatic triggers", "description": "Adopt one severity scale as a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade with recorded rationale. Start high when uncertain.\n\n- SEV1: money moved wrongly/duplicated/lost, ledger integrity in doubt, confirmed security/data exposure, both regions impaired, payment processing halted. Full role page, bridge in 5 min, status page in 15 min, regulator assessment within 1 hour, mandatory postmortem.\n- SEV2: material degradation of a payment journey, settlement window at risk, large/strategic customer fully down, SLA breach likely. Commander and SMEs paged, status page in 30 min, mandatory postmortem.\n- SEV3: narrow or single-customer impact with workaround. Owning team leads, business-hours comms, postmortem on request.\n- SEV4: no customer impact; ticket only, never pages.\n- Auto-escalations: any ledger cluster incident, any SEV3 open >2h, any unknown impact after 15 min, any cross-team incident become at least SEV2.\n- Publish a decision tree with 12 worked examples from the 31 real incidents.", "dependencies": ["S3", "S5"]}, {"step_id": "S7", "title": "Incident roles, authority, and handoffs", "description": "Codify roles so coordination never depends on seniority or heroics. The Incident Commander coordinates; responders fix only services they own.\n\n- IC: owns severity, priorities, roles, cadence, closure; pre-authorized to freeze deploys, rollback, disable features, shift traffic, invoke failover, put ledger read-only, commit spend; does not type in terminals; keeps command when VP joins.\n- Communications Lead: sole author for status page, account managers, executives, and hand-off to Legal for regulators; speaks only from commander-approved facts.\n- Scribe: timestamped record of observations, decisions, owners, state changes; human validates for SEV1/SEV2.\n- Subject-Matter Responders: diagnose and mitigate only their owned services.\n- Executive Duty Officer (SEV1): removes obstacles, shields IC from exec questions, owns board/regulator escalation; does not take command unless formally transferred.\n- Rule: IC claimed within 5 min and announced; distinct people for IC, Comms, technical lead at SEV1/SEV2; every handover announced verbally and in writing with time.\n- Financial controls survive: IC coordinates ledger recovery but cannot bypass dual control, reconciliation or privileged access.", "dependencies": ["S6"]}, {"step_id": "S8", "title": "24x7 command and responder coverage model", "description": "Do not create 28 fragile night rotations. Centralize coordination in a trained command corps, keep technical ownership local.\n\n- Layer A: incident command corps of ~30 certified volunteers with primary/secondary 24x7; paired Comms Lead pool ~18 and scribe pool as entry.\n- Layer B: consolidate 28 teams into 10-12 critical-path domains (ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes, data/reporting, integrations) each with 24x7 primary+secondary and minimum six trained responders.\n- Layer C: all other teams business-hours on-call with written after-hours escalation lists held by managers.\n- Platform on-call is safety net for unknown-owner pages, never the permanent owner of someone else's defect.\n- Evaluate follow-the-sun only as a 12-month option, not year-1 dependency.\n- Acknowledgement discipline: 5 min at SEV1/SEV2 with automatic failover.", "dependencies": ["S5", "S7"]}, {"step_id": "S9", "title": "Paid on-call and fatigue safeguards", "description": "Unpaid on-call in New York is a retention and legal risk. Pay must be in payroll before any mandatory night rotation starts.\n\n- Indicative: ~$1,000 per primary 24x7 week, secondary ~$400, business-hours ~$250, duty commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond one hour.\n- Mandatory recovery: paid recovery day after >2h overnight work, SEV1, or qualifying SEV2; managers arrange coverage.\n- HR/Legal/Finance publish amounts, tax treatment, FLSA/NY wage-hour status within 14 days.\n- Load rules: no primary more often than 1 in 6, no consecutive primary/secondary weeks, no primary on two rotations, no on-call during PTO, annual cap.\n- If a rotation averages >2 out-of-hours pages per person per week for 4 weeks, trigger staffing or alert-remediation review.\n- Count on-call as ~15% delivery load; document exemption path for health/caring without career penalty.", "dependencies": ["S8", "S4"]}, {"step_id": "S10", "title": "Service readiness bar and runbooks", "description": "A service must earn the right to page a human at 3 a.m. Define a minimum readiness bar and major incident playbooks before any Tier 0/1 service goes live with on-call.\n\n- Require for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switch, escalation contacts, data-loss/latency impact.\n- Write playbooks for top failures: shared PostgreSQL failure/corruption, cross-region failover, Kubernetes control-plane loss, processor/sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, duplicate payment.\n- Treat ledger cluster as largest structural risk: documented failover, read-only degraded mode, reconciliation after recovery, executive-signed RPO/RTO.\n- Start a parallel workstream on blast-radius reduction: tenant/function partitioning, read replicas, isolation of non-critical readers.\n- Enforcement: no readiness sign-off, no paging at night unless manager accepts a dated exception in writing.\n- Runbooks are peer-reviewed, version-controlled, marked stale if not exercised twice a year.", "dependencies": ["S5", "S8"]}, {"step_id": "S11", "title": "Alert quality standard and page budget", "description": "Make alert quality a condition of paging a human. Burn down noise deliberately rather than by mass silencing.\n\n- Every paging alert must declare owning team, affected service, customer/SLO impact, expected action, dashboard, runbook, dedup key, severity mapping, escalation policy. Anything failing becomes a ticket.\n- Page on symptoms of customer harm (SLO burn rate, payment success rate, queue age vs settlement deadlines), not raw CPU/memory.\n- Run new alerts in shadow mode 7 days; test firing and recovery.\n- Set page budget of 2 out-of-hours pages per person per week; breach blocks new alert creation and triggers tuning sprint.\n- Auto-quarantine rules firing >5 times/month without action or >70% no-action acks; return to owner with correction deadline.\n- Never silently disable: verify compensating detection, record decision, name owner, review missed detections monthly.\n- Target 3,400 to <500 pages/month, actionability >75% within six months, no loss of Tier 0/1 detection.", "dependencies": ["S3", "S5"]}, {"step_id": "S12", "title": "Detection uplift on money path", "description": "Stop customers telling you first by detecting on payment outcomes and ledger truth.\n\n- Define SLOs and business SLIs per customer journey: initiation success, authorization latency, settlement timeliness, reconciliation break rate, API availability, reporting freshness; set internal targets stricter than 99.95%.\n- Run external synthetic end-to-end payments every 60 seconds from both regions, covering all critical payment methods.\n- Add ledger assurance: continuous double-entry balance checks, duplicate detection, replication lag, failover-readiness, settlement-window countdown.\n- Add per-customer anomaly detection for top 100 accounts; five similar support tickets in 10 minutes auto-creates triage incident.\n- Route partner and processor notifications into declaration path within 5 minutes.\n- Record detection source on every incident; treat `customer detected first` as a named defect class with mandatory action.", "dependencies": ["S5", "S11"]}, {"step_id": "S13", "title": "Consolidate to one paging and incident platform", "description": "Collapse six alerting tools into one operational system of record: one queue, one timeline, one audit trail, without creating a monitoring gap.\n\n- Select one paging/scheduling platform, one incident record, one hosted status page; time-box selection to 2 weeks.\n- Ingest all six sources first, deduplicate/correlate; retire legacy path only after named owners, successful end-to-end test, and 2 weeks verified operation.\n- One-command declaration in Slack creates channel/bridge, pages commander, sets severity, opens timeline, starts clock.\n- Route every page through service catalog: service label -> owning domain schedule -> escalation policy.\n- Capture evidence automatically: declaration, acks, role assignments, severity changes, decisions, comms, mitigated/resolved times, postmortem link; retain 12 months.\n- Test out-of-band SMS/phone paging, mobile fallback, offline runbook weekly; ensure works during one AWS region/chat/identity provider failure.\n- Set hard date after which pages outside this tool create no on-call obligation.", "dependencies": ["S6", "S8", "S11", "S12"]}, {"step_id": "S14", "title": "Escalation paths and acknowledgement SLAs", "description": "Write one unskippable path from signal to named commander in under 5 minutes; default action is never waiting.\n\n- Converge all entry points on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.\n- SEV1/SEV2 ladder: owning primary immediately -> secondary at 5 min -> domain manager + Duty Commander at 10 -> Executive Duty Officer at 15. SEV2 requires command assigned within 15 min.\n- If no one claims command in 5 min, platform assigns and announces it; assignee may hand over but not decline.\n- Commander may page any domain's on-call directly with 10-min ack obligation; this reciprocity makes single-team ownership viable.\n- If ownership unclear after 10 min, commander keeps incident and names temporary owner; missing catalog entry logged as control defect.\n- Pre-authorize regional failover, ledger read-only, payment suspension, partner-bank notification so IC never waits for an executive.", "dependencies": ["S7", "S8", "S13"]}, {"step_id": "S15", "title": "Standardize internal and customer communications", "description": "Replace whoever is around with a timed, owned, pre-approved process; the Communications Lead is single author and never writes from scratch under pressure.\n\n- Internal: one working channel/bridge plus one read-only broadcast for executives/Support/Sales; cadence SEV1 15-30 min, SEV2 60 min even if nothing changed; fixed template with impact, action, next update, IC/Comms names.\n- Executives ask questions only of Executive Duty Officer; publish this as signed behavior rule.\n- Status page: SEV1 within 15 min, SEV2 within 30 min; updates every 30/60 min; resolution notice within 30 min of verified recovery; customer-facing summary within 5 business days for SEV1/qualifying SEV2.\n- Pre-approve 12-15 templates with Legal for degradation, settlement delay, API errors, regional failure, data-integrity investigation, security event.\n- Top 100 accounts get named account manager call/email within 30 min of SEV1 with briefing pack; all 2,100 subscribed by default.\n- Language rules: state impact and next update; never speculate on cause, recovery time, data integrity or blame.", "dependencies": ["S6", "S7", "S13"]}, {"step_id": "S16", "title": "Regulatory, partner, and financial impact workflows", "description": "In payments some incidents start a legal clock at detection. Build obligation assessment into the process and tie incidents to money.\n\n- Legal/Compliance produce obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC, PCI, money-transmitter, sponsor-bank/card-network windows, cyber-insurer notice.\n- Mandatory reportability checkpoint on every SEV1 and security-related SEV2 within 1 hour, recorded even when `not reportable`, with decision-maker and evidence.\n- Maintain 24x7 contact matrix for regulators, sponsor banks, networks, outside counsel, insurer.\n- Agree availability measurement method per contract; compute affected minutes per customer from incident record and journey telemetry; propose credit schedule within 5 business days.\n- Attribute credits to root-cause families; track failed payment volume, delayed value, reconciliation breaks; target annual credits from $1.3M to under $400k.", "dependencies": ["S6", "S15"]}, {"step_id": "S17", "title": "Blameless postmortem standard and action tracking", "description": "Replace various formats with one mandatory format, fixed deadlines, and enforceable actions. Closing a ticket without evidence does not close the action.\n\n- Mandatory for every SEV1/SEV2, customer-first detection, incident >2h, repeat of known cause, credit-generating/contract breach, ledger near-miss.\n- Draft within 3 business days, review within 5, publish within 10; IC owns delivery, owning manager accountable.\n- One template: summary, customer/financial impact, detection source/gap, timeline, response analysis, contributing conditions, what worked, actions.\n- Two mandatory questions: why did a customer see this first? why did mitigation take as long as it did?\n- Blameless in writing: systems and context, no individual named as cause, never used in performance reviews; HR handles personnel/misconduct separately.\n- Every action gets owner, priority, due date, verification method, ticket auto-created; classes containment 7 days, corrective 30, strategic 90; reserve 15% engineering capacity.\n- Escalation: manager at +7 days, director at +14, CTO dashboard at +30; overdue P0 can block related releases.\n- Weekly Incident Review Board ratifies severity, challenges quality, monitors actions; target 90% closure on time within two quarters.", "dependencies": ["S6", "S7"]}, {"step_id": "S18", "title": "Train and certify incident roles", "description": "Command is a skill, not a title. Certify before duty; use paid working time.\n\n- All employees: 30-min module on recognizing impact, declaring, finding channel/status page.\n- Responder: half-day on severity, escalation, runbook, financial integrity; mandatory before joining rotation.\n- Scribe: 2 hours timeline discipline; entry point.\n- Incident Commander: 2 days + shadowing; command presence, delegation, severity calls, running room, handover; certification requires two shadowed incidents and one simulated SEV1.\n- Communications Lead: 1 day on status writing, customer tiering, legal boundaries, regulator triggers.\n- Domain responders demonstrate dashboard, runbook, rollback, failover, access before primary; two shadow shifts; no primary in first 90 days.\n- Certification valid 12 months, renewed by simulation; register is audit artifact.", "dependencies": ["S7", "S13", "S15"]}, {"step_id": "S19", "title": "Simulations and game days", "description": "Rehearse process before real SEV1; exercise tooling failure and ledger scenarios.\n\n- Monthly tabletop per group reusing an incident from baseline, rotating commander.\n- Quarterly game day in staging or tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes degradation.\n- Twice-yearly unannounced paging drill to measure real overnight ack times.\n- Annually one combined operational/security exercise and one regulator-notification exercise with Legal/CISO.\n- Exercise status page down, chat/paging provider unavailable.\n- Never inject uncontrolled change into production ledger; validate backups/restore/RPO/RTO on replicas.\n- Every exercise produces tracked actions; publish time-to-commander, time-to-first-update, time-to-mitigation.", "dependencies": ["S10", "S14", "S15", "S18"]}, {"step_id": "S20", "title": "Pilot on critical path", "description": "Prove full process on highest-risk surface with willing teams for six weeks before full rollout.\n\n- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, Support intake, including two of the 12 already on-call teams.\n- Activate severity scale, command corps, single paging tool, page budget, status page policy, mandatory postmortems, paid rotations.\n- Parallel old paths for one week then cut over; real incidents use new process only.\n- Program lead attends every SEV2+ as coach, never shadow commander; review every pilot page within one business day.\n- Exit gate: commander named within 5 min in 95% cases, first status update on time, MTTD under 10 min for pilot services, pages halved, no unpaid page, all required postmortems on time, positive sentiment.\n- Publish one-page result company-wide.", "dependencies": ["S9", "S11", "S12", "S13", "S14", "S15", "S17", "S18", "S19"]}, {"step_id": "S21", "title": "Metrics, dashboards, and review cadence", "description": "Instrument the process so improvement is visible and audit sees monitoring/review evidence. Report median and p90, never averages alone.\n\n- Response: time to detect/declare/commander/ack/mitigate/resolve split by severity, tier, journey, region, detection source.\n- Quality: customer-first rate, status page timeliness, update cadence, missed escalations, role conflicts, alert actionability, out-of-hours pages per person.\n- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat factors.\n- Business: availability per journey, error budget burn, failed payment volume, reconciliation breaks, SLA credits.\n- People: rotation size, frequency, recovery days, sentiment, attrition.\n- Weekly Incident Review Board; monthly reliability review per team; monthly executive review to CEO; quarterly control review with Security/Compliance/Risk; quarterly board summary.\n- Anti-gaming: reconcile incident counts monthly against support tickets, credits, status page history; team scorecards direct help, never individual penalties.", "dependencies": ["S13", "S17"]}, {"step_id": "S22", "title": "Wave rollout to all 28 teams with readiness gates", "description": "Roll out in four waves by customer risk, each passing an explicit gate rather than a date. Complete all teams by week 24 to leave ~3 months of evidence before audit fieldwork.\n\n- Wave 1 remaining Tier 0, Wave 2 Tier 1, Wave 3 Tier 2, Wave 4 Tier 3/internal.\n- Per-team onboarding kit: catalog entry complete, alerts migrated within budget, runbooks at readiness bar, rotation staffed with 6 trained responders or time-limited exception, one IC candidate nominated, one tabletop passed, payroll active.\n- Named coach per wave for three weeks; director signs gate.\n- Legacy paging paths disabled per wave, not kept as comfort fallback.\n- Publish live adoption scoreboard.\n- Any team unable to staff fair rotation gets headcount, service reassignment, or explicit executive risk acceptance.", "dependencies": ["S20", "S21"]}, {"step_id": "S23", "title": "Culture, fairness, and continuous feedback", "description": "Run this in parallel from day one. Engineers judge fairness; executives judge results.\n\n- Repeat the deal in every forum: paid on-call, paged only for owned services, trained commander, real sprint capacity.\n- Publicly close the two historical nobody-in-charge incidents with what would be different now.\n- Weekly office hours first 8 weeks, Slack support channel 4-hour SLA, per-team champion, weekly newsletter with metrics and bad news.\n- Incident response contribution in promotion criteria; quarterly awards for best postmortem and biggest noise reduction; public thanks after SEV1.\n- Pulse survey at 60, 120, 365 days; if fairness or load red, pause expansion until fixed.", "dependencies": ["S4", "S9"]}, {"step_id": "S24", "title": "SOC 2 evidence design and internal dry-run", "description": "Make evidence a by-product of operations; test before external auditor.\n\n- Map with Compliance to Trust Services Criteria CC7.2-CC7.5, CC2.2/CC2.3, A1.2; confirm observation window early.\n- Evidence set automated and indexed: versioned policies/exception, rotation schedules, compensation activation, incident records with timestamps/roles, paging/ack logs, status history, reportability decisions, postmortems, action closure with verification, training/certification register, drill records, access reviews.\n- Internal Audit tests process in months 4 and 6, sampling incidents end-to-end.\n- Run formal mock audit month 7 using same evidence populations; include interviews with IC, engineer, Support, Compliance.\n- Correct failures through tracked actions, never by editing history.\n- Freeze process wording after month 7; any change logged as exception.", "dependencies": ["S13", "S17", "S21", "S22"]}, {"step_id": "S25", "title": "Risk register and contingency planning", "description": "Name likely failure modes and pre-commit responses; review monthly with sponsor.\n\n- Too few commander volunteers: make command rostered duty for managers and staff engineers until pool reaches 30.\n- Compensation not approved in time: fallback to time-off-in-lieu plus phased stipend, never launch mandatory night on-call uncompensated.\n- Tool migration slips: cut scope on incident record layer, never on paging consolidation.\n- Noise pruning hides real failure: demote to ticket first, observe 30 days, keep recovery path, monthly missed-detection review.\n- Burnout in experienced teams: weekly load monitoring with page caps.\n- Major SEV1 mid-rollout: program lead becomes full-time responder, wave schedule slips one wave, sponsor told same day.\n- Shared ledger concentration: if blast-radius workstream slips, escalate to board as accepted risk with dated plan.", "dependencies": ["S1"]}, {"step_id": "S26", "title": "Inspect and adapt; year-two sustainability", "description": "Prevent decay after audit by revising on data and assigning permanent owners.\n\n- At 90 days live, revise policy using measurements: severity calibration, Layer B/C membership from page data, uncovered shifts, commander burn, missed updates, action closure, survey results.\n- Assign permanent owners for policy, paging platform, status page, service catalog, training, metrics.\n- Re-baseline targets every six months; shift from lagging to leading metrics: error budget burn, near-miss rate, drill performance.\n- Year-two candidates: follow-the-sun coverage, automated mitigation for top recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, completion of ledger blast-radius reduction.\n- Keep annual policy review, certification renewal, exercise calendar, board reporting as permanent commitments.", "dependencies": ["S22", "S24"]}], "estimated_complexity": "high", "success_metrics": "- Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.\n- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.\n- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.\n- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.\n- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.\n- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update cadence compliance above 95%.\n- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, no loss of Tier 0/1 detection.\n- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.\n- 100% Tier 0/1 services have named owning team, tier, escalation policy, dashboard and runbook by day 30; all 180 services by day 120.\n- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10-12 domain rotations of at least six trained responders.\n- 30 certified incident commanders and 18 certified communications leads active.\n- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.\n- 100% required postmortems drafted in 3 business days, reviewed in 5, published in 10 from month 2.\n- Postmortem action closure rises from 17% to 90% closed by due date within two quarters; all 53 legacy actions triaged within 30 days.\n- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.\n- SLA credits fall from $1.3M to under $400k annualised within 12 months; customer-impacting incidents fall from 31 to under 15 per year.\n- Regulatory reportability assessed and recorded within 1 hour for 100% SEV1/security SEV2 including not reportable.\n- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two cross-company exercises completed before audit, each producing tracked actions.\n- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% sampled incidents; SOC 2 Type II incident response controls pass zero exceptions.\n- On-call fairness sentiment above 70% agreement at 120 days; no increase in attrition among responders."}Round 3 — refinement 3 of 3
Round 3 converged hard: all five plans now run the same skeleton (charter → 7-day floor → forensic baseline → fairness contract → catalog/tiers → severity → roles → three-layer 24x7 → paid on-call → alert budget → detection → single paging platform → escalation → comms → regulatory/credits → postmortems/actions → training → exercises → policy → pilot → gated waves → metrics → SOC 2 → 90-day inspect), with near-identical numbers ($1,000 primary week, 2 out-of-hours pages/responder/week, 15% reserved capacity, 30 ICs / 18 comms leads). Differentiation now comes from a few genuine additions — P1's change intelligence and Support-as-detection tier, P4's Duty Triage Desk carve-out and vendor-incident class, P2's defined noise/actionable metrics and severity modifiers — while P3 and P5 mostly merge material authored by P1 and P4.
What still separates them
- Unique content: only P1 has change intelligence/deployment safety (step 13), Support and account managers as an instrumented detection tier (step 20) and a concurrency doctrine with a Multi-Incident Coordinator (step 8); only P4 has a vendor-incident class for processor/bank/cloud failures (step 13) and a rule for who writes the 3 a.m. status page (step 8); only P2 defines "actionable page" and "noise" and uses five modifiers instead of flags (steps 12, 6). P3 and P5 contribute nothing the others do not have.
- Overnight first line: P1 (step 9), P4 (step 8) and P5 (step 8) fund a paid Duty Triage Desk; P4 adds the safety carve-out that it never holds suspected ledger, payment-halt or security pages. P2 (step 8) and P3 (step 9) refuse a triage desk and route unknown-owner pages to Duty Command plus platform, P2 additionally reclassifying lower-tier services that can cause severe overnight harm.
- Schedule: P2 (step 1) compresses design to weeks 1–4, pilot 5–10, rollout by week 18, control tests in months 4 and 6; P1, P3, P4 and P5 keep design weeks 1–6, pilot 7–12, rollout 13–24 and a month-5 internal test. P2 buys a longer observation window at the cost of a 4-week design for 180 services.
- Failure handling and metric honesty: P1 (step 34), P3 (step 28) and P5 (step 29) keep a standing monthly risk register; P4 compresses contingencies into bullets in step 26; P2 has none. On claims, P2 targets "no unresolved material exception" and P1 warns credits may rise before they fall (step 22), while P3 and P5 still promise "SOC 2 passes with zero exceptions".
Who took what from whom
- P1's replay test (R2 step 11) went everywhere: P2 step 13 plus a replay-coverage metric, P3 step 12, P4 steps 3 and 11 (as a "replay catalog" with a gap owner), P5 step 11. It is now the field's standard proof that detection actually improved.
- P4's severity flag (R2 step 6) was taken by P1 (step 7, financial-integrity/security flag forcing dual control and the reportability checkpoint) and generalised by P2 into five modifiers — FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, VENDOR (step 6). P3 and P5 declined it and kept four levels plus auto-escalation.
- P1's Duty Triage Desk (R2 step 8) was adopted by P4 (step 8) and P5 (step 8); P4 improved it with the carve-out that ledger, payment-halt and security events page commander and domains in parallel, plus a half-day triage-desk training module (step 18). P2 and P3 explicitly did not.
- P4's "confirm the SOC 2 observation window in week 1" (R2 step 1) was taken by P3 (step 1) and P5 (step 3), and expanded by P1 into a full week-one audit, legal and evidence scoping step (step 4) covering privilege, legal hold and the obligation matrix.
- P4's staffing arithmetic (R2 step 8 — 260 engineers cannot sustain 28 night rotas, never a two-person rotation) was taken verbatim by P3 (step 9) and restated by P1 (step 9, ~100 of 260 engineers carry a night obligation at one week in six). Nobody took P2's R2 caution against merging small teams for schedule convenience — all five consolidate into 10–12 domains.
The calls of this round
Influences: who took what from whom
| Round 3 ↓ · round 2 → | Proposal 1 | Proposal 2 | Proposal 3 | Proposal 4 | Proposal 5 | New steps |
|---|---|---|---|---|---|---|
| Proposal 1 |
kept30 | same titles0 analyst sees+3 / −2 | same titles1 analyst sees+0 / −0 | same titles0 analyst sees+3 / −0 | same titles0 analyst sees+0 / −0 | new4 |
| Proposal 2 |
same titles0 analyst sees+1 / −3 | kept18 | same titles1 analyst sees+1 / −0 | same titles7 analyst sees+3 / −0 | same titles0 analyst sees+0 / −0 | new2 |
| Proposal 3 |
same titles2 analyst sees+1 / −1 | same titles2 analyst sees+1 / −1 | kept14 | same titles11 analyst sees+3 / −1 | same titles0 analyst sees+0 / −0 | new0 |
| Proposal 4 |
same titles1 analyst sees+3 / −1 | same titles4 analyst sees+2 / −2 | same titles2 analyst sees+0 / −0 | kept17 | same titles0 analyst sees+0 / −0 | new2 |
| Proposal 5 |
same titles28 analyst sees+5 / −1 | same titles0 analyst sees+0 / −1 | same titles0 analyst sees+0 / −0 | same titles1 analyst sees+1 / −1 | kept1 | new0 |
Four substantive new steps and three hardening additions, all addressing gaps the field had left open. It is now the most complete plan, at the cost of being the longest (35 steps).
- New step 4: week-one audit, legal and evidence scoping — observation window, TSC mapping, retention, legal hold and which postmortem content is privileged, with monthly evidence sampling from month 1.
- New step 13: change intelligence — deployments, config, flags and migrations streamed into the timeline, a "what changed in the last 60 minutes" panel at declaration, tested rollback executable by the on-call without the author, no deploys in the settlement window; backed by a new metric (change correlation identified within 10 minutes in 90% of cases).
- New step 20: Support and account management as a detection tier — declaration macro, five-tickets-in-ten-minutes clustering, auto-generated affected-customer list, and "signal was in Support before monitoring" published as a detection defect. This attacks the 40% customer-first number directly.
- Step 22 adds an honest incentive guardrail: better detection will raise credits before it lowers them, and no commercial pressure may influence severity.
- Step 8 adds a concurrency doctrine (secondary commander, surge roster, Multi-Incident Coordinator arbitrating ledger/DB/freeze) with a matching concurrency drill in step 26 and in the metrics.
- Step 9 publishes the staffing arithmetic (~100 of 260 engineers on a night obligation) so the rejection of 28 rotations is argued on numbers.
- 35 steps with dense dependency chains (step 28 depends on seven prior steps) — the heaviest plan to actually run inside eight months.
- The on-call compensation metric changed to "before any mandatory night rotation", which is correct, but step 2's interim 24x7 roster still runs for weeks on retroactive pay only.
- Proposal 4 : Attach a financial-integrity or security flag to any severity instead of inventing a fifth level.
- Proposal 4 : Confirm the SOC 2 Type II observation window with the auditor in week 1 so the interim process counts as evidence.
- Proposal 2 : Legal hold, confidentiality and privileged handling of incident records, plus MFA and access reviews on the incident platform.
- Proposal 2 : A surge roster for simultaneous incidents.
- Proposal 2 : Initial internal brief within 10 minutes for SEV1 and 15 for SEV2, and a monitoring notice after mitigation.
- Proposal 4 : Staffing arithmetic and the ban on two-person 24x7 rotations.
- Proposal 2 : Do not merge small teams merely for schedule convenience.
- Proposal 2 : Finish company rollout by week 18 with control tests in months 4 and 6.
+ Audit, legal and evidence scoping in week one+ Change intelligence and deployment safety on the critical path+ Support and account management as a detection and intake tier+ SLA credit workflow, true cost model and incentive guardrailsSLA credit and financial impact workflow
The plan produced
1. Charter the programme: one owner, one mandate, funded, dated against the audit
Turn the CEO email into a chartered company programme with a single accountable owner and authority over all 28 teams. Incident response stops being a per-team preference and becomes a company operating process.
- Name the CTO as executive sponsor and a full-time Head of Reliability & Incident Management as accountable owner, supported by a programme office of three: programme lead, incident-platform engineer, reliability analyst.
- Form a decision group (Engineering, Platform/SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables on day one: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, named service ownership, and paid on-call.
- Publish the timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal control test month 5, mock audit month 7, SOC 2 fieldwork month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.3M/yr, training and exercises, plus 15% of engineering capacity reserved for reliability work.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (after 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence immediately.
- Publish a one-page interim severity card and a single declaration path: one chat command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim 24x7 Duty Incident Commander roster (primary plus backup) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document and one naming convention for every major incident, starting now.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (after 1)
Rebuild the facts before designing anything. This is the design input, the executive narrative and the frozen "before" picture for the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Correlate incidents with deployments, config changes and feature flags to quantify how many started with a change we made.
- Quantify true cost beyond credits: failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost, churn risk on named accounts.
- Freeze the baselines in a signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Audit, legal and evidence scoping in week one (after 1) new
Design evidence as a by-product of operating, never as a reconstruction before fieldwork. Settle the scope and the legal handling of incident records now, not in month seven.
- Confirm with the external auditor the Type II observation window, the incident population definition, and the evidence they will sample. Everything from the day-7 floor onwards must count.
- Map incident response to the Trust Services Criteria with Compliance: CC7.2–CC7.5 (monitoring, identification, response, recovery), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), A1.2 (availability).
- Define the evidence set and where it is produced automatically: incident records, paging and acknowledgement logs, role assignments, status-page history, reportability decisions, postmortems, action closure with verification, training register, drill records, access reviews.
- Agree retention, confidentiality, legal hold and access rules. Decide with counsel which postmortem content is privileged and how privileged material is segregated without making the ordinary postmortem secret.
- Start the obligation matrix with Legal: NYDFS 23 NYCRR 500, state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network contractual windows, cyber-insurer notice.
- Begin monthly evidence sampling from month 1 so control operation is visible long before the audit.
5. Listening tour, resistance map and the written on-call deal (after 1, 3)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written deal signed by the sponsor.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
- State the hard gate publicly: no new mandatory night rotation starts before compensation, training, runbooks and staffing rules are live.
6. Service catalog, ownership, tiering and customer-journey map (after 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- Record for every service: one accountable team, engineering manager, chat channel, escalation policy, dashboards, runbook link, dependency list, regions and data stores.
- Tier by business impact, not technology. Tier 0: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record SLO, RTO, RPO, active-active or region-bound, failover method, and the dependencies that make nominal two-region redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
7. Severity standard, declaration rights and incident lifecycle (after 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- SEV1 (crisis): money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning responders and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment started within 1 hour; mandatory postmortem.
- SEV2 (critical): material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and responders paged, comms and scribe if customer-visible, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- SEV3 (contained): narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, credit-generating or a repeat.
- SEV4: no customer impact. Ticket only; never pages a human.
- Attach a financial-integrity flag or a security flag to any severity. The flag forces dual control, Legal engagement and the reportability checkpoint without inventing a fifth level.
- Anchor on payments reality alongside error rates: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring. Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including a customer report.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
8. Roles, authority, concurrency and handover discipline (after 7)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- Incident Commander: owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend. Retains command when a VP joins.
- Communications Lead: the single voice to the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1 and SEV2.
- Subject-Matter Responders: engineers of the owning team. They mitigate; they do not run the room.
- Executive Duty Officer (SEV1): removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, Finance and Vendor Management join on defined triggers rather than by invitation.
- Rules: command claimed within 5 minutes and stated in channel; distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Concurrency doctrine: two simultaneous SEV1/SEV2 incidents activate the secondary commander and a surge roster; a designated Multi-Incident Coordinator arbitrates shared resources such as the ledger, the database platform and the deploy freeze.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation or privileged-access rules.
9. Three-layer 24x7 coverage: command corps, domain rotations, overnight triage desk (after 5, 6, 8)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- Layer A — Incident Command corps: ~30 certified commanders drawn from across the 28 teams plus managers, giving roughly one primary week per person every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- Layer B — domain responder rotations: consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement and reconciliation, auth, API edge, Kubernetes platform, data and reporting, partner integrations). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- Layer C — everyone else: business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- Duty Triage Desk: a small paid overnight first-line rotation owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Publish the staffing arithmetic: 10–12 domain rotations of six to eight people plus a 30-person command corps means roughly 100 of 260 engineers carry a night obligation, at about one week in six. Twenty-eight independent rotations would be unstaffable and is therefore rejected on the numbers.
- Teams too small for a fair rotation get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: a human acknowledgement within 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
10. Paid on-call, New York labour compliance and fatigue safeguards (after 5, 9)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour, and a premium for the overnight triage desk.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per responder per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Hard gate: no mandatory night rotation starts before its compensation is live in payroll. Interim duty since day 7 is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons without career penalty, and on-call load is published quarterly by team.
11. Alert quality contract and page budget (after 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a page budget of two out-of-hours pages per responder per week. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression can never hide an incident.
12. Detection uplift on the money path, validated by incident replay (after 6, 11)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Run the replay test: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action and a missed-detection review.
13. Change intelligence and deployment safety on the critical path (after 6, 12) new
Most of these incidents start with something we changed. Make change the first hypothesis the tooling answers, and make changes safer to reverse.
- Stream every deployment, configuration change, feature-flag flip, schema migration and infrastructure change into the incident timeline with service labels and owners.
- Give the commander an automatic "what changed in the last 60 minutes on the affected journey" panel at declaration time.
- Require Tier 0/1 changes to be progressively delivered with a documented rollback that is tested, time-bounded and executable by the on-call responder without the author.
- Treat ledger schema migrations and settlement-affecting changes as a separate class: dual approval, rehearsed rollback, no deploys inside the settlement window.
- Enforce a change freeze during SEV1 and SEV2, lifted only by the commander and logged.
- Report change-correlated incidents monthly; a rising ratio is a signal to strengthen release safety, not to blame a team.
14. One pager, one incident record, one status page — migrated without a detection gap (after 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy paging path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access, MFA and periodic access review.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page delivered outside this platform creates no on-call obligation.
15. Escalation ladder and the five-minute command rule (after 8, 9, 14)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts and support entitlements: AWS and database premium support, processors, sponsor banks, card networks.
16. Live execution doctrine and payments safety rules (after 8, 14, 15)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- Payments guards: protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed in advance rather than improvised.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and automatic reopening if impact recurs.
17. Service readiness bar, major-incident playbooks and ledger blast-radius reduction (after 6, 9, 16)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human, and the shared ledger is the largest structural risk in the estate.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as a concentration risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers, stronger regional independence — with board-visible quarterly milestones.
- Enforcement: no readiness sign-off, no night paging alerts unless the manager accepts the gap in writing with an expiry date and a compensating control. Never respond to a gap by turning detection off.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
18. Internal communications protocol (after 8, 14)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, Sales and Security.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change. First internal brief within 10 minutes of SEV1 and 15 of SEV2.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
19. Customer communications, status page and account-manager outreach (after 7, 18) from P3 step 17
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; monitoring notice after mitigation; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, monitoring, resolution.
- Component-level status page mapped to customer journeys rather than internal service names, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1 and 60 of SEV2, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state observed impact, affected capabilities, any workaround and the next update time; never speculate on cause, recovery time, data integrity or blame.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
- Record any legally required restriction or delay of public detail, its approver, and the alternative stakeholder plan.
20. Support and account management as a detection and intake tier (after 7, 12, 19) new
Customers detected 40% of incidents first, which means the front line already holds the signal. Turn Support and account managers into an instrumented detection channel rather than a bystander.
- Give Support explicit declaration rights, a one-page trigger card, and a macro that opens an incident candidate directly in the incident platform.
- Automate the clustering rule: five similar tickets or calls in ten minutes auto-creates a triage incident assigned to the duty commander.
- Route credible partner, processor and sponsor-bank notifications into the same declaration path within five minutes.
- Generate the affected-customer list automatically from journey telemetry and the incident record, and push it to Support, CS and the account-manager briefing.
- Train Support and account managers on approved language and prohibit independent technical explanations to customers.
- Measure and publish "signal was in Support before it was in monitoring" as a detection defect, and feed each instance into the detection backlog.
21. Regulatory, partner and legal notification playbook (after 4, 7, 19)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Complete the obligation matrix started in scoping and have counsel validate the triggers, deadlines, channels and submitting authority for each obligation.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded even when the answer is "not reportable", with facts considered, decision-maker, timestamp and reassessment trigger.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, card networks, outside counsel, cyber-insurer, critical vendors.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the operational facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly inside the exercise programme, including one full regulator-notification simulation per year.
22. SLA credit workflow, true cost model and incentive guardrails (after 7, 19) new
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component, including how partial degradation counts.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Guardrail against the perverse incentive: better detection will surface incidents that previously went unbilled, so credits may rise before they fall. Publish this expectation to the executive team in advance, and make it a written rule that no finance or commercial pressure may influence severity or declaration.
23. Blameless postmortem standard and Incident Review Board (after 4, 7, 8)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger; any control failure.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning engineering director is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did?
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions where security, privacy or privileged content requires it.
24. Action ownership, reserved capacity and enforcement (after 14, 23)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, accountable manager, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness through tests, telemetry, exercises or production evidence before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
25. Training and certification academy (after 8, 15, 18, 23)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days plus shadowing): command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, the 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Comms Lead (1 day): status writing, customer tiering, contractual clocks, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days of joining a domain.
- Certification is valid 12 months and renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
26. Exercise programme: tabletops, game days, night drills and vendor rehearsals (after 14, 17, 25)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander. First company-wide tabletop within 30 days.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced overnight paging drill to measure real acknowledgement times, plus a concurrency drill with two simultaneous incidents.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down. Rehearse joint escalation with AWS, a processor and a sponsor bank.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
27. Publish Incident Management Policy v1 and the exception register (after 7, 8, 9, 10, 11, 15, 18, 19, 21, 23, 24)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, escalation and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix, Readiness Bar.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
28. Pilot on the payments critical path (after 10, 12, 13, 14, 17, 25, 27)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement and reconciliation, PostgreSQL platform, Kubernetes platform, API edge, auth and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, action tracking, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The programme lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time in 95% of cases, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
29. Alert noise burn-down campaign (after 11, 14, 28)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → under 1,500 pages a month by day 90, under 500 with noise below 15% by month 6.
30. Metrics, dashboards, review cadence and anti-gaming (after 14, 23, 28)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- Response: time from first impact to detect, declare, commander, acknowledge, mitigate, resolve — split by severity, tier, journey, region and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors, change-correlated incident ratio.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
31. Wave rollout to all 28 teams with readiness gates (after 28, 30)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the programme office for three weeks; the director signs the gate.
- A failed gate is rescheduled, never waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard by team, tier and control gap.
32. Change management, fairness and pager culture (after 5, 10, 28)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
33. SOC 2 evidence by design, internal control test and mock audit (after 4, 27, 30, 31)
The real process is the evidence. Never build a parallel audit process, and never reconstruct records after the fact.
- Maintain the evidence set defined in scoping, produced automatically and indexed: versioned policies and exceptions, catalog records, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Sample incidents monthly from first signal through verified corrective action, deliberately including customer-reported events, downgraded incidents and missed timelines.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling end to end and reporting gaps to the sponsor.
- Run a formal mock audit in month 7 using the populations, evidence requests and interviews the auditor will use: a commander, a random engineer, Support, Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
34. Programme risk register and pre-committed contingencies (after 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps and immediate staffing intervention.
- A major SEV1 mid-rollout: the programme lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
35. Ninety-day inspect-and-adapt, then year-two sustainability (after 30, 31, 33)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, exercise calendar and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to leading indicators: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and quarterly board reporting as permanent commitments owned by the Head of Reliability.
- A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90 a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95% internally and externally.
- Top-100 account outreach completed within 30 minutes of SEV1 in 95% of qualifying cases, using the approved briefing pack.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- Change-correlated incidents are identified within 10 minutes of declaration in 90% of cases, and the change-correlated share of incidents declines quarter over quarter.
- SLA credits fall to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, one concurrency drill and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and a tested read-only mode by month 6, with milestones reported to the board quarterly.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: d55e2704-ef8a-4e9c-8fd5-e6d956728a8d, Agent: opus5_refine_1, LLM: anthropic/claude-opus-5):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (32):
1. Charter the program: one owner, one mandate, funded, dated before the audit
Turn the CEO email into a chartered company program with a single accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the **burning-platform narrative** and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the written on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (depends on: 5, 7)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 8, 4)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (depends on: 10, 5)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (depends on: 7, 12, 13)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
18. SLA credit and financial impact workflow (depends on: 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Use the credit delta as the standing business case for on-call pay and reserved reliability capacity.
19. Blameless postmortem standard and Incident Review Board (depends on: 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, reserved capacity and enforcement (depends on: 19, 12)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
21. Runbooks, readiness bar and ledger blast-radius reduction (depends on: 5, 8)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (depends on: 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (depends on: 22, 12, 21)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
24. Publish Incident Management Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 11, 12, 21, 22, 9)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Alert noise burn-down campaign (depends on: 10, 12, 25)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
28. Wave rollout to all 28 teams with readiness gates (depends on: 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
29. Change management, fairness and pager culture (depends on: 4, 9, 25)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (depends on: 24, 26, 28)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 26, 28, 30)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 2 (ID: ddb335b2-98d9-46f8-afa0-88fca2e25749, Agent: gpt5.6-sol_refine_2, LLM: openai/gpt-5.6-sol):
Estimated Complexity: high
Success Metrics: - By day 7, every suspected major incident uses one incident record, one coordination channel, and a named commander.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes; at least 95% receive command within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and technical coverage.
- By week 18, all 180 services have an owner, tier, tested escalation path, and coverage appropriate to their risk.
- No mandatory night rotation starts before compensation, training, access, runbooks, and staffing controls are active.
- Every direct 24x7 technical rotation has at least six qualified responders or a documented, expiring executive exception.
- No responder is routinely assigned primary duty more often than one week in six or simultaneously assigned to two primary rotations.
- At least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of qualifying SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet their required internal and customer update cadence by month 3.
- Monthly human pages fall from 3,400 alert events to no more than 1,500 by day 90 and no more than 500 actionable pages by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, with no decline in Tier 0/1 detection coverage.
- Average after-hours load remains at or below two pages per responder per week; every sustained breach produces a remediation plan.
- All six monitoring sources route human pages through the single paging platform by week 12; direct legacy paging paths are disabled by week 18.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Monthly contracted availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- Quarterly on-call surveys reach at least 75% favorable responses on fairness, ownership boundaries, compensation, and sustainability by month 6.
- The month-7 mock audit finds no unowned high-risk control gap, and at least 95% of sampled incidents contain complete operating evidence.
Steps (26):
1. Establish the mandate, owner, funding, and schedule
Launch incident management as a **company operating program within 48 hours**. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Form a decision group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Approve the non-negotiables: one severity scale, one human-paging platform, one incident record, paid on-call, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, and incident actions.
- Fund tooling, compensation, training, exercises, and reliability work. Use the $1.3M in credits as the minimum financial comparison.
- Set milestones: interim process by day 7, policy and critical-path pilot by week 6, Tier 0/1 coverage by week 10, company rollout by week 18, internal audit in month 6, and mock audit in month 7.
2. Install a seven-day incident-response floor (depends on: 1)
Do not wait for policy design or tooling migration. Put a minimum process into operation immediately and begin retaining evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communication schedule.
- Create one monitored declaration path through chat, telephone, and the current paging environment.
- Use one incident channel, bridge, timeline document, and naming convention for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced on-call engineers and engineering leaders.
- Require a named commander within 10 minutes. The duty engineering director assumes command if nobody else does.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Compensate interim duty retroactively under the permanent policy.
- Hold a 15-minute daily operational review until the permanent process is live.
3. Create the factual, legal, and audit baseline (depends on: 1)
Build one defensible baseline for process design, executive decisions, and SOC 2 testing. Confirm the auditor's expected Type II observation period immediately.
- Reconstruct all 31 incidents from first customer impact through resolution, communications, credits, postmortem, and corrective action.
- Analyze the two incidents with unclear command minute by minute.
- Identify the missing detection signal for every customer-first incident.
- Inventory the six alert sources, 3,400 monthly alert events, noisy rules, duplicates, missing owners, and missing runbooks.
- Record the current rotations, unpaid work, after-hours load, and teams without coverage.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M credits, 85% noise, and 11 of 64 actions closed.
- Map incident-response controls to applicable SOC 2 criteria with Compliance and the auditor.
- Establish retention, confidentiality, legal-hold, and access requirements for incident evidence.
4. Co-design the fairness contract with engineers (depends on: 1)
Treat pager resistance as a valid design constraint. Make the employment and ownership bargain explicit before expanding on-call.
- Interview representatives from all 28 teams and the Support, Customer Success, Security, and Operations groups.
- Distinguish objections involving unpaid work, unfamiliar code, bad alerts, weak runbooks, sleep disruption, or blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for services they own or have formally accepted, supported by trained command, and given capacity to remove recurring defects.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Measure baseline trust, fairness, fatigue, and psychological safety. Repeat the survey at days 60 and 120, then quarterly.
5. Build the service catalog and criticality model (depends on: 3)
Make a machine-readable catalog the source of truth for routing, escalation, customer impact, and audit evidence. Every production service must have one accountable owner.
- Record the owning team, manager, product capability, escalation policy, communication channel, dashboard, runbook, dependencies, regions, and data stores for all 180 services.
- Map payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, and reporting to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing services whose failure can create material contractual impact.
- Classify Tier 2 as internal or deferrable services, and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional mode, recovery method, and contractual obligations for Tier 0 and Tier 1.
- Give orphan services an owner or approved decommission date within 30 days.
- Maintain separate but coordinated ownership for the ledger application and PostgreSQL platform.
6. Adopt one severity scale and incident lifecycle (depends on: 3, 5)
Use four impact-based levels. Classify on actual or credible customer, financial, security, regulatory, and contractual harm rather than organizational seniority.
- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; both regions impaired; or a core payment journey broadly unavailable. Page every role immediately, open a bridge, notify executives, publish customer status within 15 minutes when applicable, begin legal assessment, and require a postmortem.
- **SEV2 — major:** material payment degradation; settlement deadline at risk; a critical customer or material cohort unavailable; regional impairment with reduced resilience; or likely SLA breach. Page command and technical roles, publish customer status within 30 minutes when customer-visible, and require a postmortem.
- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads; page only when immediate action is necessary.
- **SEV4 — operational event:** no current customer impact and no urgent risk. Create a ticket and handle during normal operations.
- Treat percentages as supporting guardrails, never reasons to under-classify integrity, settlement, security, or contractual risk.
- Anyone may declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Automatically use at least SEV2 posture for suspected ledger-integrity events, cross-team incidents with unknown ownership, and materially unknown impact lasting 15 minutes.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed. Resolution requires stability, backlog recovery, and necessary reconciliation.
7. Define roles, authority, and handoffs (depends on: 6)
Separate command, communications, recordkeeping, and technical repair. One named person must hold command throughout every SEV1 and SEV2.
- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not serve as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- **Subject-Matter Responders:** diagnose and mitigate only services for which they have ownership, access, training, or formally accepted support responsibility.
- **Executive Duty Officer:** removes organizational barriers and makes exceptional business decisions without displacing the commander.
- Security, Legal, Compliance, Vendor Management, and Finance join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Require distinct commander, communications, scribe, and primary technical lead for SEV1. Communications and scribe may combine temporarily for bounded SEV2 incidents.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, and reconciliation requirements for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, with the exact transfer time.
8. Create sustainable 24x7 coverage across 28 teams (depends on: 4, 5, 7)
Use central command coverage and risk-based technical coverage rather than creating 28 fragile night rotations. All services receive a response path, but only critical domains maintain direct overnight technical rotations.
- Create a 24x7 Incident Command corps of 24–30 certified people, with primary and secondary coverage at all times.
- Create 18–24 trained Communications Leads and a similarly sized scribe pool using Support, Customer Operations, Engineering Operations, and qualified managers.
- Maintain a surge roster for simultaneous incidents and a 24x7 Executive Duty Officer schedule.
- Group Tier 0 and Tier 1 ownership into approximately 8–12 coherent domains only where responders have training, access, runbooks, and explicit acceptance.
- Staff each critical domain with primary and secondary responders and at least six qualified people. Target no more than one primary week in six.
- Give Tier 2 and Tier 3 services business-hours coverage plus a maintained manager and director escalation path.
- When lower-tier impact becomes SEV1 or SEV2, central command activates the manager escalation path and obtains the necessary owner.
- Route unknown-owner pages to Duty Command and platform triage temporarily. Log each one as a catalog control failure.
- Do not merge small teams merely for schedule convenience. Provide training, reassign services, add staffing, or decommission unsupported services.
9. Approve compensation and fatigue protections (depends on: 4, 8)
End unpaid on-call before expanding mandatory night coverage. HR, Finance, Payroll, and employment counsel should approve the policy within 14 days.
- Use market-validated weekly bands, initially budgeting approximately $800–$1,200 for Tier 0/1 primary duty and $300–$500 for secondary duty.
- Budget approximately $900–$1,300 for Duty Incident Commander weeks and $400–$800 for Communications Lead or scribe duty, adjusted for actual burden.
- Pay holiday premiums and compensate active after-hours work according to exempt or non-exempt status and applicable federal and New York rules.
- Provide a protected paid recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal sprint commitments by approximately 15% during a primary on-call week.
- Prohibit simultaneous primary assignments, consecutive primary weeks, on-call during leave, and invisible schedule swaps.
- Allow responders to declare themselves temporarily unfit after disruptive night work without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Budget roughly $0.9M–$1.2M annually, then refine using actual rotation count, employment classification, and activation data.
10. Establish one paging and incident system of record (depends on: 3, 5, 8)
Monitoring tools may remain specialized, but all human pages must enter one controlled platform. This removes conflicting schedules and creates one evidence trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page.
- Ingest events from the six existing monitoring tools before disabling their direct paging paths.
- Route pages through service-catalog ownership and deduplicate related events.
- Provide a single declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, severity changes, decisions, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status-page publishing, and conference facilities.
- Require role-based access, MFA, access reviews, and immutable or tamper-evident history.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, or an AWS region.
- Retire each legacy human-paging path only after ownership review, end-to-end tests, and two weeks of verified operation.
11. Enforce an alert-quality contract and page budget (depends on: 3, 5, 10)
Treat every page as a production interface with an owner and an expected action. Noise reduction must not create detection gaps.
- Require every paging rule to identify the service, owning team, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Page only when prompt human judgment or intervention can materially reduce customer, financial, security, or contractual risk.
- Route informational, capacity, and non-urgent infrastructure conditions to dashboards or ticket queues.
- Prefer payment-outcome, settlement-risk, queue-age, and error-budget-burn alerts over raw CPU, memory, pod, or log thresholds.
- Run new paging rules in shadow mode for at least seven days unless an emergency exception is approved.
- Review rules with repeated no-action acknowledgements, low actionability, or excessive firing within two business days.
- Set a load budget of no more than two after-hours pages per responder per week, measured over four weeks.
- Require an owner, compensating detection, and recorded approval before suppressing or deleting a rule.
- Burn down the top 50 noisy rules first. Review missed detections and noise together so teams cannot improve metrics by becoming blind.
12. Detect payment and ledger failures before customers (depends on: 5, 11)
Move detection from host health to customer journeys and financial outcomes. Use internal SLOs with enough headroom to protect the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, API access, webhooks, and reporting freshness.
- Run external synthetic transactions through critical payment journeys at least every minute and from paths independent of the production platform.
- Validate each AWS region and expose dependencies that undermine nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed value relative to settlement deadlines, not only queue depth.
- Add tenant and cohort anomaly detection for high-value customers and major payment methods.
- Convert high-priority Support, account-manager, processor, sponsor-bank, and network reports into incident candidates within five minutes.
- Record the detection source for every incident. Treat customer-first detection as a mandatory missed-detection review.
13. Codify detection, escalation, and live execution (depends on: 7, 10)
Create one time-bound path from the first credible signal to named command and mitigation. Notification delivery does not count as human acknowledgement.
- Page the owning critical-domain primary and Duty Incident Commander immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at five minutes, manager at 10, and director at 15.
- Escalate an unclaimed command page to backup commander at five minutes. The Executive Duty Officer assumes command at 10 minutes until a certified handoff occurs.
- Require Support and account managers to use the same declaration path as automated monitoring and engineers.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Begin with a standard statement of severity, known impact, assigned roles, current objective, workstreams, and next update time.
- Freeze unrelated changes during SEV1 and normally during SEV2. Record any exception.
- Prefer reversible mitigation such as rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require formal command handoff for long incidents, shift changes, or fatigue. Do not leave an incident unowned during transfer.
- Reconcile the ledger and safely drain backlogs before resolving payment or ledger incidents.
14. Standardize internal incident communications (depends on: 7, 10, 13)
Give responders one working room and stakeholders one controlled information source. Executives must not interrupt the technical command path.
- Create one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue an initial internal brief within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within 15 minutes for SEV1 and 30 minutes for SEV2.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record material decisions and outbound messages in the incident timeline.
15. Standardize status-page and customer communications (depends on: 6, 10, 14)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate impact before root cause is known.
- Publish customer status within 15 minutes of declaring a customer-visible SEV1 and within 30 minutes for customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice promptly after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map status-page components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, data-integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, known workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Give affected strategic accounts direct account-manager outreach within 30 minutes for SEV1 and 60 minutes for SEV2.
- Require account managers to use the approved briefing and prohibit independent technical explanations.
- Provide a customer-facing incident summary within five business days for SEV1 and qualifying SEV2 events.
- Record any legally necessary delay or restriction of public detail, its approver, and the alternative communication plan.
16. Operationalize legal, regulatory, contractual, and credit decisions (depends on: 3, 6, 15)
Some payments incidents start external notification clocks. Make assessment mandatory without assuming that every operational incident is reportable.
- Build a counsel-validated matrix covering applicable NYDFS rules, state breach laws, GLBA or FTC obligations, PCI requirements, money-transmitter obligations, sponsor-bank and network contracts, cyber insurance, and customer contracts.
- Engage Legal and Compliance immediately for every SEV1, security event, and suspected ledger-integrity event.
- Record an initial reportability assessment within one hour for SEV1 and within two hours for other potentially reportable events.
- Record non-reportable decisions as evidence, including facts considered, approver, timestamp, and reassessment trigger.
- Maintain tested 24x7 contacts for counsel, regulators, sponsor banks, networks, insurers, and critical vendors.
- Encode customer-specific notice deadlines and channels in the customer record used by the Communications Lead.
- Have Legal own regulatory text and submission. Keep technical command with the Incident Commander.
- Have Finance calculate affected minutes, delayed value, likely credits, and contractual exposure within five business days.
- Track credits by incident and recurring cause to support reliability investment decisions.
17. Make postmortems mandatory, consistent, and blameless (depends on: 6, 7, 10)
Use one learning standard with fixed deadlines. Keep postmortems separate from performance, misconduct, and disciplinary processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact over two hours, SLA credits, contractual breach, repeated contributing factors, control failures, and ledger-integrity near misses.
- Produce the factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the Incident Commander responsible for the timeline and the owning engineering director accountable for completion.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit analysis of why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the incident commander.
- Describe decisions in the context and information available at the time. Do not name an individual as the root cause.
- Publish broadly useful findings internally while restricting security, privacy, personnel, or privileged material appropriately.
18. Give corrective actions enforceable ownership (depends on: 10, 17)
Treat incident actions as risk commitments, not suggestions. A ticket is not complete until the expected risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Escalate overdue high-risk actions to the manager after seven days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit overdue actions to block related releases where the unrepaired condition could reproduce severe impact.
- Verify effectiveness through tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days. Complete, re-plan, or formally accept each risk.
19. Set the service-readiness bar and critical playbooks (depends on: 5, 8, 11, 12)
A critical service must be supportable at 3 a.m. before its team is placed on direct overnight coverage. Existing critical detection must remain active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and data-integrity constraints.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional failure, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document controlled failover, split-brain prevention, replay protection, backlog recovery, and post-recovery reconciliation.
- Exercise critical runbooks at least twice per year and after material changes.
- Block new Tier 0/1 releases and new paging rules when readiness requirements are missing.
- Handle existing gaps through a named owner, compensating control, executive-approved expiry date, and remediation plan rather than disabling detection.
20. Publish policy and certify every response role (depends on: 6, 7, 8, 9, 11, 13, 14, 15, 16, 17, 18)
Convert the operating design into concise, signed documents and practical training. Training and exercises occur during paid working time.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, and exception standards.
- Provide one-page severity, role, authority, escalation, and communication cards inside the incident tool.
- Train all employees to recognize and declare incidents.
- Train all engineers in severity, acknowledgement, evidence preservation, handoff, and financial-integrity precautions.
- Certify responders only after they demonstrate access, dashboards, runbooks, rollback, and escalation competence.
- Certify Incident Commanders through formal instruction, simulation, and at least two shadowed incidents or exercises.
- Train Communications Leads in status writing, account segmentation, contractual clocks, and legal escalation.
- Train scribes in timeline quality and fact-versus-hypothesis labeling.
- Require two shadow shifts before independent primary duty and annual recertification.
- Maintain the training and certification register as operational and audit evidence.
21. Pilot on the payment critical path (depends on: 9, 10, 12, 19, 20)
Run a four-to-six-week pilot across the highest-risk journey before expanding. Use real incidents and exercises to correct the model quickly.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, API edge, authentication, settlement, and Support intake.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently assuming command.
- Correct critical process or tool defects within 48 hours and update the standard visibly.
- Exit only after 95% timely command assignment, 95% communication compliance, no unpaid pages, complete required postmortems, and at least 50% lower noise.
22. Roll out by customer journey and risk (depends on: 21)
Expand in controlled waves while the interim process remains active company-wide. Complete rollout early enough to accumulate operating evidence before the audit.
- Roll out remaining Tier 0 domains first, then Tier 1, Tier 2, and Tier 3.
- Use two-to-three-week waves with a named program coach and director sign-off.
- Gate each team on catalog ownership, appropriate coverage, compensation, trained responders, tested escalation, alert quality, runbooks, access, and a passed tabletop.
- Require six or more responders only for direct 24x7 technical rotations. Apply business-hours coverage and manager escalation to lower tiers.
- Reschedule failed gates or approve a time-limited executive exception with a compensating control.
- Disable legacy human-paging paths after verified cutover for each wave.
- Publish an internal adoption dashboard by team, service tier, and control gap.
- Finish critical coverage by week 10 and all 28 teams by week 18.
23. Exercise command, communications, regional recovery, and fallbacks (depends on: 19, 20)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous security and operational events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors where appropriate.
- Conduct at least one unannounced after-hours paging test before the audit.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure time to acknowledgement, command, customer notice, mitigation decision, handoff, and recovery.
- Create tracked actions for every material exercise finding.
24. Measure performance and operate fixed review forums (depends on: 3, 10, 17, 18)
Use a balanced scorecard that exposes weak controls without rewarding hidden incidents or suppressed alerts. Report medians and 90th percentiles rather than averages alone.
- Measure time from first impact to detection, declaration, acknowledgement, command assignment, mitigation, recovery, and resolution.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track pages, actionability, duplicates, after-hours load, missed detections, and page-budget breaches.
- Track postmortem timeliness, action age, closure by due date, verified effectiveness, and recurring contributing factors.
- Track availability by customer journey, error-budget burn, failed or delayed payment value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight activations, recovery days, schedule exceptions, sentiment, and attrition.
- Hold a weekly Incident Review Board for incidents, postmortems, control failures, noisy alerts, and overdue actions.
- Hold a monthly executive reliability review for trends, funding, contractual exposure, and accepted risks.
- Hold a quarterly control and resilience review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile incident records monthly against support cases, status history, credits, and customer complaints to detect under-reporting.
25. Prove SOC 2 operating effectiveness before fieldwork (depends on: 20, 22, 23, 24)
Generate audit evidence through normal operation rather than reconstructing it later. Test both control design and consistent execution.
- Maintain approved and versioned policies, exceptions, catalog records, schedules, compensation activation, access reviews, training, incidents, communications, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, non-reportable legal decisions, missed timelines, and exercises in the population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the populations, evidence requests, and interviews expected from the external auditor.
- Record deviations honestly with owners, remediation dates, and compensating controls. Never rewrite historical records.
- Confirm that evidence retention covers the full auditor-defined observation period.
- Brief commanders, responders, Support, and Compliance on the real process without scripting inaccurate answers.
26. Institutionalize improvement and reduce structural risk (depends on: 22, 24, 25)
Prevent the program from decaying after the audit. Use incident evidence to drive permanent ownership and architectural investment.
- Assign permanent owners for policy, service catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review severity thresholds, staffing, compensation, and communication timings annually and after material process failures.
- Recertify command and communications personnel annually.
- Review recurring failure families quarterly and require executive decisions where corrective work repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fairness, fatigue, and tool friction.
- Report severe incidents, credits, overdue high-risk actions, and resilience investment to the board or risk committee quarterly.
- Treat the shared ledger cluster as a strategic concentration risk. Fund failover assurance, blast-radius reduction, isolation of non-critical readers, and stronger regional independence.
- Evaluate follow-the-sun command or technical coverage using one year of page-load and staffing data.
- Build year-two plans for automated mitigation, deployment safety, graceful degradation, and error-budget release controls.
Previous Proposal 3 (ID: f7ee763b-3b76-4f07-9e7d-5164701bef1c, Agent: qwen3.8-max_refine_3, LLM: alibaba/qwen3.8-max):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (32):
1. Establish executive mandate, program office, and funding
Turn the CEO email into a **chartered company program** with one accountable owner and authority over all 28 teams.
- Appoint the CTO as executive sponsor and a Head of Reliability as program owner with full-time authority.
- Create a permanent program office: one program lead, one platform engineer, one analyst.
- Form an eight-person steering group: Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours.
- Lock the non-negotiables: one severity scale, one paging platform, one postmortem format, paid on-call, mandatory action tracking, named service ownership.
- Approve budget anchored against the $1.3M in credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs.
- Reserve 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Publish a one-page charter stating incident response is a company operating process, not a per-team choice.
2. Install a seven-day interim command bridge (depends on: 1)
Do not leave the company unprotected while the permanent process is designed. Put a **crude but real** command structure in place within one week.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Direct Support to escalate credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical actions: complete, re-plan, or formally risk-accept, prioritizing ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Build the forensic baseline of incidents, alerts, and money lost (depends on: 1)
Rebuild the facts before designing anything. This becomes both the **design input** and the before picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific missing signal. This list drives the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Freeze baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the on-call fairness deal (depends on: 1)
Engineer pushback is the **largest delivery risk**. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS, and Sales in two weeks. Separate the real objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey covering fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
5. Map compliance, evidence, and audit requirements from day one (depends on: 1)
Design evidence as a **by-product of operations**, not a reconstruction before the auditor arrives. Confirm the SOC 2 observation window immediately.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for communications; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure, training register, drill records, reportability decisions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve all policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has a **named owner**. Build a machine-readable catalog as the single source of truth.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact: Tier 0 (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for each Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
7. Adopt the severity scale, declaration rules, and lifecycle (depends on: 3, 6)
Replace judgment calls with a **lookup table**. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter.
- SEV1 (crisis): money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of all roles and executive; bridge in 5 minutes; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 (critical): material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- SEV3 (major, contained): narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected, Declared, Triaged, Mitigated (customer impact ends), Monitoring, Resolved (backlog processed and ledger reconciled), Reviewed.
- Publish a decision tree with 12 worked examples from the real 31 incidents.
8. Define incident roles, authority, and handover discipline (depends on: 7)
Solve nobody-in-charge-for-an-hour by making command **explicit, single-holder, transferable, and logged**. Separate coordination from debugging.
- Incident Commander: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates for SEV1/SEV2.
- Subject-Matter Responders: engineers of the owning team; they mitigate, they do not run the room.
- Executive Duty Officer (SEV1): removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Design the three-layer 24x7 coverage model (depends on: 6, 8)
Do not create 28 night rotations. **Centralise coordination** in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- Layer A, Incident Command corps: approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of approximately 18 from Support, CS, and engineering management. A scribe pool used as the training entry point.
- Layer B, critical-path domain rotations: consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- Layer C, everyone else: business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing: a US-only paid night rotation now, a follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
10. Approve on-call compensation, labor compliance, and fatigue safeguards (depends on: 9, 4)
Unpaid on-call in New York is both a **retention problem and a legal exposure**. Pay must be live in payroll before any mandatory night rotation starts.
- Indicative scheme: approximately $1,000 per primary 24x7 week, $400 secondary, $250 for business-hours rotations, a separate $1,200 Duty Commander stipend, holiday premiums, approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- HR, Finance, and employment counsel publish amounts, eligibility, tax treatment, and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a 15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
11. Set the alert quality standard, page budget, and noise burn-down (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes. Make alert quality a **condition of being allowed to page a human**.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly.
- Target 3,400 to under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
12. Build payment-outcome detection and ledger assurance (depends on: 11, 6)
Stop customers telling you first. Detection must be driven by **payment outcomes and ledger truth**, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer detected first becomes a named defect class with a mandatory tracked action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Consolidate to one paging platform, one incident record, one status page (depends on: 7, 9, 11)
Collapse six alerting tools into a **single operational system of record** so there is one queue, one timeline, and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is unavailable. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
14. Codify the escalation ladder and five-minute command rule (depends on: 8, 9, 13)
Write one unskippable path from something looks wrong to **someone is in charge**. The default action is never waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately, secondary at 5 minutes, domain manager and Duty Commander at 10, Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension, and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
15. Write the live incident execution doctrine and major-incident playbooks (depends on: 8, 13, 14)
Give responders one short operating procedure from the first minutes through closure. Priority is **limiting customer and financial harm**, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the single largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window, an explicit handback to the owning team, Support, and Customer Success, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
16. Standardize internal communications protocol (depends on: 8, 13)
Standardise the internal picture so executives, Support, and Sales are informed **without interrupting** the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
17. Build customer communications, status page, and account-manager outreach (depends on: 7, 16)
Replace whoever is around with a **timed, owned, pre-approved process**. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook, and RSS. All 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity, or blame.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
18. Create the regulatory, partner, and legal notification playbook (depends on: 7, 17)
In payments, some incidents start a **legal clock at detection**. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded even when the answer is not reportable, with evidence, decision-maker, and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
19. Operationalize SLA credit and financial impact workflow (depends on: 7, 17)
Tie incidents to money so severity, credits, and investment decisions **stay honest**, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one. Use the delta as the standing business case for on-call pay and reliability capacity.
20. Establish the blameless postmortem standard and Incident Review Board (depends on: 7, 8)
Replace some incidents, various formats with **one mandatory format, fixed deadlines, and a forum with teeth**.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly top-five recurring causes analysis.
21. Enforce action ownership, capacity reservation, and tracking (depends on: 20, 13)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of **customer commitments** and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
22. Build the training and certification academy (depends on: 8, 14, 16, 20)
Command is a skill, not a title. **Certify before assigning duty**, and use paid working time for all of it.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days + shadowing): command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Comms Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum; never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Run the exercise programme: tabletops, game days, and unannounced drills (depends on: 22, 13, 15)
The process must meet a **simulated SEV1** before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger. Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (depends on: 7, 8, 9, 10, 11, 14, 16, 17, 18, 20, 21)
Collapse the design into a document people will **actually open mid-outage**, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 12, 13, 15, 22, 10)
Prove the process on the **highest-risk surface** with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment.
- Publish a one-page result to the whole company.
26. Run the alert noise burn-down campaign (depends on: 11, 13, 25)
Run noise reduction as a **visible, quota-driven campaign** in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness. Review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 to 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
27. Establish metrics, dashboards, review cadence, and anti-gaming (depends on: 13, 20, 25)
Instrument the process itself so improvement is visible and the auditor sees **evidence of monitoring and review**. Report median and 90th percentile, never averages alone.
- Response: time to detect, declare, commander, acknowledgement, mitigate, resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
28. Wave rollout to all 28 teams with readiness gates (depends on: 25, 27)
Roll out in four waves of six to eight teams every two to three weeks, ordered by **customer risk**. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. Failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Publish a live adoption scoreboard.
29. Run change management, fairness, and pager culture programme (depends on: 4, 10, 25)
Run this from day one in parallel. Engineers judge the process on **fairness**; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. Build SOC 2 evidence by design, internal testing, and mock audit (depends on: 24, 27, 28)
Design the evidence as a **by-product of doing the work**. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any change as an exception.
31. Maintain the program risk register and contingencies (depends on: 1)
Name the ways this programme fails and **pre-commit the response**. Review it monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 27, 28, 30)
Guard against the classic failure: the process **decays once the audit is signed**. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 4 (ID: 101b9773-a36a-47bf-8af6-14fe9a48abb2, Agent: grok4.6_refine_4, LLM: xai/grok-4.6):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (26):
1. Charter the program and start the audit clock
Convert the CEO email into a company operating process with one owner, a budget, and an observation window that starts this week.
Incident response is no longer a per-team choice.
- Name the CTO as sponsor and a Head of Reliability as the single accountable owner, with a three-person program office.
- Form a small decision group: Engineering, SRE, Support, Customer Success, Security, Legal, Finance, and HR. The sponsor decides within 48 hours.
- Lock non-negotiables: one severity scale, one paging path, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, and reserved engineering capacity against the $1.3M in SLA credits.
- Timeline: operating floor in 7 days, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, mock audit in month 7, audit in month 8.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for new tools or the final policy. Put a crude but real process in place this week so the next outage has a named commander.
This week is the first audit evidence.
- Publish a one-page interim severity card and one declaration path: Slack command, phone number, and existing pagers, all reaching the same duty person.
- Staff interim primary and backup Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in the incident channel.
- Tell Support to declare from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline (depends on: 1)
Rebuild the facts before locking design. This is both the design input and the before picture for the CEO and the auditor.
Freeze the numbers so they cannot drift during design.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisy rules and every rule with no owner or runbook.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11 of 64 actions closed, on-call in 12 of 28 teams.
4. Map resistance and publish the fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
Publish the deal in writing before any new mandatory pager is assigned.
- Interview all 28 teams plus Support, Customer Success, and Sales in two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the **fairness contract**: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and money-path tiers (depends on: 3)
You cannot page the right person across 180 services until every service has one owner. A wrong owner recreates the pager objection.
The catalog is the single source of truth for paging, impact, status-page components, and audit.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to services, data stores, regions, and third parties, including sponsor banks and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
6. Lock severity levels and what each one triggers (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by who reported it or how hard the fix looks.
Anyone may declare. Nobody is punished for over-declaring. Only the Incident Commander may downgrade, with the evidence recorded.
- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 major, contained: narrow or single-customer impact with a workaround. Owning team leads; business-hours comms; postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages.
- Attach a financial-integrity or security flag to any severity. The flag forces dual-control, Legal, and the regulatory checkpoint without inventing a fifth level.
- Auto-escalate: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer impact ends) → Monitoring → Resolved (backlog processed and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged.
Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. A human validates for SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1 and SEV2. Every handover is announced verbally and in writing with the exact time.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
8. Staff 24x7 with a command corps, not 28 night rotations (depends on: 5, 7)
Do not create 28 night rotations. That is what engineers are rejecting.
Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified volunteers from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to seven months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate 28 teams into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat Lisbon or APAC follow-the-sun as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
9. Pay on-call, meet New York labor rules, and cap fatigue (depends on: 4, 8)
Unpaid on-call in New York is a retention problem and a legal exposure. Pay must be in payroll before any new mandatory night rotation starts.
Publish the numbers. Then ask people to sign up.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
10. Set the alert-quality bar and a hard page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them.
Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers do (depends on: 5, 10)
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail.
Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One Slack command creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is down. Test weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Write the five-minute path from signal to command (depends on: 7, 8, 12)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting.
If nobody claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Codify live execution, readiness, and ledger playbooks (depends on: 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving a root cause.
A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down. No command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers.
15. Run one communications clock for internals, customers, and regulators (depends on: 6, 7, 12)
Replace whoever is around with one timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is not reportable.
- Obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
16. Tie incidents to SLA credits and true financial cost (depends on: 6, 15)
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
Credit calculation is an output of the incident record, not a negotiation.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
17. Make blameless postmortems mandatory and consistent (depends on: 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth.
The discipline lives in the deadlines and the review, not the template.
- Mandatory for every SEV1 and SEV2, any incident detected by a customer first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed.
18. Track actions as risk commitments with reserved capacity (depends on: 12, 17)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
Closing a ticket without evidence of effectiveness does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Target 90% of high-priority actions closed by due date within two quarters.
19. Train and certify every role before independent duty (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
The certification register is an audit artefact.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
20. Rehearse with tabletops, game days, and night drills (depends on: 12, 14, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
21. Publish Policy v1 and the signed fairness contract (depends on: 6, 7, 8, 9, 10, 13, 15, 17, 18)
Collapse the design into a document people will actually open mid-outage, and make it official.
Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
22. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 19, 21)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse. Publish a one-page result to the whole company.
23. Instrument metrics, reviews, and anti-gaming (depends on: 12, 17, 22)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review.
Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
24. Roll out in risk-ordered waves with readiness gates (depends on: 22, 23)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave, not as a background hope. Give every team its noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
25. Produce SOC 2 evidence by operating, then mock-audit (depends on: 21, 23, 24)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions. Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (depends on: 23, 24, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
Name the ways this programme fails and pre-commit the response.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 5 (ID: 289def1b-8f78-402b-85b3-0d9e18661842, Agent: deepseek-v4-pro_refine_5, LLM: deepseek/deepseek-v4-pro):
Estimated Complexity: high
Success Metrics: - Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, no loss of Tier 0/1 detection.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% Tier 0/1 services have named owning team, tier, escalation policy, dashboard and runbook by day 30; all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10-12 domain rotations of at least six trained responders.
- 30 certified incident commanders and 18 certified communications leads active.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% required postmortems drafted in 3 business days, reviewed in 5, published in 10 from month 2.
- Postmortem action closure rises from 17% to 90% closed by due date within two quarters; all 53 legacy actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k annualised within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% SEV1/security SEV2 including not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two cross-company exercises completed before audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% sampled incidents; SOC 2 Type II incident response controls pass zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days; no increase in attrition among responders.
Steps (26):
1. Executive mandate, budget, and governance
Secure a written CTO/CEO charter making incident management a company operating process, not a team option. Name one accountable Head of Reliability and a small steering group with authority to decide within 48 hours.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking, and protected engineering capacity.
- Approve budget against the $1.3M annual credits: tooling $150-250k/yr, on-call compensation, training, and 2-3 program FTEs.
- Set timeline: interim process in week 1, design weeks 1-6, pilot weeks 7-12, rollout weeks 13-24, mock audit month 7, SOC 2 at month 8.
- Freeze current baselines: 31 incidents, 22-min MTTD, 40% customer-first detection, 3h10 MTTM, $1.3M credits, 3,400 alerts/month at 85% noise, 11/64 actions closed.
2. Interim 7-day command (depends on: 1)
Put a crude but real process in place immediately so no incident remains unowned while the permanent design proceeds.
- Publish a one-page interim severity card, one declaration path (phone, Slack command, pager), and one incident channel/bridge/timeline naming convention.
- Staff an interim 24x7 duty commander with primary and backup from engineering managers and senior SREs; compensate retroactively under final policy.
- Require a named incident commander within 10 minutes of any suspected major incident, announced in channel.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions in 30 days, prioritizing ledger integrity, duplicate payment, regional failover, security and detection gaps.
- Hold a daily 15-minute operations review until the permanent process is live.
3. Forensic baseline and alert estate analysis (depends on: 1)
Reconstruct the true before picture from the last 12 months; it drives design and serves as audit baseline.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact start/detect/declare/commander/mitigate/resolve, who led, credits paid, contributing factors.
- For every customer-first detection, name the missing signal; this becomes the detection backlog.
- Reconstruct the two `nobody in charge` incidents minute by minute; use them as the burning-platform narrative and test case.
- Profile the 3,400 monthly alerts by tool, team, rule, outcome; identify top 50 noisy rules and every rule with no owner or runbook.
- Freeze all baseline metrics in one signed document for the executive and the auditor.
4. Listening tour and fairness contract (depends on: 1, 3)
Treat engineer pushback as the main delivery risk and convert it into a written deal that removes the objection to carrying a pager for other teams' code.
- Interview all 28 teams plus Support, CS and Sales in two weeks to separate objections: unpaid work, nights, unfamiliar code, missing runbooks, or fear of blame.
- Harvest practices from the 12 teams already on-call; they supply pilot teams and first commanders.
- Publish the deal: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Recruit 10-15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey and repeat at 60, 120 and 365 days.
5. Service ownership catalog and criticality tiers (depends on: 3)
Create a machine-readable catalog as the single source of truth for paging, routing, status page components, and audit evidence.
- Assign one accountable team, engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list per service.
- Tier by business impact: Tier 0 money movement/ledger/auth/shared PostgreSQL; Tier 1 customer-facing degradable; Tier 2 internal/batch; Tier 3 non-critical.
- Map customer journeys (initiate, authorize, settle, reconcile, report, onboard) to services, data stores, regions, third parties.
- Give orphan services an owner within 30 days or a decommission date approved by the sponsor; Tier 0 without an owner is an executive escalation.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
6. Severity scale, declaration rules, and automatic triggers (depends on: 3, 5)
Adopt one severity scale as a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade with recorded rationale. Start high when uncertain.
- SEV1: money moved wrongly/duplicated/lost, ledger integrity in doubt, confirmed security/data exposure, both regions impaired, payment processing halted. Full role page, bridge in 5 min, status page in 15 min, regulator assessment within 1 hour, mandatory postmortem.
- SEV2: material degradation of a payment journey, settlement window at risk, large/strategic customer fully down, SLA breach likely. Commander and SMEs paged, status page in 30 min, mandatory postmortem.
- SEV3: narrow or single-customer impact with workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any ledger cluster incident, any SEV3 open >2h, any unknown impact after 15 min, any cross-team incident become at least SEV2.
- Publish a decision tree with 12 worked examples from the 31 real incidents.
7. Incident roles, authority, and handoffs (depends on: 6)
Codify roles so coordination never depends on seniority or heroics. The Incident Commander coordinates; responders fix only services they own.
- IC: owns severity, priorities, roles, cadence, closure; pre-authorized to freeze deploys, rollback, disable features, shift traffic, invoke failover, put ledger read-only, commit spend; does not type in terminals; keeps command when VP joins.
- Communications Lead: sole author for status page, account managers, executives, and hand-off to Legal for regulators; speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, state changes; human validates for SEV1/SEV2.
- Subject-Matter Responders: diagnose and mitigate only their owned services.
- Executive Duty Officer (SEV1): removes obstacles, shields IC from exec questions, owns board/regulator escalation; does not take command unless formally transferred.
- Rule: IC claimed within 5 min and announced; distinct people for IC, Comms, technical lead at SEV1/SEV2; every handover announced verbally and in writing with time.
- Financial controls survive: IC coordinates ledger recovery but cannot bypass dual control, reconciliation or privileged access.
8. 24x7 command and responder coverage model (depends on: 5, 7)
Do not create 28 fragile night rotations. Centralize coordination in a trained command corps, keep technical ownership local.
- Layer A: incident command corps of ~30 certified volunteers with primary/secondary 24x7; paired Comms Lead pool ~18 and scribe pool as entry.
- Layer B: consolidate 28 teams into 10-12 critical-path domains (ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes, data/reporting, integrations) each with 24x7 primary+secondary and minimum six trained responders.
- Layer C: all other teams business-hours on-call with written after-hours escalation lists held by managers.
- Platform on-call is safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate follow-the-sun only as a 12-month option, not year-1 dependency.
- Acknowledgement discipline: 5 min at SEV1/SEV2 with automatic failover.
9. Paid on-call and fatigue safeguards (depends on: 8, 4)
Unpaid on-call in New York is a retention and legal risk. Pay must be in payroll before any mandatory night rotation starts.
- Indicative: ~$1,000 per primary 24x7 week, secondary ~$400, business-hours ~$250, duty commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond one hour.
- Mandatory recovery: paid recovery day after >2h overnight work, SEV1, or qualifying SEV2; managers arrange coverage.
- HR/Legal/Finance publish amounts, tax treatment, FLSA/NY wage-hour status within 14 days.
- Load rules: no primary more often than 1 in 6, no consecutive primary/secondary weeks, no primary on two rotations, no on-call during PTO, annual cap.
- If a rotation averages >2 out-of-hours pages per person per week for 4 weeks, trigger staffing or alert-remediation review.
- Count on-call as ~15% delivery load; document exemption path for health/caring without career penalty.
10. Service readiness bar and runbooks (depends on: 5, 8)
A service must earn the right to page a human at 3 a.m. Define a minimum readiness bar and major incident playbooks before any Tier 0/1 service goes live with on-call.
- Require for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switch, escalation contacts, data-loss/latency impact.
- Write playbooks for top failures: shared PostgreSQL failure/corruption, cross-region failover, Kubernetes control-plane loss, processor/sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, duplicate payment.
- Treat ledger cluster as largest structural risk: documented failover, read-only degraded mode, reconciliation after recovery, executive-signed RPO/RTO.
- Start a parallel workstream on blast-radius reduction: tenant/function partitioning, read replicas, isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging at night unless manager accepts a dated exception in writing.
- Runbooks are peer-reviewed, version-controlled, marked stale if not exercised twice a year.
11. Alert quality standard and page budget (depends on: 3, 5)
Make alert quality a condition of paging a human. Burn down noise deliberately rather than by mass silencing.
- Every paging alert must declare owning team, affected service, customer/SLO impact, expected action, dashboard, runbook, dedup key, severity mapping, escalation policy. Anything failing becomes a ticket.
- Page on symptoms of customer harm (SLO burn rate, payment success rate, queue age vs settlement deadlines), not raw CPU/memory.
- Run new alerts in shadow mode 7 days; test firing and recovery.
- Set page budget of 2 out-of-hours pages per person per week; breach blocks new alert creation and triggers tuning sprint.
- Auto-quarantine rules firing >5 times/month without action or >70% no-action acks; return to owner with correction deadline.
- Never silently disable: verify compensating detection, record decision, name owner, review missed detections monthly.
- Target 3,400 to <500 pages/month, actionability >75% within six months, no loss of Tier 0/1 detection.
12. Detection uplift on money path (depends on: 5, 11)
Stop customers telling you first by detecting on payment outcomes and ledger truth.
- Define SLOs and business SLIs per customer journey: initiation success, authorization latency, settlement timeliness, reconciliation break rate, API availability, reporting freshness; set internal targets stricter than 99.95%.
- Run external synthetic end-to-end payments every 60 seconds from both regions, covering all critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate detection, replication lag, failover-readiness, settlement-window countdown.
- Add per-customer anomaly detection for top 100 accounts; five similar support tickets in 10 minutes auto-creates triage incident.
- Route partner and processor notifications into declaration path within 5 minutes.
- Record detection source on every incident; treat `customer detected first` as a named defect class with mandatory action.
13. Consolidate to one paging and incident platform (depends on: 6, 8, 11, 12)
Collapse six alerting tools into one operational system of record: one queue, one timeline, one audit trail, without creating a monitoring gap.
- Select one paging/scheduling platform, one incident record, one hosted status page; time-box selection to 2 weeks.
- Ingest all six sources first, deduplicate/correlate; retire legacy path only after named owners, successful end-to-end test, and 2 weeks verified operation.
- One-command declaration in Slack creates channel/bridge, pages commander, sets severity, opens timeline, starts clock.
- Route every page through service catalog: service label -> owning domain schedule -> escalation policy.
- Capture evidence automatically: declaration, acks, role assignments, severity changes, decisions, comms, mitigated/resolved times, postmortem link; retain 12 months.
- Test out-of-band SMS/phone paging, mobile fallback, offline runbook weekly; ensure works during one AWS region/chat/identity provider failure.
- Set hard date after which pages outside this tool create no on-call obligation.
14. Escalation paths and acknowledgement SLAs (depends on: 7, 8, 13)
Write one unskippable path from signal to named commander in under 5 minutes; default action is never waiting.
- Converge all entry points on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately -> secondary at 5 min -> domain manager + Duty Commander at 10 -> Executive Duty Officer at 15. SEV2 requires command assigned within 15 min.
- If no one claims command in 5 min, platform assigns and announces it; assignee may hand over but not decline.
- Commander may page any domain's on-call directly with 10-min ack obligation; this reciprocity makes single-team ownership viable.
- If ownership unclear after 10 min, commander keeps incident and names temporary owner; missing catalog entry logged as control defect.
- Pre-authorize regional failover, ledger read-only, payment suspension, partner-bank notification so IC never waits for an executive.
15. Standardize internal and customer communications (depends on: 6, 7, 13)
Replace whoever is around with a timed, owned, pre-approved process; the Communications Lead is single author and never writes from scratch under pressure.
- Internal: one working channel/bridge plus one read-only broadcast for executives/Support/Sales; cadence SEV1 15-30 min, SEV2 60 min even if nothing changed; fixed template with impact, action, next update, IC/Comms names.
- Executives ask questions only of Executive Duty Officer; publish this as signed behavior rule.
- Status page: SEV1 within 15 min, SEV2 within 30 min; updates every 30/60 min; resolution notice within 30 min of verified recovery; customer-facing summary within 5 business days for SEV1/qualifying SEV2.
- Pre-approve 12-15 templates with Legal for degradation, settlement delay, API errors, regional failure, data-integrity investigation, security event.
- Top 100 accounts get named account manager call/email within 30 min of SEV1 with briefing pack; all 2,100 subscribed by default.
- Language rules: state impact and next update; never speculate on cause, recovery time, data integrity or blame.
16. Regulatory, partner, and financial impact workflows (depends on: 6, 15)
In payments some incidents start a legal clock at detection. Build obligation assessment into the process and tie incidents to money.
- Legal/Compliance produce obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC, PCI, money-transmitter, sponsor-bank/card-network windows, cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and security-related SEV2 within 1 hour, recorded even when `not reportable`, with decision-maker and evidence.
- Maintain 24x7 contact matrix for regulators, sponsor banks, networks, outside counsel, insurer.
- Agree availability measurement method per contract; compute affected minutes per customer from incident record and journey telemetry; propose credit schedule within 5 business days.
- Attribute credits to root-cause families; track failed payment volume, delayed value, reconciliation breaks; target annual credits from $1.3M to under $400k.
17. Blameless postmortem standard and action tracking (depends on: 6, 7)
Replace various formats with one mandatory format, fixed deadlines, and enforceable actions. Closing a ticket without evidence does not close the action.
- Mandatory for every SEV1/SEV2, customer-first detection, incident >2h, repeat of known cause, credit-generating/contract breach, ledger near-miss.
- Draft within 3 business days, review within 5, publish within 10; IC owns delivery, owning manager accountable.
- One template: summary, customer/financial impact, detection source/gap, timeline, response analysis, contributing conditions, what worked, actions.
- Two mandatory questions: why did a customer see this first? why did mitigation take as long as it did?
- Blameless in writing: systems and context, no individual named as cause, never used in performance reviews; HR handles personnel/misconduct separately.
- Every action gets owner, priority, due date, verification method, ticket auto-created; classes containment 7 days, corrective 30, strategic 90; reserve 15% engineering capacity.
- Escalation: manager at +7 days, director at +14, CTO dashboard at +30; overdue P0 can block related releases.
- Weekly Incident Review Board ratifies severity, challenges quality, monitors actions; target 90% closure on time within two quarters.
18. Train and certify incident roles (depends on: 7, 13, 15)
Command is a skill, not a title. Certify before duty; use paid working time.
- All employees: 30-min module on recognizing impact, declaring, finding channel/status page.
- Responder: half-day on severity, escalation, runbook, financial integrity; mandatory before joining rotation.
- Scribe: 2 hours timeline discipline; entry point.
- Incident Commander: 2 days + shadowing; command presence, delegation, severity calls, running room, handover; certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead: 1 day on status writing, customer tiering, legal boundaries, regulator triggers.
- Domain responders demonstrate dashboard, runbook, rollback, failover, access before primary; two shadow shifts; no primary in first 90 days.
- Certification valid 12 months, renewed by simulation; register is audit artifact.
19. Simulations and game days (depends on: 10, 14, 15, 18)
Rehearse process before real SEV1; exercise tooling failure and ledger scenarios.
- Monthly tabletop per group reusing an incident from baseline, rotating commander.
- Quarterly game day in staging or tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes degradation.
- Twice-yearly unannounced paging drill to measure real overnight ack times.
- Annually one combined operational/security exercise and one regulator-notification exercise with Legal/CISO.
- Exercise status page down, chat/paging provider unavailable.
- Never inject uncontrolled change into production ledger; validate backups/restore/RPO/RTO on replicas.
- Every exercise produces tracked actions; publish time-to-commander, time-to-first-update, time-to-mitigation.
20. Pilot on critical path (depends on: 9, 11, 12, 13, 14, 15, 17, 18, 19)
Prove full process on highest-risk surface with willing teams for six weeks before full rollout.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, Support intake, including two of the 12 already on-call teams.
- Activate severity scale, command corps, single paging tool, page budget, status page policy, mandatory postmortems, paid rotations.
- Parallel old paths for one week then cut over; real incidents use new process only.
- Program lead attends every SEV2+ as coach, never shadow commander; review every pilot page within one business day.
- Exit gate: commander named within 5 min in 95% cases, first status update on time, MTTD under 10 min for pilot services, pages halved, no unpaid page, all required postmortems on time, positive sentiment.
- Publish one-page result company-wide.
21. Metrics, dashboards, and review cadence (depends on: 13, 17)
Instrument the process so improvement is visible and audit sees monitoring/review evidence. Report median and p90, never averages alone.
- Response: time to detect/declare/commander/ack/mitigate/resolve split by severity, tier, journey, region, detection source.
- Quality: customer-first rate, status page timeliness, update cadence, missed escalations, role conflicts, alert actionability, out-of-hours pages per person.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat factors.
- Business: availability per journey, error budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, frequency, recovery days, sentiment, attrition.
- Weekly Incident Review Board; monthly reliability review per team; monthly executive review to CEO; quarterly control review with Security/Compliance/Risk; quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, status page history; team scorecards direct help, never individual penalties.
22. Wave rollout to all 28 teams with readiness gates (depends on: 20, 21)
Roll out in four waves by customer risk, each passing an explicit gate rather than a date. Complete all teams by week 24 to leave ~3 months of evidence before audit fieldwork.
- Wave 1 remaining Tier 0, Wave 2 Tier 1, Wave 3 Tier 2, Wave 4 Tier 3/internal.
- Per-team onboarding kit: catalog entry complete, alerts migrated within budget, runbooks at readiness bar, rotation staffed with 6 trained responders or time-limited exception, one IC candidate nominated, one tabletop passed, payroll active.
- Named coach per wave for three weeks; director signs gate.
- Legacy paging paths disabled per wave, not kept as comfort fallback.
- Publish live adoption scoreboard.
- Any team unable to staff fair rotation gets headcount, service reassignment, or explicit executive risk acceptance.
23. Culture, fairness, and continuous feedback (depends on: 4, 9)
Run this in parallel from day one. Engineers judge fairness; executives judge results.
- Repeat the deal in every forum: paid on-call, paged only for owned services, trained commander, real sprint capacity.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
- Weekly office hours first 8 weeks, Slack support channel 4-hour SLA, per-team champion, weekly newsletter with metrics and bad news.
- Incident response contribution in promotion criteria; quarterly awards for best postmortem and biggest noise reduction; public thanks after SEV1.
- Pulse survey at 60, 120, 365 days; if fairness or load red, pause expansion until fixed.
24. SOC 2 evidence design and internal dry-run (depends on: 13, 17, 21, 22)
Make evidence a by-product of operations; test before external auditor.
- Map with Compliance to Trust Services Criteria CC7.2-CC7.5, CC2.2/CC2.3, A1.2; confirm observation window early.
- Evidence set automated and indexed: versioned policies/exception, rotation schedules, compensation activation, incident records with timestamps/roles, paging/ack logs, status history, reportability decisions, postmortems, action closure with verification, training/certification register, drill records, access reviews.
- Internal Audit tests process in months 4 and 6, sampling incidents end-to-end.
- Run formal mock audit month 7 using same evidence populations; include interviews with IC, engineer, Support, Compliance.
- Correct failures through tracked actions, never by editing history.
- Freeze process wording after month 7; any change logged as exception.
25. Risk register and contingency planning (depends on: 1)
Name likely failure modes and pre-commit responses; review monthly with sponsor.
- Too few commander volunteers: make command rostered duty for managers and staff engineers until pool reaches 30.
- Compensation not approved in time: fallback to time-off-in-lieu plus phased stipend, never launch mandatory night on-call uncompensated.
- Tool migration slips: cut scope on incident record layer, never on paging consolidation.
- Noise pruning hides real failure: demote to ticket first, observe 30 days, keep recovery path, monthly missed-detection review.
- Burnout in experienced teams: weekly load monitoring with page caps.
- Major SEV1 mid-rollout: program lead becomes full-time responder, wave schedule slips one wave, sponsor told same day.
- Shared ledger concentration: if blast-radius workstream slips, escalate to board as accepted risk with dated plan.
26. Inspect and adapt; year-two sustainability (depends on: 22, 24)
Prevent decay after audit by revising on data and assigning permanent owners.
- At 90 days live, revise policy using measurements: severity calibration, Layer B/C membership from page data, uncovered shifts, commander burn, missed updates, action closure, survey results.
- Assign permanent owners for policy, paging platform, status page, service catalog, training, metrics.
- Re-baseline targets every six months; shift from lagging to leading metrics: error budget burn, near-miss rate, drill performance.
- Year-two candidates: follow-the-sun coverage, automated mitigation for top recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, completion of ledger blast-radius reduction.
- Keep annual policy review, certification renewal, exercise calendar, board reporting as permanent commitments.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Charter the programme: one owner, one mandate, funded, dated against the audit", "description": "Turn the CEO email into a chartered company programme with a single accountable owner and authority over all 28 teams. Incident response stops being a per-team preference and becomes a **company operating process**.\n\n- Name the CTO as executive sponsor and a full-time Head of Reliability & Incident Management as accountable owner, supported by a programme office of three: programme lead, incident-platform engineer, reliability analyst.\n- Form a decision group (Engineering, Platform/SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.\n- Lock the non-negotiables on day one: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, named service ownership, and **paid on-call**.\n- Publish the timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal control test month 5, mock audit month 7, SOC 2 fieldwork month 8.\n- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.3M/yr, training and exercises, plus 15% of engineering capacity reserved for reliability work.\n- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.", "dependencies": []}, {"step_id": "S2", "title": "Seven-day operating floor so the next outage already has an owner", "description": "Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence immediately.\n\n- Publish a one-page interim severity card and a single declaration path: one chat command, one phone number, one pager service, all reaching the same duty person.\n- Stand up an interim 24x7 Duty Incident Commander roster (primary plus backup) from engineering managers and senior SREs. Pay it retroactively under the final policy.\n- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.\n- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.\n- Use one channel, one bridge, one timeline document and one naming convention for every major incident, starting now.\n- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.\n- Hold a 15-minute daily operations review until the permanent process is live.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Forensic baseline of incidents, alerts and money lost", "description": "Rebuild the facts before designing anything. This is the design input, the executive narrative and the frozen \"before\" picture for the auditor.\n\n- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.\n- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the **detection backlog**.\n- Reconstruct the two \"nobody in charge\" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.\n- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.\n- Correlate incidents with deployments, config changes and feature flags to quantify how many started with a change we made.\n- Quantify true cost beyond credits: failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost, churn risk on named accounts.\n- Freeze the baselines in a signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Audit, legal and evidence scoping in week one", "description": "Design evidence as a by-product of operating, never as a reconstruction before fieldwork. Settle the scope and the legal handling of incident records now, not in month seven.\n\n- Confirm with the external auditor the Type II observation window, the incident population definition, and the evidence they will sample. Everything from the day-7 floor onwards must count.\n- Map incident response to the Trust Services Criteria with Compliance: CC7.2–CC7.5 (monitoring, identification, response, recovery), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), A1.2 (availability).\n- Define the evidence set and where it is produced automatically: incident records, paging and acknowledgement logs, role assignments, status-page history, reportability decisions, postmortems, action closure with verification, training register, drill records, access reviews.\n- Agree retention, confidentiality, legal hold and access rules. Decide with counsel which postmortem content is privileged and how privileged material is segregated **without making the ordinary postmortem secret**.\n- Start the obligation matrix with Legal: NYDFS 23 NYCRR 500, state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network contractual windows, cyber-insurer notice.\n- Begin monthly evidence sampling from month 1 so control operation is visible long before the audit.", "dependencies": ["S1"]}, {"step_id": "S5", "title": "Listening tour, resistance map and the written on-call deal", "description": "Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written deal signed by the sponsor.\n\n- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.\n- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.\n- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.\n- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.\n- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.\n- State the hard gate publicly: no new mandatory night rotation starts before compensation, training, runbooks and staffing rules are live.", "dependencies": ["S1", "S3"]}, {"step_id": "S6", "title": "Service catalog, ownership, tiering and customer-journey map", "description": "You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.\n\n- Record for every service: one accountable team, engineering manager, chat channel, escalation policy, dashboards, runbook link, dependency list, regions and data stores.\n- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.\n- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.\n- Record SLO, RTO, RPO, active-active or region-bound, failover method, and the dependencies that make nominal two-region redundancy fake.\n- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.\n- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.", "dependencies": ["S3"]}, {"step_id": "S7", "title": "Severity standard, declaration rights and incident lifecycle", "description": "Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.\n\n- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning responders and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment started within 1 hour; mandatory postmortem.\n- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and responders paged, comms and scribe if customer-visible, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.\n- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, credit-generating or a repeat.\n- **SEV4:** no customer impact. Ticket only; never pages a human.\n- Attach a financial-integrity flag or a security flag to any severity. The flag forces dual control, Legal engagement and the reportability checkpoint without inventing a fifth level.\n- Anchor on payments reality alongside error rates: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.\n- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.\n- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.\n- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including a customer report.\n- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.", "dependencies": ["S3", "S6"]}, {"step_id": "S8", "title": "Roles, authority, concurrency and handover discipline", "description": "Solve \"nobody in charge for an hour\" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.\n\n- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend. Retains command when a VP joins.\n- **Communications Lead:** the single voice to the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.\n- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1 and SEV2.\n- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.\n- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.\n- Security, Legal, Compliance, Finance and Vendor Management join on defined triggers rather than by invitation.\n- Rules: command claimed within 5 minutes and stated in channel; distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.\n- **Concurrency doctrine:** two simultaneous SEV1/SEV2 incidents activate the secondary commander and a surge roster; a designated Multi-Incident Coordinator arbitrates shared resources such as the ledger, the database platform and the deploy freeze.\n- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.", "dependencies": ["S7"]}, {"step_id": "S9", "title": "Three-layer 24x7 coverage: command corps, domain rotations, overnight triage desk", "description": "Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.\n\n- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving roughly one primary week per person every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.\n- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement and reconciliation, auth, API edge, Kubernetes platform, data and reporting, partner integrations). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.\n- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.\n- **Duty Triage Desk:** a small paid overnight first-line rotation owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to \"I will not carry a pager for other teams' code.\"\n- Publish the staffing arithmetic: 10–12 domain rotations of six to eight people plus a 30-person command corps means roughly 100 of 260 engineers carry a night obligation, at about one week in six. Twenty-eight independent rotations would be unstaffable and is therefore rejected on the numbers.\n- Teams too small for a fair rotation get headcount, service reassignment, or a time-limited executive exception. **Never a two-person 24x7 rotation.**\n- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.\n- Acknowledgement discipline: a human acknowledgement within 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.\n- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.", "dependencies": ["S5", "S6", "S8"]}, {"step_id": "S10", "title": "Paid on-call, New York labour compliance and fatigue safeguards", "description": "Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.\n\n- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour, and a premium for the overnight triage desk.\n- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.\n- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.\n- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.\n- A rotation averaging more than two out-of-hours pages per responder per week for four weeks triggers a mandatory staffing or alert-remediation review.\n- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty since day 7 is paid retroactively.\n- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons without career penalty, and on-call load is published quarterly by team.", "dependencies": ["S5", "S9"]}, {"step_id": "S11", "title": "Alert quality contract and page budget", "description": "3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.\n\n- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.\n- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.\n- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.\n- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.\n- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.\n- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression can never hide an incident.", "dependencies": ["S3", "S6"]}, {"step_id": "S12", "title": "Detection uplift on the money path, validated by incident replay", "description": "The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.\n\n- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.\n- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.\n- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.\n- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.\n- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish \"replay coverage\" as a leading metric and close the gaps it exposes.\n- Record detection source on every incident. \"Customer detected first\" becomes a named defect class with a mandatory tracked action and a missed-detection review.", "dependencies": ["S6", "S11"]}, {"step_id": "S13", "title": "Change intelligence and deployment safety on the critical path", "description": "Most of these incidents start with something we changed. Make change the first hypothesis the tooling answers, and make changes safer to reverse.\n\n- Stream every deployment, configuration change, feature-flag flip, schema migration and infrastructure change into the incident timeline with service labels and owners.\n- Give the commander an automatic \"what changed in the last 60 minutes on the affected journey\" panel at declaration time.\n- Require Tier 0/1 changes to be progressively delivered with a documented rollback that is tested, time-bounded and executable by the on-call responder without the author.\n- Treat ledger schema migrations and settlement-affecting changes as a separate class: dual approval, rehearsed rollback, no deploys inside the settlement window.\n- Enforce a change freeze during SEV1 and SEV2, lifted only by the commander and logged.\n- Report change-correlated incidents monthly; a rising ratio is a signal to strengthen release safety, not to blame a team.", "dependencies": ["S6", "S12"]}, {"step_id": "S14", "title": "One pager, one incident record, one status page — migrated without a detection gap", "description": "Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail.\n\n- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.\n- Ingest all six sources first, deduplicate and correlate, then retire a legacy paging path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.\n- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.\n- Route every page through the service catalog: service label → owning domain schedule → escalation policy.\n- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access, MFA and periodic access review.\n- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.\n- Set a hard date after which a page delivered outside this platform creates no on-call obligation.", "dependencies": ["S7", "S9", "S11"]}, {"step_id": "S15", "title": "Escalation ladder and the five-minute command rule", "description": "Write one unskippable path from \"something looks wrong\" to \"someone is in charge\", and make the default action never be waiting.\n\n- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.\n- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.\n- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.\n- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.\n- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.\n- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.\n- Maintain and test quarterly the external escalation contacts and support entitlements: AWS and database premium support, processors, sponsor banks, card networks.", "dependencies": ["S8", "S9", "S14"]}, {"step_id": "S16", "title": "Live execution doctrine and payments safety rules", "description": "Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.\n\n- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.\n- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.\n- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.\n- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.\n- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed in advance rather than improvised.\n- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and automatic reopening if impact recurs.", "dependencies": ["S8", "S14", "S15"]}, {"step_id": "S17", "title": "Service readiness bar, major-incident playbooks and ledger blast-radius reduction", "description": "Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human, and the shared ledger is the largest structural risk in the estate.\n\n- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.\n- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.\n- Treat the **shared ledger cluster as a concentration risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.\n- Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers, stronger regional independence — with board-visible quarterly milestones.\n- Enforcement: no readiness sign-off, no night paging alerts unless the manager accepts the gap in writing with an expiry date and a compensating control. Never respond to a gap by turning detection off.\n- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.", "dependencies": ["S6", "S9", "S16"]}, {"step_id": "S18", "title": "Internal communications protocol", "description": "Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.\n\n- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, Sales and Security.\n- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change. First internal brief within 10 minutes of SEV1 and 15 of SEV2.\n- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.\n- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.\n- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.\n- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.", "dependencies": ["S8", "S14"]}, {"step_id": "S19", "title": "Customer communications, status page and account-manager outreach", "description": "Replace \"whoever is around\" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.\n\n- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; monitoring notice after mitigation; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.\n- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, monitoring, resolution.\n- Component-level status page mapped to customer journeys rather than internal service names, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.\n- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1 and 60 of SEV2, using one approved briefing pack. The long tail gets the status page plus proactive email.\n- Language rules: state observed impact, affected capabilities, any workaround and the next update time; **never speculate on cause, recovery time, data integrity or blame**.\n- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.\n- Record any legally required restriction or delay of public detail, its approver, and the alternative stakeholder plan.", "dependencies": ["S7", "S18"]}, {"step_id": "S20", "title": "Support and account management as a detection and intake tier", "description": "Customers detected 40% of incidents first, which means the front line already holds the signal. Turn Support and account managers into an instrumented detection channel rather than a bystander.\n\n- Give Support explicit declaration rights, a one-page trigger card, and a macro that opens an incident candidate directly in the incident platform.\n- Automate the clustering rule: five similar tickets or calls in ten minutes auto-creates a triage incident assigned to the duty commander.\n- Route credible partner, processor and sponsor-bank notifications into the same declaration path within five minutes.\n- Generate the affected-customer list automatically from journey telemetry and the incident record, and push it to Support, CS and the account-manager briefing.\n- Train Support and account managers on approved language and prohibit independent technical explanations to customers.\n- Measure and publish \"signal was in Support before it was in monitoring\" as a detection defect, and feed each instance into the detection backlog.", "dependencies": ["S7", "S12", "S19"]}, {"step_id": "S21", "title": "Regulatory, partner and legal notification playbook", "description": "In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.\n\n- Complete the obligation matrix started in scoping and have counsel validate the triggers, deadlines, channels and submitting authority for each obligation.\n- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is \"not reportable\"**, with facts considered, decision-maker, timestamp and reassessment trigger.\n- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, card networks, outside counsel, cyber-insurer, critical vendors.\n- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the operational facts.\n- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.\n- Rehearse the playbook quarterly inside the exercise programme, including one full regulator-notification simulation per year.", "dependencies": ["S4", "S7", "S19"]}, {"step_id": "S22", "title": "SLA credit workflow, true cost model and incentive guardrails", "description": "Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.\n\n- Agree with Legal and Finance the availability measurement method per contract and per component, including how partial degradation counts.\n- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.\n- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.\n- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.\n- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.\n- Guardrail against the perverse incentive: better detection will surface incidents that previously went unbilled, so credits may rise before they fall. Publish this expectation to the executive team in advance, and make it a written rule that no finance or commercial pressure may influence severity or declaration.", "dependencies": ["S7", "S19"]}, {"step_id": "S23", "title": "Blameless postmortem standard and Incident Review Board", "description": "Replace \"some incidents, various formats\" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.\n\n- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger; any control failure.\n- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning engineering director is accountable.\n- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.\n- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**\n- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.\n- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.\n- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly \"top five recurring causes\" analysis, with restricted versions where security, privacy or privileged content requires it.", "dependencies": ["S4", "S7", "S8"]}, {"step_id": "S24", "title": "Action ownership, reserved capacity and enforcement", "description": "Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.\n\n- Every action gets one named individual owner, accountable manager, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.\n- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.\n- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.\n- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.\n- Verify effectiveness through tests, telemetry, exercises or production evidence before closing. A closed ticket without evidence does not close the action.\n- Report closure rate by team monthly and include it in engineering manager objectives.", "dependencies": ["S14", "S23"]}, {"step_id": "S25", "title": "Training and certification academy", "description": "Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.\n\n- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.\n- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.\n- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.\n- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, the 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.\n- **Comms Lead (1 day):** status writing, customer tiering, contractual clocks, legal boundaries, regulator triggers, what never to say.\n- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days of joining a domain.\n- Certification is valid 12 months and renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.", "dependencies": ["S8", "S15", "S18", "S23"]}, {"step_id": "S26", "title": "Exercise programme: tabletops, game days, night drills and vendor rehearsals", "description": "The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.\n\n- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander. First company-wide tabletop within 30 days.\n- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.\n- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times, plus a concurrency drill with two simultaneous incidents.\n- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.\n- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down. Rehearse joint escalation with AWS, a processor and a sponsor bank.\n- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.\n- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.", "dependencies": ["S14", "S17", "S25"]}, {"step_id": "S27", "title": "Publish Incident Management Policy v1 and the exception register", "description": "Collapse the design into a document people will actually open mid-outage, and make it official.\n\n- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, escalation and communications timings. Linked directly from the incident tool.\n- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.\n- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix, Readiness Bar.\n- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.\n- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.", "dependencies": ["S7", "S8", "S9", "S10", "S11", "S15", "S18", "S19", "S21", "S23", "S24"]}, {"step_id": "S28", "title": "Pilot on the payments critical path", "description": "Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.\n\n- Scope: ledger, payments orchestration, settlement and reconciliation, PostgreSQL platform, Kubernetes platform, API edge, auth and Support intake — including two of the 12 teams already on-call.\n- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, action tracking, and paid rotations processed through payroll.\n- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.\n- The programme lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.\n- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.\n- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time in 95% of cases, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.", "dependencies": ["S10", "S12", "S13", "S14", "S17", "S25", "S27"]}, {"step_id": "S29", "title": "Alert noise burn-down campaign", "description": "Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.\n\n- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.\n- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.\n- Publish a weekly noise leaderboard that names systems, never people.\n- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.\n- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.\n- Milestones: 3,400 → under 1,500 pages a month by day 90, under 500 with noise below 15% by month 6.", "dependencies": ["S11", "S14", "S28"]}, {"step_id": "S30", "title": "Metrics, dashboards, review cadence and anti-gaming", "description": "Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.\n\n- **Response:** time from first impact to detect, declare, commander, acknowledge, mitigate, resolve — split by severity, tier, journey, region and detection source.\n- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.\n- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.\n- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors, change-correlated incident ratio.\n- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.\n- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.\n- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.\n- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.", "dependencies": ["S14", "S23", "S28"]}, {"step_id": "S31", "title": "Wave rollout to all 28 teams with readiness gates", "description": "Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.\n\n- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.\n- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.\n- Each wave gets a named coach from the programme office for three weeks; the director signs the gate.\n- **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.\n- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.\n- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard by team, tier and control gap.", "dependencies": ["S28", "S30"]}, {"step_id": "S32", "title": "Change management, fairness and pager culture", "description": "Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.\n\n- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.\n- Publicly close the two historical \"nobody in charge\" incidents with a written account of what would be different now.\n- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.\n- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.\n- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.\n- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.", "dependencies": ["S5", "S10", "S28"]}, {"step_id": "S33", "title": "SOC 2 evidence by design, internal control test and mock audit", "description": "The real process is the evidence. Never build a parallel audit process, and never reconstruct records after the fact.\n\n- Maintain the evidence set defined in scoping, produced automatically and indexed: versioned policies and exceptions, catalog records, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including \"not reportable\", postmortems, action closure with verification, training register, drill records, access reviews.\n- Sample incidents monthly from first signal through verified corrective action, deliberately including customer-reported events, downgraded incidents and missed timelines.\n- Internal Audit or an independent control owner tests design and operation in month 5, sampling end to end and reporting gaps to the sponsor.\n- Run a formal mock audit in month 7 using the populations, evidence requests and interviews the auditor will use: a commander, a random engineer, Support, Compliance.\n- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.\n- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.", "dependencies": ["S4", "S27", "S30", "S31"]}, {"step_id": "S34", "title": "Programme risk register and pre-committed contingencies", "description": "Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.\n\n- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.\n- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.\n- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.\n- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.\n- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps and immediate staffing intervention.\n- **A major SEV1 mid-rollout:** the programme lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.\n- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.", "dependencies": ["S1"]}, {"step_id": "S35", "title": "Ninety-day inspect-and-adapt, then year-two sustainability", "description": "Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.\n\n- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.\n- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.\n- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, exercise calendar and metrics, independent of the audit cycle.\n- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.\n- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.\n- Keep the annual policy review, certification renewal, exercise calendar and quarterly board reporting as permanent commitments owned by the Head of Reliability.", "dependencies": ["S30", "S31", "S33"]}], "estimated_complexity": "high", "success_metrics": "- A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.\n- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.\n- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.\n- Replay coverage: by day 90 a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.\n- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.\n- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.\n- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95% internally and externally.\n- Top-100 account outreach completed within 30 minutes of SEV1 in 95% of qualifying cases, using the approved briefing pack.\n- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.\n- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.\n- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.\n- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.\n- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.\n- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.\n- Paid on-call live in payroll before any mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.\n- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.\n- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.\n- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.\n- Change-correlated incidents are identified within 10 minutes of declaration in 90% of cases, and the change-correlated share of incidents declines quarter over quarter.\n- SLA credits fall to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.\n- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.\n- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of \"not reportable\".\n- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, one concurrency drill and two full cross-company exercises completed before the audit, each producing tracked actions.\n- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.\n- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.\n- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and a tested read-only mode by month 6, with milestones reported to the board quarterly."}It decomposed its round-2 mega-steps into discrete, individually gated steps and added the sharpest measurement definitions in the field. It remains the only plan without a risk register.
- Round-2 compressions were split: fairness contract (step 4), live lifecycle (step 10), escalation ladder (step 14), credits (step 17), postmortems (step 18), corrective actions (step 19), policy (step 23), scorecard (step 24), rollout (step 26), inspect (step 28).
- Step 12 defines an actionable page and defines noise (duplicate, non-urgent, unactionable, stale, test-generated, misrouted) and counts "human notification episodes" rather than raw alert events — the only plan whose noise target is unambiguously measurable.
- Step 6 replaces a single flag with five modifiers (FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, VENDOR) that invoke specialist controls without distorting customer-impact severity.
- Step 24 (scorecard) is now instrumented before the pilot (step 25 depends on 24), so pilot failures are visible immediately.
- Step 12 adds "never disable an existing critical detector solely because its metadata or runbook is incomplete"; step 19 prefers hazard-removing actions over "add monitoring" or retraining.
- Step 16 auto-enrols customers on the status page only where contracts and consent permit, and metrics target "no unresolved material exception" rather than the others' absolute "zero exceptions".
- Still no standing risk register or pre-committed contingencies; the only failure handling is per-gate exceptions and step 28.
- Step 16 now carries internal-to-customer comms, the obligation matrix and reportability in 12 bullets — a new compression that undoes some of the decomposition gains.
- Dropped the round-2 aggregate on-call budget estimate ($0.9M–$1.2M/yr), leaving only per-week bands against the $1.3M credit comparison.
- Design compressed to weeks 1–4 for a 180-service catalog, 28-team listening tour and tool selection — the most optimistic schedule of the five.
- Proposal 1 : Replay the 31 historical incidents to name the detector that would now fire and at which minute.
- Proposal 4 : A flag attached to any severity for integrity or security events.
- Proposal 4 : A standalone, published on-call fairness contract before any mandatory night rotation.
- Proposal 4 : Publish a signed Policy v1 plus an exception register as a discrete step.
- Proposal 3 : A blameless postmortem standard with fixed deadlines and a weekly Incident Review Board as its own step.
- Proposal 1 : A paid overnight Duty Triage Desk that owns the first ten minutes of any page.
- Proposal 1 : Subscribe all 2,100 customers to the status page by default.
- Proposal 1 : Claim SOC 2 incident-response controls will pass with zero exceptions.
+ Charter the program and fund immediate action+ Publish the on-call fairness contract+ Codify the live incident lifecycle+ Implement the detection and escalation ladder+ Tie incidents to SLA credits and financial exposure+ Establish mandatory blameless postmortems+ Publish the signed policy and control set+ Instrument the scorecard and review forums+ Roll out by risk with readiness gates+ Inspect, adapt, and institutionalize ownershipEstablish the mandate, owner, funding, and scheduleCo-design the fairness contract with engineersCodify detection, escalation, and live executionOperationalize legal, regulatory, contractual, and credit decisionsMake postmortems mandatory, consistent, and blamelessRoll out by customer journey and riskMeasure performance and operate fixed review forumsInstitutionalize improvement and reduce structural risk
The plan produced
1. Charter the program and fund immediate action from P4 step 1
Make incident management a company operating process within 48 hours. Give one accountable leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Include Engineering, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit in a decision group. The sponsor resolves blocked decisions within 48 hours.
- Approve non-negotiables: one severity model, one human-paging path, one incident record, paid on-call, named service ownership, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, resilience, and incident actions.
- Fund tooling, compensation, training, exercises, and program staff. Compare the cost with the existing $1.3M in annual SLA credits.
- Set the schedule: operating floor by day 7, design during weeks 1–4, pilot during weeks 5–10, rollout during weeks 11–18, control tests in months 4 and 6, and mock audit in month 7.
2. Install a seven-day operating floor (after 1) from P4 step 2
Do not wait for the final policy or new tooling. Put a minimum viable process into operation immediately and retain evidence from the first day.
- Publish a one-page interim severity guide, declaration procedure, role card, and communications clock.
- Provide one monitored declaration route through chat, telephone, and the current paging environment.
- Create one channel, bridge, timeline, and incident identifier for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced managers and engineers.
- Require a named commander within 10 minutes. The duty engineering director assumes command if the command page is unclaimed.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Stop uncompensated mandatory after-hours expansion. Pay interim duty under a temporary stipend, retroactive to program launch.
- Hold a 15-minute daily operational review until the permanent process is active.
3. Create the factual, contractual, and control baseline (after 1)
Build one defensible baseline for process design, investment decisions, and SOC 2 testing. Preserve the original data so progress cannot be created by changing definitions.
- Reconstruct all 31 incidents from impact start through detection, declaration, command, mitigation, recovery, communications, credits, and corrective actions.
- Reconstruct the two incidents with unclear command minute by minute.
- Identify the missing signal for every customer-first detection.
- Inventory all six alert sources, 3,400 monthly alert events, duplicates, noisy rules, missing owners, and missing runbooks.
- Record current rotations, unpaid duty, overnight activations, schedule size, and uncovered services.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, 85% noise, $1.3M in credits, and 11 of 64 actions closed.
- Inventory customer-specific availability definitions, notice periods, credit terms, sponsor-bank obligations, and other incident-related contracts.
- Confirm the required SOC 2 Type II observation period and evidence expectations with the auditor during week 1.
4. Publish the on-call fairness contract (after 1) from P4 step 21
Treat pager resistance as a legitimate design constraint. Adoption depends on a written agreement that separates command from technical ownership.
- Interview representatives from all 28 teams and from Support, Customer Success, Security, and Operations.
- Separate concerns about unpaid work, sleep loss, unfamiliar systems, noisy alerts, inadequate runbooks, and blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for systems they own or have formally accepted and trained to support, and assisted by a separate Incident Commander.
- State that Platform may temporarily triage unknown ownership but does not inherit another team's service.
- Provide confidential accommodations for health, disability, pregnancy, or caregiving constraints without career penalty.
- Measure baseline trust, fairness, fatigue, and alert confidence. Repeat at days 60 and 120, then quarterly.
5. Build the service catalog and customer-journey map (after 3)
Make a machine-readable catalog the source of truth for routing, impact analysis, status-page components, and control evidence. Every production service must have one accountable owner.
- Record the owning team, manager, business capability, repository, channel, dashboard, runbook, escalation policy, dependencies, regions, and data stores for all 180 services.
- Map initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, reporting, and onboarding to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing systems that can create material contractual impact.
- Classify Tier 2 as deferrable internal or batch systems and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional recovery mode, contractual commitments, and critical vendors for Tier 0 and Tier 1.
- Keep separate but coordinated ownership for the ledger application and PostgreSQL platform.
- Give orphan services an owner or approved decommission date within 30 days. Treat unowned Tier 0 services as release-blocking executive risks.
6. Adopt the severity model and incident modifiers (after 3, 5)
Classify incidents by credible customer, financial, security, regulatory, and contractual harm. Start at the higher plausible severity while scope or integrity remains unknown.
- SEV1 — crisis: incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; broad loss of a core payment journey; both regions impaired; or processing must be stopped. Immediately page all response roles and the Executive Duty Officer. Open the bridge within 5 minutes, freeze unrelated changes, issue internal notice within 10 minutes, publish applicable customer status within 15 minutes, start legal assessment within 1 hour, and require a postmortem.
- SEV2 — major: material payment degradation; settlement deadline at risk; regional impairment with reduced resilience; a critical customer or material cohort unavailable; or an SLA breach is likely. Page command and technical roles immediately. Issue internal notice within 15 minutes, applicable customer status within 30 minutes, and require a postmortem.
- SEV3 — limited: narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads. Page only when immediate action can reduce harm.
- SEV4 — operational event: no current customer impact and no credible imminent harm. Create a ticket and handle during normal hours.
- Add FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, and VENDOR modifiers. These invoke specialist controls without distorting customer-impact severity.
- Use at least SEV2 posture for unknown impact lasting 15 minutes, credible ledger-integrity risk, or a cross-domain incident without clear ownership.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Permit anyone to declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Publish a decision tree and examples based on the 31 historical incidents.
7. Define roles, authority, and handoffs (after 6)
Separate coordination, communications, recordkeeping, and technical repair. One named person holds command continuously throughout every SEV1 and SEV2.
- Incident Commander: owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not act as the primary technical operator.
- Communications Lead: owns internal broadcasts, status-page updates, account-manager briefs, executive updates, and coordination with Legal.
- Scribe: maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- Subject-Matter Responders: diagnose and mitigate only systems for which they have ownership, access, training, or a formally accepted support agreement.
- Executive Duty Officer: removes organizational obstacles and makes exceptional business decisions without displacing the commander.
- Add Security, Legal, Compliance, Finance, and Vendor Management on modifier-specific triggers.
- Require distinct commander, Communications Lead, scribe, and technical lead for SEV1. Communications and scribe may combine for the first 10 minutes of a bounded SEV2.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, reconciliation, and change evidence for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, including the exact transfer time and unresolved risks.
8. Create sustainable 24x7 coverage (after 4, 5, 7)
Use central command coverage and risk-based technical coverage instead of creating 28 fragile night rotations. Responders do not carry pagers for unfamiliar code.
- Create a 24x7 Incident Command corps of approximately 30 certified people with primary and backup schedules. Two schedules create about 104 weekly assignments a year, or roughly three to four weeks per person annually.
- Create a 24x7 Communications pool of 18–24 trained people from Support, Customer Operations, Engineering Operations, and management.
- Create a similarly sized scribe pool. The backup commander temporarily records the first minutes if a scribe has not joined.
- Maintain a 24x7 Executive Duty Officer schedule and specialist contact paths for Security and Legal.
- Group Tier 0 and Tier 1 systems into roughly 8–12 coherent response domains only where responders share training, access, runbooks, and explicit support acceptance.
- Staff every critical domain with primary and secondary responders and at least six qualified people. Target eight where overnight activation is frequent.
- Give Tier 2 and Tier 3 services business-hours ownership plus tested manager and director escalation.
- Reclassify any lower-tier service capable of causing severe overnight harm rather than hiding the risk behind manager callback.
- Route unknown-owner incidents to Duty Command and Platform temporarily. Record each occurrence as a catalog control defect.
- Prohibit simultaneous primary assignments and two-person 24x7 rotations.
9. Implement compensation and fatigue protections (after 4, 8)
End unpaid on-call before expanding mandatory coverage. Use fixed duty compensation so responders are not rewarded for alert volume.
- Use planning bands of $900–$1,200 per Tier 0/1 primary week and $300–$500 per secondary week.
- Use planning bands of $1,000–$1,300 per Duty Commander week and $400–$700 for Communications or scribe primary duty.
- Pay holiday premiums. Compensate all legally compensable active and waiting time for non-exempt staff, including overtime where required.
- Have HR, Finance, Payroll, and employment counsel approve final bands, tax handling, FLSA classification, New York wage-hour treatment, and schedule constraints within 14 days.
- Provide a protected recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal delivery commitments by about 15% during a primary week.
- Prohibit consecutive primary weeks, on-call during leave, hidden schedule swaps, and primary duty more often than one week in six.
- Allow responders to declare temporary fatigue-related unfitness without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Make active payroll setup, training, access, and readiness hard gates before any new mandatory night rotation starts.
10. Codify the live incident lifecycle (after 6, 7)
Give every incident the same operational flow from first signal through verified recovery. The first objective is limiting customer and financial harm, not proving root cause.
- Use the states Detected, Declared, Triaged, Mitigating, Mitigated, Monitoring, Resolved, and Reviewed.
- Record impact start separately from detection. Use the earliest defensible evidence and revise it transparently when later facts emerge.
- Open with a standard command message: severity, known impact, assigned roles, immediate objective, workstreams, and next update time.
- Freeze unrelated production changes during SEV1 and normally during SEV2. Record every exception.
- Prefer reversible mitigation: rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, partner rerouting, or controlled processing suspension.
- Separate mitigation and diagnosis workstreams when staffing allows.
- Keep decisions in the shared incident record rather than direct messages.
- Require formal command and technical handoffs for shift changes, fatigue, or incidents exceeding four hours.
- For payment incidents, verify backlog handling, duplicate protection, customer state, settlement exposure, and ledger reconciliation before resolution.
- Require a severity-specific stability period and explicit handback to the owning team, Support, and Customer Success.
11. Establish one paging and incident system of record (after 5, 7, 8)
Monitoring tools may remain specialized, but every human page and major-incident record must enter one controlled platform. This provides consistent routing and an audit trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page within two weeks.
- Ingest events from all six monitoring tools before disabling their direct human-paging routes.
- Route pages using the service catalog and deduplicate events belonging to the same symptom.
- Provide one declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, decisions, severity changes, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status publishing, and conference facilities.
- Apply MFA, role-based access, periodic access reviews, and tamper-evident history.
- Keep privileged legal or security material in restricted linked records rather than exposing it in the general timeline.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, the status provider, or an AWS region.
- Retire each legacy paging route only after ownership review, end-to-end testing, and two weeks of verified operation.
12. Enforce the alert-quality contract (after 3, 5)
Treat every human page as a production interface with an owner and a required action. Measure human notification episodes rather than raw monitoring events.
- Require every paging rule to identify the service, owner, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Define an actionable page as one that causes or materially informs a timely intervention or risk decision.
- Define noise as duplicate, non-urgent, unactionable, stale, test-generated, or incorrectly routed notification.
- Page on payment outcomes, error-budget burn, queue age against deadlines, and financial-integrity risk rather than raw CPU, memory, pod, or log thresholds.
- Run new rules in shadow mode for seven days unless a documented emergency exception applies.
- Review repeated no-action pages within two business days.
- Set a page budget of no more than two after-hours notification episodes per responder per week, measured over four weeks.
- Make a sustained budget breach trigger a tuning sprint and block additional non-emergency paging rules.
- Require compensating detection and central approval before suppressing a Tier 0 or Tier 1 rule.
- Never disable an existing critical detector solely because its metadata or runbook is incomplete. Track the gap with a dated remediation owner.
13. Detect payment failures before customers (after 5, 12)
Move detection from infrastructure health to customer journeys and ledger truth. Validate coverage against actual historical failures.
- Define SLIs and internal SLOs for initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, APIs, webhooks, and reporting freshness.
- Set internal objectives with enough headroom to protect the contractual 99.95% availability commitment.
- Run external synthetic transactions through critical journeys at least once per minute from paths independent of the production platform.
- Test each region and expose dependencies that defeat nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed payment value relative to settlement deadlines.
- Add tenant and cohort anomaly detection for high-value customers and critical payment methods.
- Convert credible Support, account-manager, processor, bank, and network reports into incident candidates within five minutes.
- Replay all 31 historical incidents. Record which current detector would fire and at what minute.
- Treat customer-first detection as a mandatory missed-detection review with a tracked action.
14. Implement the detection and escalation ladder (after 7, 8, 11) new
Create one time-bound path from the first credible signal to named command and the correct technical owner. Device delivery does not count as acknowledgement.
- Converge automated alerts, engineer observations, support cases, account-manager reports, partner notices, and customer calls on the same declaration path.
- Page the Duty Incident Commander and owning critical-domain primary immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at 5 minutes, manager at 10 minutes, and director at 15 minutes.
- Escalate unclaimed command to the backup commander at 5 minutes. The Executive Duty Officer assumes temporary command at 10 minutes until a certified transfer occurs.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Keep command with the current commander when service ownership remains unclear. Assign a temporary technical lead and record the ownership gap.
- Maintain tested external escalation routes for AWS, database support, processors, sponsor banks, networks, and critical vendors.
- Test the full declaration, acknowledgement, fallback, conference, and status-publishing path weekly.
- Treat missed acknowledgement, failed routing, and unowned incidents as control failures requiring review.
15. Standardize internal communications (after 7, 10, 11)
Give responders one working room and stakeholders one controlled source of truth. Executives must not interrupt the technical command path.
- Maintain one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue the initial internal notice within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within the same initial-notice window.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record every material decision and outbound message in the incident timeline.
- Require the executive team to sign the communication behavior rules.
16. Standardize customer, account, and regulatory communications (after 3, 6, 15)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate observed impact before root cause is known.
- Publish customer status within 15 minutes of a customer-visible SEV1 and within 30 minutes of a customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map public status components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, available workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Have account managers contact affected strategic accounts within 30 minutes for SEV1 and 60 minutes for SEV2 using the approved briefing.
- Offer status subscriptions to all customers. Auto-enroll only where contracts, consent, and applicable communication rules permit.
- Encode customer-specific notice deadlines and channels in the customer record.
- Have Legal and Compliance maintain a counsel-validated matrix covering applicable NYDFS, breach, GLBA or FTC, PCI, money-transmitter, sponsor-bank, network, insurance, and customer obligations.
- Complete and record a reportability assessment within one hour for every SEV1 and every security, privacy, or integrity-related SEV2, including decisions of not reportable.
- Let Legal own regulatory text and submission while the Incident Commander owns operational facts. Record any legally required restriction of public detail and its alternative stakeholder plan.
17. Tie incidents to SLA credits and financial exposure (after 3, 16) from P4 step 16
Link the incident record to contractual and financial outcomes. Finance should not discover outages through later credit claims.
- Define the authoritative availability calculation for each contract and customer journey with Legal and Finance.
- Calculate affected customers and minutes from incident scope and journey telemetry.
- Produce a preliminary credit and contractual-exposure estimate within five business days of resolution.
- Record failed payment count, failed or delayed value, settlement exposure, reconciliation breaks, support effort, and engineering effort.
- Establish a documented approval path for proactive credits and claims-based credits.
- Attribute credits and financial harm to recurring failure families.
- Use the quarterly credit analysis to prioritize detection, resilience, and architectural investment.
18. Establish mandatory blameless postmortems (after 6, 7) from P3 step 20
Use one learning standard with fixed deadlines. Keep learning separate from disciplinary and misconduct processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact lasting more than two hours, SLA credits, contractual breach, repeated contributing factors, major control failures, and ledger-integrity near misses.
- Produce a factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the owning engineering director accountable for completion. The commander owns response analysis, and the scribe supplies the timeline.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit answers to why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the commander or primary technical responder.
- Describe decisions using the context and information available at the time. Do not name an individual as the root cause.
- Keep HR, misconduct, and personnel matters in separate processes.
- Publish broadly useful findings internally while restricting security, privacy, personnel, and privileged content appropriately.
19. Make corrective actions enforceable risk commitments (after 18)
An action is not complete when its ticket is closed. It is complete when the intended risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within 7 days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Prefer actions that remove hazards or reduce blast radius over vague actions such as retraining or adding monitoring.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of engineering capacity for approved reliability work.
- Escalate overdue high-risk actions to the manager after 7 days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit unrepaired severe conditions to block related releases.
- Verify effectiveness using tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days as complete, re-planned, superseded with evidence, or formally risk-accepted.
20. Set the service-readiness bar and payments playbooks (after 5, 10, 12, 13)
A critical service must be supportable at 3 a.m. before it enters direct overnight coverage. Existing critical detection remains active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and integrity constraints.
- Require responders to demonstrate access and safe execution before independent primary duty.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional impairment, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document split-brain prevention, replay protection, failover, controlled backlog recovery, and post-recovery reconciliation.
- Require executive-approved RTO and RPO for the shared ledger cluster.
- Exercise critical runbooks at least twice a year and after material changes.
- Block new Tier 0 or Tier 1 releases and paging rules when readiness requirements are missing.
- Handle existing gaps using named owners, compensating controls, executive-approved expiry dates, and remediation plans.
- Run a parallel architecture workstream to reduce ledger blast radius, isolate non-critical readers, and strengthen regional independence.
21. Train and certify every response role (after 7, 14, 15, 18, 20)
Command and communication are learned skills. Use paid working time and certify people before independent duty.
- Give all employees a 30-minute module on recognizing impact, declaring incidents, and locating the status page.
- Give responders a half-day module on severity, acknowledgement, escalation, evidence, runbooks, and financial-integrity precautions.
- Train scribes for two hours on timeline quality and fact-versus-hypothesis labeling.
- Train Incident Commanders for two days on delegation, uncertainty, severity, mitigation strategy, fatigue, handoffs, and executive management.
- Require commander candidates to complete a simulated SEV1 and two shadowed incidents or exercises.
- Train Communications Leads for one day on status writing, customer segmentation, legal boundaries, and contractual clocks.
- Require domain responders to demonstrate dashboards, access, rollback, failover, escalation, and relevant playbooks.
- Require two shadow shifts before independent primary duty.
- Renew certification annually through simulation.
- Maintain the training, assessment, and certification register as operational and audit evidence.
- Nominate an incident-management champion in each of the 28 teams.
22. Exercise command, recovery, and tool failure (after 11, 20, 21)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first cross-company command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous operational and security events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors when relevant.
- Conduct at least one unannounced after-hours paging test before the audit and two annually thereafter.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure acknowledgement, command, customer notice, mitigation decision, handoff, and recovery times.
- Create tracked corrective actions for every material exercise finding.
23. Publish the signed policy and control set (after 6, 7, 8, 9, 10, 12, 14, 15, 16, 18, 19) from P4 step 21
Convert the design into concise documents that people can use during an incident. The actual operating process must also be the documented and audited process.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, evidence, and exception standards.
- Include one-page cards for severity, roles, authority, escalation, and communication timings inside the incident tool.
- State explicitly that responders support only owned or formally accepted and trained service portfolios.
- Include the compensation structure, fatigue rules, declaration rights, and non-retaliation commitment.
- Obtain approval from the CTO, HR, Legal, Security, Compliance, and Internal Audit.
- Announce the policy at an all-hands and through team briefings.
- Create an exception register with owner, rationale, compensating control, approver, review date, and expiry.
- Version every policy change. Do not rewrite historical records when the process changes.
24. Instrument the scorecard and review forums (after 3, 11, 18, 19) new
Measure the process before the pilot so failures become visible immediately. Report medians and 90th percentiles rather than averages alone.
- Measure impact-to-detection, detection-to-declaration, declaration-to-command, acknowledgement, mitigation, recovery, and resolution.
- Split results by severity, service tier, customer journey, region, detection source, and business-hours status.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track notification episodes, actionability, duplicates, after-hours load, missed detection, routing errors, and page-budget breaches.
- Track postmortem timeliness, action age, due-date performance, verified effectiveness, and repeated contributing factors.
- Track journey availability, error-budget burn, failed or delayed value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight work, recovery days, exceptions, sentiment, and responder attrition.
- Hold a weekly Incident Review Board chaired by the Head of Reliability with relevant directors.
- Hold a monthly executive reliability review and a quarterly control and resilience review with Internal Audit.
- Reconcile incident records monthly against support cases, customer complaints, status history, credits, and major operational anomalies to detect under-reporting.
- Use team-level scorecards to direct help and investment. Never penalize an individual for good-faith declaration.
25. Pilot the complete process on the payment path (after 9, 11, 13, 20, 21, 23, 24)
Run a four-to-six-week pilot across the highest-risk journey before expanding. The interim operating floor remains active for the rest of the company.
- Include payment orchestration, ledger application, PostgreSQL platform, API edge, authentication, settlement, reconciliation, Kubernetes platform, and Support intake.
- Include teams with existing on-call experience and teams new to the model.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking together.
- Run legacy and new paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently taking command.
- Correct critical process or tooling defects within 48 hours.
- Exit only after 95% timely command assignment, 95% communications compliance, no unpaid pages, complete required postmortems, tested fallbacks, and at least 50% lower pilot noise.
- Publish the pilot results, defects, and policy changes company-wide.
26. Roll out by risk with readiness gates (after 25) from P4 step 24
Expand in controlled waves and finish early enough to accumulate operating evidence before the audit. A calendar date does not override a failed readiness gate.
- Roll out remaining Tier 0 domains first, followed by Tier 1, Tier 2, and Tier 3.
- Use four waves of six to eight teams, each lasting two to three weeks.
- Gate each service on catalog ownership, appropriate coverage, active compensation, trained responders, access, tested escalation, alert quality, runbooks, and a passed tabletop.
- Require at least six responders only for direct 24x7 technical rotations. Use business-hours coverage for lower tiers.
- Give each wave a named coach and director sign-off.
- Reschedule failed gates or use a time-limited executive exception with compensating controls. Do not create silent waivers.
- Disable legacy human-paging routes after verified cutover for each wave.
- Run a quota-based noise reduction sprint in every wave, starting with the highest-volume rules.
- Pair every suppression with a compensating-detection check.
- Publish an internal adoption dashboard by team, service tier, coverage, and control gap.
- Complete critical coverage by approximately week 12 and all 28 teams by week 18.
27. Prove SOC 2 operating effectiveness (after 22, 23, 24, 26)
Generate evidence through normal operation rather than reconstructing it before fieldwork. Test both control design and consistent execution.
- Map controls to the applicable Trust Services Criteria with Compliance and the auditor, including monitoring, incident identification, response, recovery, communications, and availability.
- Retain approved policies, exceptions, service ownership, schedules, compensation activation, access reviews, training, incidents, communications, reportability decisions, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, missed timelines, non-reportable decisions, and exercises in the testing population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the evidence populations and interviews expected from the external auditor.
- Correct deviations through tracked actions with owners and dates. Never edit history to create apparent compliance.
- Verify that evidence retention covers the full auditor-defined observation period.
- Brief commanders, engineers, Support, and Compliance on the actual process without scripting inaccurate answers.
28. Inspect, adapt, and institutionalize ownership (after 26, 27) from P4 step 26
Prevent the process from decaying after rollout or the audit. Change it using measured operating evidence rather than opinion.
- Review the policy after 90 days of live operation using severity calibration, page load, missed detection, communication compliance, action closure, fatigue, and survey results.
- Remove steps that create work without reducing risk. Add controls only where incidents, exercises, or evidence show a gap.
- Reassess Tier 0 and Tier 1 classification and domain boundaries every six months.
- Assign permanent owners for policy, catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review compensation bands, rotation burden, accommodations, and staffing annually.
- Report severe incidents, credits, overdue high-risk actions, and ledger concentration risk to the board or risk committee quarterly.
- Maintain the ledger blast-radius program as an executive risk until failover, degraded mode, reconciliation, and regional independence meet approved objectives.
- Evaluate follow-the-sun coverage using one year of actual activation and staffing data.
- Build year-two plans for automated mitigation, safer deployments, graceful degradation, and error-budget release controls.
- By day 7, every suspected major incident uses one record, one coordination channel, and a named Incident Commander within 10 minutes.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes.
- By month 3, at least 95% of SEV1 and SEV2 incidents have a named commander within 5 minutes.
- By month 3, at least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and appropriate technical coverage.
- By week 18, all 180 services have an owner, tier, escalation path, and tested coverage model.
- No new mandatory night rotation begins before compensation, training, access, runbooks, and minimum staffing are active.
- Every direct 24x7 technical rotation has at least six qualified responders or an approved, expiring executive exception.
- No responder is routinely primary more often than one week in six or assigned to two simultaneous primary rotations.
- Median impact-to-detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- By day 90, incident replay identifies a current detector for at least 90% of the 31 historical incidents, including its expected detection minute.
- Median time to mitigation falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of applicable SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet the required internal and customer update cadence by month 3.
- Monthly human notification episodes fall from the current 3,400 alert events to no more than 1,500 by day 90 and 500 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without reducing Tier 0 or Tier 1 replay coverage.
- Average after-hours load remains at or below two notification episodes per responder per week; every sustained breach receives a dated remediation plan.
- All six monitoring sources route human pages through the controlled paging platform by week 12, with direct legacy routes retired by week 18.
- All required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 from month 3 onward.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due dates with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Reportability is assessed and recorded within 1 hour for 100% of SEV1 and qualifying SEV2 incidents, including not-reportable decisions.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Approved ledger RTO and RPO, controlled failover, read-only mode, and post-recovery reconciliation are exercised before month 6.
- Quarterly responder surveys reach at least 75% favorable responses for fairness, ownership boundaries, compensation, and sustainability by month 6.
- Monthly contracted availability meets or exceeds 99.95% by month 6 using the contractually authoritative measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- The month-7 mock audit finds no unowned high-risk control gap and at least 95% of sampled incidents contain complete operating evidence.
- The SOC 2 Type II incident-response controls complete external testing without an unresolved material exception.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: d55e2704-ef8a-4e9c-8fd5-e6d956728a8d, Agent: opus5_refine_1, LLM: anthropic/claude-opus-5):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (32):
1. Charter the program: one owner, one mandate, funded, dated before the audit
Turn the CEO email into a chartered company program with a single accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the **burning-platform narrative** and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the written on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (depends on: 5, 7)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 8, 4)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (depends on: 10, 5)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (depends on: 7, 12, 13)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
18. SLA credit and financial impact workflow (depends on: 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Use the credit delta as the standing business case for on-call pay and reserved reliability capacity.
19. Blameless postmortem standard and Incident Review Board (depends on: 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, reserved capacity and enforcement (depends on: 19, 12)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
21. Runbooks, readiness bar and ledger blast-radius reduction (depends on: 5, 8)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (depends on: 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (depends on: 22, 12, 21)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
24. Publish Incident Management Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 11, 12, 21, 22, 9)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Alert noise burn-down campaign (depends on: 10, 12, 25)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
28. Wave rollout to all 28 teams with readiness gates (depends on: 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
29. Change management, fairness and pager culture (depends on: 4, 9, 25)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (depends on: 24, 26, 28)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 26, 28, 30)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 2 (ID: ddb335b2-98d9-46f8-afa0-88fca2e25749, Agent: gpt5.6-sol_refine_2, LLM: openai/gpt-5.6-sol):
Estimated Complexity: high
Success Metrics: - By day 7, every suspected major incident uses one incident record, one coordination channel, and a named commander.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes; at least 95% receive command within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and technical coverage.
- By week 18, all 180 services have an owner, tier, tested escalation path, and coverage appropriate to their risk.
- No mandatory night rotation starts before compensation, training, access, runbooks, and staffing controls are active.
- Every direct 24x7 technical rotation has at least six qualified responders or a documented, expiring executive exception.
- No responder is routinely assigned primary duty more often than one week in six or simultaneously assigned to two primary rotations.
- At least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of qualifying SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet their required internal and customer update cadence by month 3.
- Monthly human pages fall from 3,400 alert events to no more than 1,500 by day 90 and no more than 500 actionable pages by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, with no decline in Tier 0/1 detection coverage.
- Average after-hours load remains at or below two pages per responder per week; every sustained breach produces a remediation plan.
- All six monitoring sources route human pages through the single paging platform by week 12; direct legacy paging paths are disabled by week 18.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Monthly contracted availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- Quarterly on-call surveys reach at least 75% favorable responses on fairness, ownership boundaries, compensation, and sustainability by month 6.
- The month-7 mock audit finds no unowned high-risk control gap, and at least 95% of sampled incidents contain complete operating evidence.
Steps (26):
1. Establish the mandate, owner, funding, and schedule
Launch incident management as a **company operating program within 48 hours**. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Form a decision group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Approve the non-negotiables: one severity scale, one human-paging platform, one incident record, paid on-call, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, and incident actions.
- Fund tooling, compensation, training, exercises, and reliability work. Use the $1.3M in credits as the minimum financial comparison.
- Set milestones: interim process by day 7, policy and critical-path pilot by week 6, Tier 0/1 coverage by week 10, company rollout by week 18, internal audit in month 6, and mock audit in month 7.
2. Install a seven-day incident-response floor (depends on: 1)
Do not wait for policy design or tooling migration. Put a minimum process into operation immediately and begin retaining evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communication schedule.
- Create one monitored declaration path through chat, telephone, and the current paging environment.
- Use one incident channel, bridge, timeline document, and naming convention for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced on-call engineers and engineering leaders.
- Require a named commander within 10 minutes. The duty engineering director assumes command if nobody else does.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Compensate interim duty retroactively under the permanent policy.
- Hold a 15-minute daily operational review until the permanent process is live.
3. Create the factual, legal, and audit baseline (depends on: 1)
Build one defensible baseline for process design, executive decisions, and SOC 2 testing. Confirm the auditor's expected Type II observation period immediately.
- Reconstruct all 31 incidents from first customer impact through resolution, communications, credits, postmortem, and corrective action.
- Analyze the two incidents with unclear command minute by minute.
- Identify the missing detection signal for every customer-first incident.
- Inventory the six alert sources, 3,400 monthly alert events, noisy rules, duplicates, missing owners, and missing runbooks.
- Record the current rotations, unpaid work, after-hours load, and teams without coverage.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M credits, 85% noise, and 11 of 64 actions closed.
- Map incident-response controls to applicable SOC 2 criteria with Compliance and the auditor.
- Establish retention, confidentiality, legal-hold, and access requirements for incident evidence.
4. Co-design the fairness contract with engineers (depends on: 1)
Treat pager resistance as a valid design constraint. Make the employment and ownership bargain explicit before expanding on-call.
- Interview representatives from all 28 teams and the Support, Customer Success, Security, and Operations groups.
- Distinguish objections involving unpaid work, unfamiliar code, bad alerts, weak runbooks, sleep disruption, or blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for services they own or have formally accepted, supported by trained command, and given capacity to remove recurring defects.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Measure baseline trust, fairness, fatigue, and psychological safety. Repeat the survey at days 60 and 120, then quarterly.
5. Build the service catalog and criticality model (depends on: 3)
Make a machine-readable catalog the source of truth for routing, escalation, customer impact, and audit evidence. Every production service must have one accountable owner.
- Record the owning team, manager, product capability, escalation policy, communication channel, dashboard, runbook, dependencies, regions, and data stores for all 180 services.
- Map payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, and reporting to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing services whose failure can create material contractual impact.
- Classify Tier 2 as internal or deferrable services, and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional mode, recovery method, and contractual obligations for Tier 0 and Tier 1.
- Give orphan services an owner or approved decommission date within 30 days.
- Maintain separate but coordinated ownership for the ledger application and PostgreSQL platform.
6. Adopt one severity scale and incident lifecycle (depends on: 3, 5)
Use four impact-based levels. Classify on actual or credible customer, financial, security, regulatory, and contractual harm rather than organizational seniority.
- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; both regions impaired; or a core payment journey broadly unavailable. Page every role immediately, open a bridge, notify executives, publish customer status within 15 minutes when applicable, begin legal assessment, and require a postmortem.
- **SEV2 — major:** material payment degradation; settlement deadline at risk; a critical customer or material cohort unavailable; regional impairment with reduced resilience; or likely SLA breach. Page command and technical roles, publish customer status within 30 minutes when customer-visible, and require a postmortem.
- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads; page only when immediate action is necessary.
- **SEV4 — operational event:** no current customer impact and no urgent risk. Create a ticket and handle during normal operations.
- Treat percentages as supporting guardrails, never reasons to under-classify integrity, settlement, security, or contractual risk.
- Anyone may declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Automatically use at least SEV2 posture for suspected ledger-integrity events, cross-team incidents with unknown ownership, and materially unknown impact lasting 15 minutes.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed. Resolution requires stability, backlog recovery, and necessary reconciliation.
7. Define roles, authority, and handoffs (depends on: 6)
Separate command, communications, recordkeeping, and technical repair. One named person must hold command throughout every SEV1 and SEV2.
- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not serve as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- **Subject-Matter Responders:** diagnose and mitigate only services for which they have ownership, access, training, or formally accepted support responsibility.
- **Executive Duty Officer:** removes organizational barriers and makes exceptional business decisions without displacing the commander.
- Security, Legal, Compliance, Vendor Management, and Finance join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Require distinct commander, communications, scribe, and primary technical lead for SEV1. Communications and scribe may combine temporarily for bounded SEV2 incidents.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, and reconciliation requirements for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, with the exact transfer time.
8. Create sustainable 24x7 coverage across 28 teams (depends on: 4, 5, 7)
Use central command coverage and risk-based technical coverage rather than creating 28 fragile night rotations. All services receive a response path, but only critical domains maintain direct overnight technical rotations.
- Create a 24x7 Incident Command corps of 24–30 certified people, with primary and secondary coverage at all times.
- Create 18–24 trained Communications Leads and a similarly sized scribe pool using Support, Customer Operations, Engineering Operations, and qualified managers.
- Maintain a surge roster for simultaneous incidents and a 24x7 Executive Duty Officer schedule.
- Group Tier 0 and Tier 1 ownership into approximately 8–12 coherent domains only where responders have training, access, runbooks, and explicit acceptance.
- Staff each critical domain with primary and secondary responders and at least six qualified people. Target no more than one primary week in six.
- Give Tier 2 and Tier 3 services business-hours coverage plus a maintained manager and director escalation path.
- When lower-tier impact becomes SEV1 or SEV2, central command activates the manager escalation path and obtains the necessary owner.
- Route unknown-owner pages to Duty Command and platform triage temporarily. Log each one as a catalog control failure.
- Do not merge small teams merely for schedule convenience. Provide training, reassign services, add staffing, or decommission unsupported services.
9. Approve compensation and fatigue protections (depends on: 4, 8)
End unpaid on-call before expanding mandatory night coverage. HR, Finance, Payroll, and employment counsel should approve the policy within 14 days.
- Use market-validated weekly bands, initially budgeting approximately $800–$1,200 for Tier 0/1 primary duty and $300–$500 for secondary duty.
- Budget approximately $900–$1,300 for Duty Incident Commander weeks and $400–$800 for Communications Lead or scribe duty, adjusted for actual burden.
- Pay holiday premiums and compensate active after-hours work according to exempt or non-exempt status and applicable federal and New York rules.
- Provide a protected paid recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal sprint commitments by approximately 15% during a primary on-call week.
- Prohibit simultaneous primary assignments, consecutive primary weeks, on-call during leave, and invisible schedule swaps.
- Allow responders to declare themselves temporarily unfit after disruptive night work without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Budget roughly $0.9M–$1.2M annually, then refine using actual rotation count, employment classification, and activation data.
10. Establish one paging and incident system of record (depends on: 3, 5, 8)
Monitoring tools may remain specialized, but all human pages must enter one controlled platform. This removes conflicting schedules and creates one evidence trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page.
- Ingest events from the six existing monitoring tools before disabling their direct paging paths.
- Route pages through service-catalog ownership and deduplicate related events.
- Provide a single declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, severity changes, decisions, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status-page publishing, and conference facilities.
- Require role-based access, MFA, access reviews, and immutable or tamper-evident history.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, or an AWS region.
- Retire each legacy human-paging path only after ownership review, end-to-end tests, and two weeks of verified operation.
11. Enforce an alert-quality contract and page budget (depends on: 3, 5, 10)
Treat every page as a production interface with an owner and an expected action. Noise reduction must not create detection gaps.
- Require every paging rule to identify the service, owning team, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Page only when prompt human judgment or intervention can materially reduce customer, financial, security, or contractual risk.
- Route informational, capacity, and non-urgent infrastructure conditions to dashboards or ticket queues.
- Prefer payment-outcome, settlement-risk, queue-age, and error-budget-burn alerts over raw CPU, memory, pod, or log thresholds.
- Run new paging rules in shadow mode for at least seven days unless an emergency exception is approved.
- Review rules with repeated no-action acknowledgements, low actionability, or excessive firing within two business days.
- Set a load budget of no more than two after-hours pages per responder per week, measured over four weeks.
- Require an owner, compensating detection, and recorded approval before suppressing or deleting a rule.
- Burn down the top 50 noisy rules first. Review missed detections and noise together so teams cannot improve metrics by becoming blind.
12. Detect payment and ledger failures before customers (depends on: 5, 11)
Move detection from host health to customer journeys and financial outcomes. Use internal SLOs with enough headroom to protect the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, API access, webhooks, and reporting freshness.
- Run external synthetic transactions through critical payment journeys at least every minute and from paths independent of the production platform.
- Validate each AWS region and expose dependencies that undermine nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed value relative to settlement deadlines, not only queue depth.
- Add tenant and cohort anomaly detection for high-value customers and major payment methods.
- Convert high-priority Support, account-manager, processor, sponsor-bank, and network reports into incident candidates within five minutes.
- Record the detection source for every incident. Treat customer-first detection as a mandatory missed-detection review.
13. Codify detection, escalation, and live execution (depends on: 7, 10)
Create one time-bound path from the first credible signal to named command and mitigation. Notification delivery does not count as human acknowledgement.
- Page the owning critical-domain primary and Duty Incident Commander immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at five minutes, manager at 10, and director at 15.
- Escalate an unclaimed command page to backup commander at five minutes. The Executive Duty Officer assumes command at 10 minutes until a certified handoff occurs.
- Require Support and account managers to use the same declaration path as automated monitoring and engineers.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Begin with a standard statement of severity, known impact, assigned roles, current objective, workstreams, and next update time.
- Freeze unrelated changes during SEV1 and normally during SEV2. Record any exception.
- Prefer reversible mitigation such as rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require formal command handoff for long incidents, shift changes, or fatigue. Do not leave an incident unowned during transfer.
- Reconcile the ledger and safely drain backlogs before resolving payment or ledger incidents.
14. Standardize internal incident communications (depends on: 7, 10, 13)
Give responders one working room and stakeholders one controlled information source. Executives must not interrupt the technical command path.
- Create one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue an initial internal brief within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within 15 minutes for SEV1 and 30 minutes for SEV2.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record material decisions and outbound messages in the incident timeline.
15. Standardize status-page and customer communications (depends on: 6, 10, 14)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate impact before root cause is known.
- Publish customer status within 15 minutes of declaring a customer-visible SEV1 and within 30 minutes for customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice promptly after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map status-page components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, data-integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, known workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Give affected strategic accounts direct account-manager outreach within 30 minutes for SEV1 and 60 minutes for SEV2.
- Require account managers to use the approved briefing and prohibit independent technical explanations.
- Provide a customer-facing incident summary within five business days for SEV1 and qualifying SEV2 events.
- Record any legally necessary delay or restriction of public detail, its approver, and the alternative communication plan.
16. Operationalize legal, regulatory, contractual, and credit decisions (depends on: 3, 6, 15)
Some payments incidents start external notification clocks. Make assessment mandatory without assuming that every operational incident is reportable.
- Build a counsel-validated matrix covering applicable NYDFS rules, state breach laws, GLBA or FTC obligations, PCI requirements, money-transmitter obligations, sponsor-bank and network contracts, cyber insurance, and customer contracts.
- Engage Legal and Compliance immediately for every SEV1, security event, and suspected ledger-integrity event.
- Record an initial reportability assessment within one hour for SEV1 and within two hours for other potentially reportable events.
- Record non-reportable decisions as evidence, including facts considered, approver, timestamp, and reassessment trigger.
- Maintain tested 24x7 contacts for counsel, regulators, sponsor banks, networks, insurers, and critical vendors.
- Encode customer-specific notice deadlines and channels in the customer record used by the Communications Lead.
- Have Legal own regulatory text and submission. Keep technical command with the Incident Commander.
- Have Finance calculate affected minutes, delayed value, likely credits, and contractual exposure within five business days.
- Track credits by incident and recurring cause to support reliability investment decisions.
17. Make postmortems mandatory, consistent, and blameless (depends on: 6, 7, 10)
Use one learning standard with fixed deadlines. Keep postmortems separate from performance, misconduct, and disciplinary processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact over two hours, SLA credits, contractual breach, repeated contributing factors, control failures, and ledger-integrity near misses.
- Produce the factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the Incident Commander responsible for the timeline and the owning engineering director accountable for completion.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit analysis of why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the incident commander.
- Describe decisions in the context and information available at the time. Do not name an individual as the root cause.
- Publish broadly useful findings internally while restricting security, privacy, personnel, or privileged material appropriately.
18. Give corrective actions enforceable ownership (depends on: 10, 17)
Treat incident actions as risk commitments, not suggestions. A ticket is not complete until the expected risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Escalate overdue high-risk actions to the manager after seven days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit overdue actions to block related releases where the unrepaired condition could reproduce severe impact.
- Verify effectiveness through tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days. Complete, re-plan, or formally accept each risk.
19. Set the service-readiness bar and critical playbooks (depends on: 5, 8, 11, 12)
A critical service must be supportable at 3 a.m. before its team is placed on direct overnight coverage. Existing critical detection must remain active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and data-integrity constraints.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional failure, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document controlled failover, split-brain prevention, replay protection, backlog recovery, and post-recovery reconciliation.
- Exercise critical runbooks at least twice per year and after material changes.
- Block new Tier 0/1 releases and new paging rules when readiness requirements are missing.
- Handle existing gaps through a named owner, compensating control, executive-approved expiry date, and remediation plan rather than disabling detection.
20. Publish policy and certify every response role (depends on: 6, 7, 8, 9, 11, 13, 14, 15, 16, 17, 18)
Convert the operating design into concise, signed documents and practical training. Training and exercises occur during paid working time.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, and exception standards.
- Provide one-page severity, role, authority, escalation, and communication cards inside the incident tool.
- Train all employees to recognize and declare incidents.
- Train all engineers in severity, acknowledgement, evidence preservation, handoff, and financial-integrity precautions.
- Certify responders only after they demonstrate access, dashboards, runbooks, rollback, and escalation competence.
- Certify Incident Commanders through formal instruction, simulation, and at least two shadowed incidents or exercises.
- Train Communications Leads in status writing, account segmentation, contractual clocks, and legal escalation.
- Train scribes in timeline quality and fact-versus-hypothesis labeling.
- Require two shadow shifts before independent primary duty and annual recertification.
- Maintain the training and certification register as operational and audit evidence.
21. Pilot on the payment critical path (depends on: 9, 10, 12, 19, 20)
Run a four-to-six-week pilot across the highest-risk journey before expanding. Use real incidents and exercises to correct the model quickly.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, API edge, authentication, settlement, and Support intake.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently assuming command.
- Correct critical process or tool defects within 48 hours and update the standard visibly.
- Exit only after 95% timely command assignment, 95% communication compliance, no unpaid pages, complete required postmortems, and at least 50% lower noise.
22. Roll out by customer journey and risk (depends on: 21)
Expand in controlled waves while the interim process remains active company-wide. Complete rollout early enough to accumulate operating evidence before the audit.
- Roll out remaining Tier 0 domains first, then Tier 1, Tier 2, and Tier 3.
- Use two-to-three-week waves with a named program coach and director sign-off.
- Gate each team on catalog ownership, appropriate coverage, compensation, trained responders, tested escalation, alert quality, runbooks, access, and a passed tabletop.
- Require six or more responders only for direct 24x7 technical rotations. Apply business-hours coverage and manager escalation to lower tiers.
- Reschedule failed gates or approve a time-limited executive exception with a compensating control.
- Disable legacy human-paging paths after verified cutover for each wave.
- Publish an internal adoption dashboard by team, service tier, and control gap.
- Finish critical coverage by week 10 and all 28 teams by week 18.
23. Exercise command, communications, regional recovery, and fallbacks (depends on: 19, 20)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous security and operational events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors where appropriate.
- Conduct at least one unannounced after-hours paging test before the audit.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure time to acknowledgement, command, customer notice, mitigation decision, handoff, and recovery.
- Create tracked actions for every material exercise finding.
24. Measure performance and operate fixed review forums (depends on: 3, 10, 17, 18)
Use a balanced scorecard that exposes weak controls without rewarding hidden incidents or suppressed alerts. Report medians and 90th percentiles rather than averages alone.
- Measure time from first impact to detection, declaration, acknowledgement, command assignment, mitigation, recovery, and resolution.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track pages, actionability, duplicates, after-hours load, missed detections, and page-budget breaches.
- Track postmortem timeliness, action age, closure by due date, verified effectiveness, and recurring contributing factors.
- Track availability by customer journey, error-budget burn, failed or delayed payment value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight activations, recovery days, schedule exceptions, sentiment, and attrition.
- Hold a weekly Incident Review Board for incidents, postmortems, control failures, noisy alerts, and overdue actions.
- Hold a monthly executive reliability review for trends, funding, contractual exposure, and accepted risks.
- Hold a quarterly control and resilience review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile incident records monthly against support cases, status history, credits, and customer complaints to detect under-reporting.
25. Prove SOC 2 operating effectiveness before fieldwork (depends on: 20, 22, 23, 24)
Generate audit evidence through normal operation rather than reconstructing it later. Test both control design and consistent execution.
- Maintain approved and versioned policies, exceptions, catalog records, schedules, compensation activation, access reviews, training, incidents, communications, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, non-reportable legal decisions, missed timelines, and exercises in the population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the populations, evidence requests, and interviews expected from the external auditor.
- Record deviations honestly with owners, remediation dates, and compensating controls. Never rewrite historical records.
- Confirm that evidence retention covers the full auditor-defined observation period.
- Brief commanders, responders, Support, and Compliance on the real process without scripting inaccurate answers.
26. Institutionalize improvement and reduce structural risk (depends on: 22, 24, 25)
Prevent the program from decaying after the audit. Use incident evidence to drive permanent ownership and architectural investment.
- Assign permanent owners for policy, service catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review severity thresholds, staffing, compensation, and communication timings annually and after material process failures.
- Recertify command and communications personnel annually.
- Review recurring failure families quarterly and require executive decisions where corrective work repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fairness, fatigue, and tool friction.
- Report severe incidents, credits, overdue high-risk actions, and resilience investment to the board or risk committee quarterly.
- Treat the shared ledger cluster as a strategic concentration risk. Fund failover assurance, blast-radius reduction, isolation of non-critical readers, and stronger regional independence.
- Evaluate follow-the-sun command or technical coverage using one year of page-load and staffing data.
- Build year-two plans for automated mitigation, deployment safety, graceful degradation, and error-budget release controls.
Previous Proposal 3 (ID: f7ee763b-3b76-4f07-9e7d-5164701bef1c, Agent: qwen3.8-max_refine_3, LLM: alibaba/qwen3.8-max):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (32):
1. Establish executive mandate, program office, and funding
Turn the CEO email into a **chartered company program** with one accountable owner and authority over all 28 teams.
- Appoint the CTO as executive sponsor and a Head of Reliability as program owner with full-time authority.
- Create a permanent program office: one program lead, one platform engineer, one analyst.
- Form an eight-person steering group: Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours.
- Lock the non-negotiables: one severity scale, one paging platform, one postmortem format, paid on-call, mandatory action tracking, named service ownership.
- Approve budget anchored against the $1.3M in credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs.
- Reserve 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Publish a one-page charter stating incident response is a company operating process, not a per-team choice.
2. Install a seven-day interim command bridge (depends on: 1)
Do not leave the company unprotected while the permanent process is designed. Put a **crude but real** command structure in place within one week.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Direct Support to escalate credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical actions: complete, re-plan, or formally risk-accept, prioritizing ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Build the forensic baseline of incidents, alerts, and money lost (depends on: 1)
Rebuild the facts before designing anything. This becomes both the **design input** and the before picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific missing signal. This list drives the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Freeze baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the on-call fairness deal (depends on: 1)
Engineer pushback is the **largest delivery risk**. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS, and Sales in two weeks. Separate the real objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey covering fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
5. Map compliance, evidence, and audit requirements from day one (depends on: 1)
Design evidence as a **by-product of operations**, not a reconstruction before the auditor arrives. Confirm the SOC 2 observation window immediately.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for communications; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure, training register, drill records, reportability decisions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve all policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has a **named owner**. Build a machine-readable catalog as the single source of truth.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact: Tier 0 (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for each Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
7. Adopt the severity scale, declaration rules, and lifecycle (depends on: 3, 6)
Replace judgment calls with a **lookup table**. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter.
- SEV1 (crisis): money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of all roles and executive; bridge in 5 minutes; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 (critical): material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- SEV3 (major, contained): narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected, Declared, Triaged, Mitigated (customer impact ends), Monitoring, Resolved (backlog processed and ledger reconciled), Reviewed.
- Publish a decision tree with 12 worked examples from the real 31 incidents.
8. Define incident roles, authority, and handover discipline (depends on: 7)
Solve nobody-in-charge-for-an-hour by making command **explicit, single-holder, transferable, and logged**. Separate coordination from debugging.
- Incident Commander: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates for SEV1/SEV2.
- Subject-Matter Responders: engineers of the owning team; they mitigate, they do not run the room.
- Executive Duty Officer (SEV1): removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Design the three-layer 24x7 coverage model (depends on: 6, 8)
Do not create 28 night rotations. **Centralise coordination** in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- Layer A, Incident Command corps: approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of approximately 18 from Support, CS, and engineering management. A scribe pool used as the training entry point.
- Layer B, critical-path domain rotations: consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- Layer C, everyone else: business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing: a US-only paid night rotation now, a follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
10. Approve on-call compensation, labor compliance, and fatigue safeguards (depends on: 9, 4)
Unpaid on-call in New York is both a **retention problem and a legal exposure**. Pay must be live in payroll before any mandatory night rotation starts.
- Indicative scheme: approximately $1,000 per primary 24x7 week, $400 secondary, $250 for business-hours rotations, a separate $1,200 Duty Commander stipend, holiday premiums, approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- HR, Finance, and employment counsel publish amounts, eligibility, tax treatment, and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a 15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
11. Set the alert quality standard, page budget, and noise burn-down (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes. Make alert quality a **condition of being allowed to page a human**.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly.
- Target 3,400 to under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
12. Build payment-outcome detection and ledger assurance (depends on: 11, 6)
Stop customers telling you first. Detection must be driven by **payment outcomes and ledger truth**, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer detected first becomes a named defect class with a mandatory tracked action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Consolidate to one paging platform, one incident record, one status page (depends on: 7, 9, 11)
Collapse six alerting tools into a **single operational system of record** so there is one queue, one timeline, and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is unavailable. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
14. Codify the escalation ladder and five-minute command rule (depends on: 8, 9, 13)
Write one unskippable path from something looks wrong to **someone is in charge**. The default action is never waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately, secondary at 5 minutes, domain manager and Duty Commander at 10, Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension, and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
15. Write the live incident execution doctrine and major-incident playbooks (depends on: 8, 13, 14)
Give responders one short operating procedure from the first minutes through closure. Priority is **limiting customer and financial harm**, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the single largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window, an explicit handback to the owning team, Support, and Customer Success, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
16. Standardize internal communications protocol (depends on: 8, 13)
Standardise the internal picture so executives, Support, and Sales are informed **without interrupting** the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
17. Build customer communications, status page, and account-manager outreach (depends on: 7, 16)
Replace whoever is around with a **timed, owned, pre-approved process**. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook, and RSS. All 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity, or blame.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
18. Create the regulatory, partner, and legal notification playbook (depends on: 7, 17)
In payments, some incidents start a **legal clock at detection**. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded even when the answer is not reportable, with evidence, decision-maker, and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
19. Operationalize SLA credit and financial impact workflow (depends on: 7, 17)
Tie incidents to money so severity, credits, and investment decisions **stay honest**, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one. Use the delta as the standing business case for on-call pay and reliability capacity.
20. Establish the blameless postmortem standard and Incident Review Board (depends on: 7, 8)
Replace some incidents, various formats with **one mandatory format, fixed deadlines, and a forum with teeth**.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly top-five recurring causes analysis.
21. Enforce action ownership, capacity reservation, and tracking (depends on: 20, 13)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of **customer commitments** and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
22. Build the training and certification academy (depends on: 8, 14, 16, 20)
Command is a skill, not a title. **Certify before assigning duty**, and use paid working time for all of it.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days + shadowing): command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Comms Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum; never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Run the exercise programme: tabletops, game days, and unannounced drills (depends on: 22, 13, 15)
The process must meet a **simulated SEV1** before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger. Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (depends on: 7, 8, 9, 10, 11, 14, 16, 17, 18, 20, 21)
Collapse the design into a document people will **actually open mid-outage**, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 12, 13, 15, 22, 10)
Prove the process on the **highest-risk surface** with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment.
- Publish a one-page result to the whole company.
26. Run the alert noise burn-down campaign (depends on: 11, 13, 25)
Run noise reduction as a **visible, quota-driven campaign** in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness. Review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 to 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
27. Establish metrics, dashboards, review cadence, and anti-gaming (depends on: 13, 20, 25)
Instrument the process itself so improvement is visible and the auditor sees **evidence of monitoring and review**. Report median and 90th percentile, never averages alone.
- Response: time to detect, declare, commander, acknowledgement, mitigate, resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
28. Wave rollout to all 28 teams with readiness gates (depends on: 25, 27)
Roll out in four waves of six to eight teams every two to three weeks, ordered by **customer risk**. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. Failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Publish a live adoption scoreboard.
29. Run change management, fairness, and pager culture programme (depends on: 4, 10, 25)
Run this from day one in parallel. Engineers judge the process on **fairness**; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. Build SOC 2 evidence by design, internal testing, and mock audit (depends on: 24, 27, 28)
Design the evidence as a **by-product of doing the work**. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any change as an exception.
31. Maintain the program risk register and contingencies (depends on: 1)
Name the ways this programme fails and **pre-commit the response**. Review it monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 27, 28, 30)
Guard against the classic failure: the process **decays once the audit is signed**. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 4 (ID: 101b9773-a36a-47bf-8af6-14fe9a48abb2, Agent: grok4.6_refine_4, LLM: xai/grok-4.6):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (26):
1. Charter the program and start the audit clock
Convert the CEO email into a company operating process with one owner, a budget, and an observation window that starts this week.
Incident response is no longer a per-team choice.
- Name the CTO as sponsor and a Head of Reliability as the single accountable owner, with a three-person program office.
- Form a small decision group: Engineering, SRE, Support, Customer Success, Security, Legal, Finance, and HR. The sponsor decides within 48 hours.
- Lock non-negotiables: one severity scale, one paging path, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, and reserved engineering capacity against the $1.3M in SLA credits.
- Timeline: operating floor in 7 days, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, mock audit in month 7, audit in month 8.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for new tools or the final policy. Put a crude but real process in place this week so the next outage has a named commander.
This week is the first audit evidence.
- Publish a one-page interim severity card and one declaration path: Slack command, phone number, and existing pagers, all reaching the same duty person.
- Staff interim primary and backup Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in the incident channel.
- Tell Support to declare from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline (depends on: 1)
Rebuild the facts before locking design. This is both the design input and the before picture for the CEO and the auditor.
Freeze the numbers so they cannot drift during design.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisy rules and every rule with no owner or runbook.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11 of 64 actions closed, on-call in 12 of 28 teams.
4. Map resistance and publish the fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
Publish the deal in writing before any new mandatory pager is assigned.
- Interview all 28 teams plus Support, Customer Success, and Sales in two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the **fairness contract**: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and money-path tiers (depends on: 3)
You cannot page the right person across 180 services until every service has one owner. A wrong owner recreates the pager objection.
The catalog is the single source of truth for paging, impact, status-page components, and audit.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to services, data stores, regions, and third parties, including sponsor banks and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
6. Lock severity levels and what each one triggers (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by who reported it or how hard the fix looks.
Anyone may declare. Nobody is punished for over-declaring. Only the Incident Commander may downgrade, with the evidence recorded.
- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 major, contained: narrow or single-customer impact with a workaround. Owning team leads; business-hours comms; postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages.
- Attach a financial-integrity or security flag to any severity. The flag forces dual-control, Legal, and the regulatory checkpoint without inventing a fifth level.
- Auto-escalate: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer impact ends) → Monitoring → Resolved (backlog processed and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged.
Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. A human validates for SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1 and SEV2. Every handover is announced verbally and in writing with the exact time.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
8. Staff 24x7 with a command corps, not 28 night rotations (depends on: 5, 7)
Do not create 28 night rotations. That is what engineers are rejecting.
Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified volunteers from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to seven months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate 28 teams into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat Lisbon or APAC follow-the-sun as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
9. Pay on-call, meet New York labor rules, and cap fatigue (depends on: 4, 8)
Unpaid on-call in New York is a retention problem and a legal exposure. Pay must be in payroll before any new mandatory night rotation starts.
Publish the numbers. Then ask people to sign up.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
10. Set the alert-quality bar and a hard page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them.
Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers do (depends on: 5, 10)
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail.
Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One Slack command creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is down. Test weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Write the five-minute path from signal to command (depends on: 7, 8, 12)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting.
If nobody claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Codify live execution, readiness, and ledger playbooks (depends on: 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving a root cause.
A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down. No command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers.
15. Run one communications clock for internals, customers, and regulators (depends on: 6, 7, 12)
Replace whoever is around with one timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is not reportable.
- Obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
16. Tie incidents to SLA credits and true financial cost (depends on: 6, 15)
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
Credit calculation is an output of the incident record, not a negotiation.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
17. Make blameless postmortems mandatory and consistent (depends on: 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth.
The discipline lives in the deadlines and the review, not the template.
- Mandatory for every SEV1 and SEV2, any incident detected by a customer first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed.
18. Track actions as risk commitments with reserved capacity (depends on: 12, 17)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
Closing a ticket without evidence of effectiveness does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Target 90% of high-priority actions closed by due date within two quarters.
19. Train and certify every role before independent duty (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
The certification register is an audit artefact.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
20. Rehearse with tabletops, game days, and night drills (depends on: 12, 14, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
21. Publish Policy v1 and the signed fairness contract (depends on: 6, 7, 8, 9, 10, 13, 15, 17, 18)
Collapse the design into a document people will actually open mid-outage, and make it official.
Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
22. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 19, 21)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse. Publish a one-page result to the whole company.
23. Instrument metrics, reviews, and anti-gaming (depends on: 12, 17, 22)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review.
Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
24. Roll out in risk-ordered waves with readiness gates (depends on: 22, 23)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave, not as a background hope. Give every team its noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
25. Produce SOC 2 evidence by operating, then mock-audit (depends on: 21, 23, 24)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions. Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (depends on: 23, 24, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
Name the ways this programme fails and pre-commit the response.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 5 (ID: 289def1b-8f78-402b-85b3-0d9e18661842, Agent: deepseek-v4-pro_refine_5, LLM: deepseek/deepseek-v4-pro):
Estimated Complexity: high
Success Metrics: - Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, no loss of Tier 0/1 detection.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% Tier 0/1 services have named owning team, tier, escalation policy, dashboard and runbook by day 30; all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10-12 domain rotations of at least six trained responders.
- 30 certified incident commanders and 18 certified communications leads active.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% required postmortems drafted in 3 business days, reviewed in 5, published in 10 from month 2.
- Postmortem action closure rises from 17% to 90% closed by due date within two quarters; all 53 legacy actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k annualised within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% SEV1/security SEV2 including not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two cross-company exercises completed before audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% sampled incidents; SOC 2 Type II incident response controls pass zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days; no increase in attrition among responders.
Steps (26):
1. Executive mandate, budget, and governance
Secure a written CTO/CEO charter making incident management a company operating process, not a team option. Name one accountable Head of Reliability and a small steering group with authority to decide within 48 hours.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking, and protected engineering capacity.
- Approve budget against the $1.3M annual credits: tooling $150-250k/yr, on-call compensation, training, and 2-3 program FTEs.
- Set timeline: interim process in week 1, design weeks 1-6, pilot weeks 7-12, rollout weeks 13-24, mock audit month 7, SOC 2 at month 8.
- Freeze current baselines: 31 incidents, 22-min MTTD, 40% customer-first detection, 3h10 MTTM, $1.3M credits, 3,400 alerts/month at 85% noise, 11/64 actions closed.
2. Interim 7-day command (depends on: 1)
Put a crude but real process in place immediately so no incident remains unowned while the permanent design proceeds.
- Publish a one-page interim severity card, one declaration path (phone, Slack command, pager), and one incident channel/bridge/timeline naming convention.
- Staff an interim 24x7 duty commander with primary and backup from engineering managers and senior SREs; compensate retroactively under final policy.
- Require a named incident commander within 10 minutes of any suspected major incident, announced in channel.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions in 30 days, prioritizing ledger integrity, duplicate payment, regional failover, security and detection gaps.
- Hold a daily 15-minute operations review until the permanent process is live.
3. Forensic baseline and alert estate analysis (depends on: 1)
Reconstruct the true before picture from the last 12 months; it drives design and serves as audit baseline.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact start/detect/declare/commander/mitigate/resolve, who led, credits paid, contributing factors.
- For every customer-first detection, name the missing signal; this becomes the detection backlog.
- Reconstruct the two `nobody in charge` incidents minute by minute; use them as the burning-platform narrative and test case.
- Profile the 3,400 monthly alerts by tool, team, rule, outcome; identify top 50 noisy rules and every rule with no owner or runbook.
- Freeze all baseline metrics in one signed document for the executive and the auditor.
4. Listening tour and fairness contract (depends on: 1, 3)
Treat engineer pushback as the main delivery risk and convert it into a written deal that removes the objection to carrying a pager for other teams' code.
- Interview all 28 teams plus Support, CS and Sales in two weeks to separate objections: unpaid work, nights, unfamiliar code, missing runbooks, or fear of blame.
- Harvest practices from the 12 teams already on-call; they supply pilot teams and first commanders.
- Publish the deal: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Recruit 10-15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey and repeat at 60, 120 and 365 days.
5. Service ownership catalog and criticality tiers (depends on: 3)
Create a machine-readable catalog as the single source of truth for paging, routing, status page components, and audit evidence.
- Assign one accountable team, engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list per service.
- Tier by business impact: Tier 0 money movement/ledger/auth/shared PostgreSQL; Tier 1 customer-facing degradable; Tier 2 internal/batch; Tier 3 non-critical.
- Map customer journeys (initiate, authorize, settle, reconcile, report, onboard) to services, data stores, regions, third parties.
- Give orphan services an owner within 30 days or a decommission date approved by the sponsor; Tier 0 without an owner is an executive escalation.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
6. Severity scale, declaration rules, and automatic triggers (depends on: 3, 5)
Adopt one severity scale as a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade with recorded rationale. Start high when uncertain.
- SEV1: money moved wrongly/duplicated/lost, ledger integrity in doubt, confirmed security/data exposure, both regions impaired, payment processing halted. Full role page, bridge in 5 min, status page in 15 min, regulator assessment within 1 hour, mandatory postmortem.
- SEV2: material degradation of a payment journey, settlement window at risk, large/strategic customer fully down, SLA breach likely. Commander and SMEs paged, status page in 30 min, mandatory postmortem.
- SEV3: narrow or single-customer impact with workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any ledger cluster incident, any SEV3 open >2h, any unknown impact after 15 min, any cross-team incident become at least SEV2.
- Publish a decision tree with 12 worked examples from the 31 real incidents.
7. Incident roles, authority, and handoffs (depends on: 6)
Codify roles so coordination never depends on seniority or heroics. The Incident Commander coordinates; responders fix only services they own.
- IC: owns severity, priorities, roles, cadence, closure; pre-authorized to freeze deploys, rollback, disable features, shift traffic, invoke failover, put ledger read-only, commit spend; does not type in terminals; keeps command when VP joins.
- Communications Lead: sole author for status page, account managers, executives, and hand-off to Legal for regulators; speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, state changes; human validates for SEV1/SEV2.
- Subject-Matter Responders: diagnose and mitigate only their owned services.
- Executive Duty Officer (SEV1): removes obstacles, shields IC from exec questions, owns board/regulator escalation; does not take command unless formally transferred.
- Rule: IC claimed within 5 min and announced; distinct people for IC, Comms, technical lead at SEV1/SEV2; every handover announced verbally and in writing with time.
- Financial controls survive: IC coordinates ledger recovery but cannot bypass dual control, reconciliation or privileged access.
8. 24x7 command and responder coverage model (depends on: 5, 7)
Do not create 28 fragile night rotations. Centralize coordination in a trained command corps, keep technical ownership local.
- Layer A: incident command corps of ~30 certified volunteers with primary/secondary 24x7; paired Comms Lead pool ~18 and scribe pool as entry.
- Layer B: consolidate 28 teams into 10-12 critical-path domains (ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes, data/reporting, integrations) each with 24x7 primary+secondary and minimum six trained responders.
- Layer C: all other teams business-hours on-call with written after-hours escalation lists held by managers.
- Platform on-call is safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate follow-the-sun only as a 12-month option, not year-1 dependency.
- Acknowledgement discipline: 5 min at SEV1/SEV2 with automatic failover.
9. Paid on-call and fatigue safeguards (depends on: 8, 4)
Unpaid on-call in New York is a retention and legal risk. Pay must be in payroll before any mandatory night rotation starts.
- Indicative: ~$1,000 per primary 24x7 week, secondary ~$400, business-hours ~$250, duty commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond one hour.
- Mandatory recovery: paid recovery day after >2h overnight work, SEV1, or qualifying SEV2; managers arrange coverage.
- HR/Legal/Finance publish amounts, tax treatment, FLSA/NY wage-hour status within 14 days.
- Load rules: no primary more often than 1 in 6, no consecutive primary/secondary weeks, no primary on two rotations, no on-call during PTO, annual cap.
- If a rotation averages >2 out-of-hours pages per person per week for 4 weeks, trigger staffing or alert-remediation review.
- Count on-call as ~15% delivery load; document exemption path for health/caring without career penalty.
10. Service readiness bar and runbooks (depends on: 5, 8)
A service must earn the right to page a human at 3 a.m. Define a minimum readiness bar and major incident playbooks before any Tier 0/1 service goes live with on-call.
- Require for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switch, escalation contacts, data-loss/latency impact.
- Write playbooks for top failures: shared PostgreSQL failure/corruption, cross-region failover, Kubernetes control-plane loss, processor/sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, duplicate payment.
- Treat ledger cluster as largest structural risk: documented failover, read-only degraded mode, reconciliation after recovery, executive-signed RPO/RTO.
- Start a parallel workstream on blast-radius reduction: tenant/function partitioning, read replicas, isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging at night unless manager accepts a dated exception in writing.
- Runbooks are peer-reviewed, version-controlled, marked stale if not exercised twice a year.
11. Alert quality standard and page budget (depends on: 3, 5)
Make alert quality a condition of paging a human. Burn down noise deliberately rather than by mass silencing.
- Every paging alert must declare owning team, affected service, customer/SLO impact, expected action, dashboard, runbook, dedup key, severity mapping, escalation policy. Anything failing becomes a ticket.
- Page on symptoms of customer harm (SLO burn rate, payment success rate, queue age vs settlement deadlines), not raw CPU/memory.
- Run new alerts in shadow mode 7 days; test firing and recovery.
- Set page budget of 2 out-of-hours pages per person per week; breach blocks new alert creation and triggers tuning sprint.
- Auto-quarantine rules firing >5 times/month without action or >70% no-action acks; return to owner with correction deadline.
- Never silently disable: verify compensating detection, record decision, name owner, review missed detections monthly.
- Target 3,400 to <500 pages/month, actionability >75% within six months, no loss of Tier 0/1 detection.
12. Detection uplift on money path (depends on: 5, 11)
Stop customers telling you first by detecting on payment outcomes and ledger truth.
- Define SLOs and business SLIs per customer journey: initiation success, authorization latency, settlement timeliness, reconciliation break rate, API availability, reporting freshness; set internal targets stricter than 99.95%.
- Run external synthetic end-to-end payments every 60 seconds from both regions, covering all critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate detection, replication lag, failover-readiness, settlement-window countdown.
- Add per-customer anomaly detection for top 100 accounts; five similar support tickets in 10 minutes auto-creates triage incident.
- Route partner and processor notifications into declaration path within 5 minutes.
- Record detection source on every incident; treat `customer detected first` as a named defect class with mandatory action.
13. Consolidate to one paging and incident platform (depends on: 6, 8, 11, 12)
Collapse six alerting tools into one operational system of record: one queue, one timeline, one audit trail, without creating a monitoring gap.
- Select one paging/scheduling platform, one incident record, one hosted status page; time-box selection to 2 weeks.
- Ingest all six sources first, deduplicate/correlate; retire legacy path only after named owners, successful end-to-end test, and 2 weeks verified operation.
- One-command declaration in Slack creates channel/bridge, pages commander, sets severity, opens timeline, starts clock.
- Route every page through service catalog: service label -> owning domain schedule -> escalation policy.
- Capture evidence automatically: declaration, acks, role assignments, severity changes, decisions, comms, mitigated/resolved times, postmortem link; retain 12 months.
- Test out-of-band SMS/phone paging, mobile fallback, offline runbook weekly; ensure works during one AWS region/chat/identity provider failure.
- Set hard date after which pages outside this tool create no on-call obligation.
14. Escalation paths and acknowledgement SLAs (depends on: 7, 8, 13)
Write one unskippable path from signal to named commander in under 5 minutes; default action is never waiting.
- Converge all entry points on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately -> secondary at 5 min -> domain manager + Duty Commander at 10 -> Executive Duty Officer at 15. SEV2 requires command assigned within 15 min.
- If no one claims command in 5 min, platform assigns and announces it; assignee may hand over but not decline.
- Commander may page any domain's on-call directly with 10-min ack obligation; this reciprocity makes single-team ownership viable.
- If ownership unclear after 10 min, commander keeps incident and names temporary owner; missing catalog entry logged as control defect.
- Pre-authorize regional failover, ledger read-only, payment suspension, partner-bank notification so IC never waits for an executive.
15. Standardize internal and customer communications (depends on: 6, 7, 13)
Replace whoever is around with a timed, owned, pre-approved process; the Communications Lead is single author and never writes from scratch under pressure.
- Internal: one working channel/bridge plus one read-only broadcast for executives/Support/Sales; cadence SEV1 15-30 min, SEV2 60 min even if nothing changed; fixed template with impact, action, next update, IC/Comms names.
- Executives ask questions only of Executive Duty Officer; publish this as signed behavior rule.
- Status page: SEV1 within 15 min, SEV2 within 30 min; updates every 30/60 min; resolution notice within 30 min of verified recovery; customer-facing summary within 5 business days for SEV1/qualifying SEV2.
- Pre-approve 12-15 templates with Legal for degradation, settlement delay, API errors, regional failure, data-integrity investigation, security event.
- Top 100 accounts get named account manager call/email within 30 min of SEV1 with briefing pack; all 2,100 subscribed by default.
- Language rules: state impact and next update; never speculate on cause, recovery time, data integrity or blame.
16. Regulatory, partner, and financial impact workflows (depends on: 6, 15)
In payments some incidents start a legal clock at detection. Build obligation assessment into the process and tie incidents to money.
- Legal/Compliance produce obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC, PCI, money-transmitter, sponsor-bank/card-network windows, cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and security-related SEV2 within 1 hour, recorded even when `not reportable`, with decision-maker and evidence.
- Maintain 24x7 contact matrix for regulators, sponsor banks, networks, outside counsel, insurer.
- Agree availability measurement method per contract; compute affected minutes per customer from incident record and journey telemetry; propose credit schedule within 5 business days.
- Attribute credits to root-cause families; track failed payment volume, delayed value, reconciliation breaks; target annual credits from $1.3M to under $400k.
17. Blameless postmortem standard and action tracking (depends on: 6, 7)
Replace various formats with one mandatory format, fixed deadlines, and enforceable actions. Closing a ticket without evidence does not close the action.
- Mandatory for every SEV1/SEV2, customer-first detection, incident >2h, repeat of known cause, credit-generating/contract breach, ledger near-miss.
- Draft within 3 business days, review within 5, publish within 10; IC owns delivery, owning manager accountable.
- One template: summary, customer/financial impact, detection source/gap, timeline, response analysis, contributing conditions, what worked, actions.
- Two mandatory questions: why did a customer see this first? why did mitigation take as long as it did?
- Blameless in writing: systems and context, no individual named as cause, never used in performance reviews; HR handles personnel/misconduct separately.
- Every action gets owner, priority, due date, verification method, ticket auto-created; classes containment 7 days, corrective 30, strategic 90; reserve 15% engineering capacity.
- Escalation: manager at +7 days, director at +14, CTO dashboard at +30; overdue P0 can block related releases.
- Weekly Incident Review Board ratifies severity, challenges quality, monitors actions; target 90% closure on time within two quarters.
18. Train and certify incident roles (depends on: 7, 13, 15)
Command is a skill, not a title. Certify before duty; use paid working time.
- All employees: 30-min module on recognizing impact, declaring, finding channel/status page.
- Responder: half-day on severity, escalation, runbook, financial integrity; mandatory before joining rotation.
- Scribe: 2 hours timeline discipline; entry point.
- Incident Commander: 2 days + shadowing; command presence, delegation, severity calls, running room, handover; certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead: 1 day on status writing, customer tiering, legal boundaries, regulator triggers.
- Domain responders demonstrate dashboard, runbook, rollback, failover, access before primary; two shadow shifts; no primary in first 90 days.
- Certification valid 12 months, renewed by simulation; register is audit artifact.
19. Simulations and game days (depends on: 10, 14, 15, 18)
Rehearse process before real SEV1; exercise tooling failure and ledger scenarios.
- Monthly tabletop per group reusing an incident from baseline, rotating commander.
- Quarterly game day in staging or tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes degradation.
- Twice-yearly unannounced paging drill to measure real overnight ack times.
- Annually one combined operational/security exercise and one regulator-notification exercise with Legal/CISO.
- Exercise status page down, chat/paging provider unavailable.
- Never inject uncontrolled change into production ledger; validate backups/restore/RPO/RTO on replicas.
- Every exercise produces tracked actions; publish time-to-commander, time-to-first-update, time-to-mitigation.
20. Pilot on critical path (depends on: 9, 11, 12, 13, 14, 15, 17, 18, 19)
Prove full process on highest-risk surface with willing teams for six weeks before full rollout.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, Support intake, including two of the 12 already on-call teams.
- Activate severity scale, command corps, single paging tool, page budget, status page policy, mandatory postmortems, paid rotations.
- Parallel old paths for one week then cut over; real incidents use new process only.
- Program lead attends every SEV2+ as coach, never shadow commander; review every pilot page within one business day.
- Exit gate: commander named within 5 min in 95% cases, first status update on time, MTTD under 10 min for pilot services, pages halved, no unpaid page, all required postmortems on time, positive sentiment.
- Publish one-page result company-wide.
21. Metrics, dashboards, and review cadence (depends on: 13, 17)
Instrument the process so improvement is visible and audit sees monitoring/review evidence. Report median and p90, never averages alone.
- Response: time to detect/declare/commander/ack/mitigate/resolve split by severity, tier, journey, region, detection source.
- Quality: customer-first rate, status page timeliness, update cadence, missed escalations, role conflicts, alert actionability, out-of-hours pages per person.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat factors.
- Business: availability per journey, error budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, frequency, recovery days, sentiment, attrition.
- Weekly Incident Review Board; monthly reliability review per team; monthly executive review to CEO; quarterly control review with Security/Compliance/Risk; quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, status page history; team scorecards direct help, never individual penalties.
22. Wave rollout to all 28 teams with readiness gates (depends on: 20, 21)
Roll out in four waves by customer risk, each passing an explicit gate rather than a date. Complete all teams by week 24 to leave ~3 months of evidence before audit fieldwork.
- Wave 1 remaining Tier 0, Wave 2 Tier 1, Wave 3 Tier 2, Wave 4 Tier 3/internal.
- Per-team onboarding kit: catalog entry complete, alerts migrated within budget, runbooks at readiness bar, rotation staffed with 6 trained responders or time-limited exception, one IC candidate nominated, one tabletop passed, payroll active.
- Named coach per wave for three weeks; director signs gate.
- Legacy paging paths disabled per wave, not kept as comfort fallback.
- Publish live adoption scoreboard.
- Any team unable to staff fair rotation gets headcount, service reassignment, or explicit executive risk acceptance.
23. Culture, fairness, and continuous feedback (depends on: 4, 9)
Run this in parallel from day one. Engineers judge fairness; executives judge results.
- Repeat the deal in every forum: paid on-call, paged only for owned services, trained commander, real sprint capacity.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
- Weekly office hours first 8 weeks, Slack support channel 4-hour SLA, per-team champion, weekly newsletter with metrics and bad news.
- Incident response contribution in promotion criteria; quarterly awards for best postmortem and biggest noise reduction; public thanks after SEV1.
- Pulse survey at 60, 120, 365 days; if fairness or load red, pause expansion until fixed.
24. SOC 2 evidence design and internal dry-run (depends on: 13, 17, 21, 22)
Make evidence a by-product of operations; test before external auditor.
- Map with Compliance to Trust Services Criteria CC7.2-CC7.5, CC2.2/CC2.3, A1.2; confirm observation window early.
- Evidence set automated and indexed: versioned policies/exception, rotation schedules, compensation activation, incident records with timestamps/roles, paging/ack logs, status history, reportability decisions, postmortems, action closure with verification, training/certification register, drill records, access reviews.
- Internal Audit tests process in months 4 and 6, sampling incidents end-to-end.
- Run formal mock audit month 7 using same evidence populations; include interviews with IC, engineer, Support, Compliance.
- Correct failures through tracked actions, never by editing history.
- Freeze process wording after month 7; any change logged as exception.
25. Risk register and contingency planning (depends on: 1)
Name likely failure modes and pre-commit responses; review monthly with sponsor.
- Too few commander volunteers: make command rostered duty for managers and staff engineers until pool reaches 30.
- Compensation not approved in time: fallback to time-off-in-lieu plus phased stipend, never launch mandatory night on-call uncompensated.
- Tool migration slips: cut scope on incident record layer, never on paging consolidation.
- Noise pruning hides real failure: demote to ticket first, observe 30 days, keep recovery path, monthly missed-detection review.
- Burnout in experienced teams: weekly load monitoring with page caps.
- Major SEV1 mid-rollout: program lead becomes full-time responder, wave schedule slips one wave, sponsor told same day.
- Shared ledger concentration: if blast-radius workstream slips, escalate to board as accepted risk with dated plan.
26. Inspect and adapt; year-two sustainability (depends on: 22, 24)
Prevent decay after audit by revising on data and assigning permanent owners.
- At 90 days live, revise policy using measurements: severity calibration, Layer B/C membership from page data, uncovered shifts, commander burn, missed updates, action closure, survey results.
- Assign permanent owners for policy, paging platform, status page, service catalog, training, metrics.
- Re-baseline targets every six months; shift from lagging to leading metrics: error budget burn, near-miss rate, drill performance.
- Year-two candidates: follow-the-sun coverage, automated mitigation for top recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, completion of ledger blast-radius reduction.
- Keep annual policy review, certification renewal, exercise calendar, board reporting as permanent commitments.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Charter the program and fund immediate action", "description": "Make incident management a **company operating process** within 48 hours. Give one accountable leader authority to set standards across all 28 teams.\n\n- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.\n- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.\n- Include Engineering, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit in a decision group. The sponsor resolves blocked decisions within 48 hours.\n- Approve non-negotiables: one severity model, one human-paging path, one incident record, paid on-call, named service ownership, mandatory postmortems, and tracked corrective actions.\n- Reserve 10%–15% of engineering capacity for detection, runbooks, resilience, and incident actions.\n- Fund tooling, compensation, training, exercises, and program staff. Compare the cost with the existing $1.3M in annual SLA credits.\n- Set the schedule: operating floor by day 7, design during weeks 1–4, pilot during weeks 5–10, rollout during weeks 11–18, control tests in months 4 and 6, and mock audit in month 7.", "dependencies": []}, {"step_id": "S2", "title": "Install a seven-day operating floor", "description": "Do not wait for the final policy or new tooling. Put a minimum viable process into operation immediately and retain evidence from the first day.\n\n- Publish a one-page interim severity guide, declaration procedure, role card, and communications clock.\n- Provide one monitored declaration route through chat, telephone, and the current paging environment.\n- Create one channel, bridge, timeline, and incident identifier for every suspected major incident.\n- Staff temporary primary and backup Duty Incident Commanders from experienced managers and engineers.\n- Require a named commander within 10 minutes. The duty engineering director assumes command if the command page is unclaimed.\n- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.\n- Stop uncompensated mandatory after-hours expansion. Pay interim duty under a temporary stipend, retroactive to program launch.\n- Hold a 15-minute daily operational review until the permanent process is active.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Create the factual, contractual, and control baseline", "description": "Build one defensible baseline for process design, investment decisions, and SOC 2 testing. Preserve the original data so progress cannot be created by changing definitions.\n\n- Reconstruct all 31 incidents from impact start through detection, declaration, command, mitigation, recovery, communications, credits, and corrective actions.\n- Reconstruct the two incidents with unclear command minute by minute.\n- Identify the missing signal for every customer-first detection.\n- Inventory all six alert sources, 3,400 monthly alert events, duplicates, noisy rules, missing owners, and missing runbooks.\n- Record current rotations, unpaid duty, overnight activations, schedule size, and uncovered services.\n- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, 85% noise, $1.3M in credits, and 11 of 64 actions closed.\n- Inventory customer-specific availability definitions, notice periods, credit terms, sponsor-bank obligations, and other incident-related contracts.\n- Confirm the required SOC 2 Type II observation period and evidence expectations with the auditor during week 1.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Publish the on-call fairness contract", "description": "Treat pager resistance as a legitimate design constraint. Adoption depends on a written agreement that separates command from technical ownership.\n\n- Interview representatives from all 28 teams and from Support, Customer Success, Security, and Operations.\n- Separate concerns about unpaid work, sleep loss, unfamiliar systems, noisy alerts, inadequate runbooks, and blame.\n- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.\n- Publish the core promise: responders are paid, paged only for systems they own or have formally accepted and trained to support, and assisted by a separate Incident Commander.\n- State that Platform may temporarily triage unknown ownership but does not inherit another team's service.\n- Provide confidential accommodations for health, disability, pregnancy, or caregiving constraints without career penalty.\n- Measure baseline trust, fairness, fatigue, and alert confidence. Repeat at days 60 and 120, then quarterly.", "dependencies": ["S1"]}, {"step_id": "S5", "title": "Build the service catalog and customer-journey map", "description": "Make a machine-readable catalog the source of truth for routing, impact analysis, status-page components, and control evidence. Every production service must have one accountable owner.\n\n- Record the owning team, manager, business capability, repository, channel, dashboard, runbook, escalation policy, dependencies, regions, and data stores for all 180 services.\n- Map initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, reporting, and onboarding to their dependencies.\n- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.\n- Classify Tier 1 as customer-facing systems that can create material contractual impact.\n- Classify Tier 2 as deferrable internal or batch systems and Tier 3 as non-critical systems.\n- Record SLO, RTO, RPO, regional recovery mode, contractual commitments, and critical vendors for Tier 0 and Tier 1.\n- Keep separate but coordinated ownership for the ledger application and PostgreSQL platform.\n- Give orphan services an owner or approved decommission date within 30 days. Treat unowned Tier 0 services as release-blocking executive risks.", "dependencies": ["S3"]}, {"step_id": "S6", "title": "Adopt the severity model and incident modifiers", "description": "Classify incidents by credible customer, financial, security, regulatory, and contractual harm. Start at the higher plausible severity while scope or integrity remains unknown.\n\n- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; broad loss of a core payment journey; both regions impaired; or processing must be stopped. Immediately page all response roles and the Executive Duty Officer. Open the bridge within 5 minutes, freeze unrelated changes, issue internal notice within 10 minutes, publish applicable customer status within 15 minutes, start legal assessment within 1 hour, and require a postmortem.\n- **SEV2 — major:** material payment degradation; settlement deadline at risk; regional impairment with reduced resilience; a critical customer or material cohort unavailable; or an SLA breach is likely. Page command and technical roles immediately. Issue internal notice within 15 minutes, applicable customer status within 30 minutes, and require a postmortem.\n- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads. Page only when immediate action can reduce harm.\n- **SEV4 — operational event:** no current customer impact and no credible imminent harm. Create a ticket and handle during normal hours.\n- Add FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, and VENDOR modifiers. These invoke specialist controls without distorting customer-impact severity.\n- Use at least SEV2 posture for unknown impact lasting 15 minutes, credible ledger-integrity risk, or a cross-domain incident without clear ownership.\n- Escalate a customer-visible SEV3 that remains unmitigated for two hours.\n- Permit anyone to declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.\n- Publish a decision tree and examples based on the 31 historical incidents.", "dependencies": ["S3", "S5"]}, {"step_id": "S7", "title": "Define roles, authority, and handoffs", "description": "Separate coordination, communications, recordkeeping, and technical repair. One named person holds command continuously throughout every SEV1 and SEV2.\n\n- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not act as the primary technical operator.\n- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, executive updates, and coordination with Legal.\n- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.\n- **Subject-Matter Responders:** diagnose and mitigate only systems for which they have ownership, access, training, or a formally accepted support agreement.\n- **Executive Duty Officer:** removes organizational obstacles and makes exceptional business decisions without displacing the commander.\n- Add Security, Legal, Compliance, Finance, and Vendor Management on modifier-specific triggers.\n- Require distinct commander, Communications Lead, scribe, and technical lead for SEV1. Communications and scribe may combine for the first 10 minutes of a bounded SEV2.\n- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.\n- Preserve dual control, privileged-access restrictions, reconciliation, and change evidence for all ledger operations.\n- Announce every role assignment and handoff verbally and in writing, including the exact transfer time and unresolved risks.", "dependencies": ["S6"]}, {"step_id": "S8", "title": "Create sustainable 24x7 coverage", "description": "Use central command coverage and risk-based technical coverage instead of creating 28 fragile night rotations. Responders do not carry pagers for unfamiliar code.\n\n- Create a 24x7 Incident Command corps of approximately 30 certified people with primary and backup schedules. Two schedules create about 104 weekly assignments a year, or roughly three to four weeks per person annually.\n- Create a 24x7 Communications pool of 18–24 trained people from Support, Customer Operations, Engineering Operations, and management.\n- Create a similarly sized scribe pool. The backup commander temporarily records the first minutes if a scribe has not joined.\n- Maintain a 24x7 Executive Duty Officer schedule and specialist contact paths for Security and Legal.\n- Group Tier 0 and Tier 1 systems into roughly 8–12 coherent response domains only where responders share training, access, runbooks, and explicit support acceptance.\n- Staff every critical domain with primary and secondary responders and at least six qualified people. Target eight where overnight activation is frequent.\n- Give Tier 2 and Tier 3 services business-hours ownership plus tested manager and director escalation.\n- Reclassify any lower-tier service capable of causing severe overnight harm rather than hiding the risk behind manager callback.\n- Route unknown-owner incidents to Duty Command and Platform temporarily. Record each occurrence as a catalog control defect.\n- Prohibit simultaneous primary assignments and two-person 24x7 rotations.", "dependencies": ["S4", "S5", "S7"]}, {"step_id": "S9", "title": "Implement compensation and fatigue protections", "description": "End unpaid on-call before expanding mandatory coverage. Use fixed duty compensation so responders are not rewarded for alert volume.\n\n- Use planning bands of $900–$1,200 per Tier 0/1 primary week and $300–$500 per secondary week.\n- Use planning bands of $1,000–$1,300 per Duty Commander week and $400–$700 for Communications or scribe primary duty.\n- Pay holiday premiums. Compensate all legally compensable active and waiting time for non-exempt staff, including overtime where required.\n- Have HR, Finance, Payroll, and employment counsel approve final bands, tax handling, FLSA classification, New York wage-hour treatment, and schedule constraints within 14 days.\n- Provide a protected recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.\n- Reduce normal delivery commitments by about 15% during a primary week.\n- Prohibit consecutive primary weeks, on-call during leave, hidden schedule swaps, and primary duty more often than one week in six.\n- Allow responders to declare temporary fatigue-related unfitness without career penalty.\n- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.\n- Make active payroll setup, training, access, and readiness hard gates before any new mandatory night rotation starts.", "dependencies": ["S4", "S8"]}, {"step_id": "S10", "title": "Codify the live incident lifecycle", "description": "Give every incident the same operational flow from first signal through verified recovery. The first objective is limiting customer and financial harm, not proving root cause.\n\n- Use the states Detected, Declared, Triaged, Mitigating, Mitigated, Monitoring, Resolved, and Reviewed.\n- Record impact start separately from detection. Use the earliest defensible evidence and revise it transparently when later facts emerge.\n- Open with a standard command message: severity, known impact, assigned roles, immediate objective, workstreams, and next update time.\n- Freeze unrelated production changes during SEV1 and normally during SEV2. Record every exception.\n- Prefer reversible mitigation: rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, partner rerouting, or controlled processing suspension.\n- Separate mitigation and diagnosis workstreams when staffing allows.\n- Keep decisions in the shared incident record rather than direct messages.\n- Require formal command and technical handoffs for shift changes, fatigue, or incidents exceeding four hours.\n- For payment incidents, verify backlog handling, duplicate protection, customer state, settlement exposure, and ledger reconciliation before resolution.\n- Require a severity-specific stability period and explicit handback to the owning team, Support, and Customer Success.", "dependencies": ["S6", "S7"]}, {"step_id": "S11", "title": "Establish one paging and incident system of record", "description": "Monitoring tools may remain specialized, but every human page and major-incident record must enter one controlled platform. This provides consistent routing and an audit trail.\n\n- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page within two weeks.\n- Ingest events from all six monitoring tools before disabling their direct human-paging routes.\n- Route pages using the service catalog and deduplicate events belonging to the same symptom.\n- Provide one declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.\n- Capture declarations, acknowledgements, escalations, role assignments, decisions, severity changes, communications, mitigation, resolution, and postmortem linkage automatically.\n- Integrate ticketing, customer-success data, status publishing, and conference facilities.\n- Apply MFA, role-based access, periodic access reviews, and tamper-evident history.\n- Keep privileged legal or security material in restricted linked records rather than exposing it in the general timeline.\n- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, the status provider, or an AWS region.\n- Retire each legacy paging route only after ownership review, end-to-end testing, and two weeks of verified operation.", "dependencies": ["S5", "S7", "S8"]}, {"step_id": "S12", "title": "Enforce the alert-quality contract", "description": "Treat every human page as a production interface with an owner and a required action. Measure human notification episodes rather than raw monitoring events.\n\n- Require every paging rule to identify the service, owner, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.\n- Define an actionable page as one that causes or materially informs a timely intervention or risk decision.\n- Define noise as duplicate, non-urgent, unactionable, stale, test-generated, or incorrectly routed notification.\n- Page on payment outcomes, error-budget burn, queue age against deadlines, and financial-integrity risk rather than raw CPU, memory, pod, or log thresholds.\n- Run new rules in shadow mode for seven days unless a documented emergency exception applies.\n- Review repeated no-action pages within two business days.\n- Set a page budget of no more than two after-hours notification episodes per responder per week, measured over four weeks.\n- Make a sustained budget breach trigger a tuning sprint and block additional non-emergency paging rules.\n- Require compensating detection and central approval before suppressing a Tier 0 or Tier 1 rule.\n- Never disable an existing critical detector solely because its metadata or runbook is incomplete. Track the gap with a dated remediation owner.", "dependencies": ["S3", "S5"]}, {"step_id": "S13", "title": "Detect payment failures before customers", "description": "Move detection from infrastructure health to customer journeys and ledger truth. Validate coverage against actual historical failures.\n\n- Define SLIs and internal SLOs for initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, APIs, webhooks, and reporting freshness.\n- Set internal objectives with enough headroom to protect the contractual 99.95% availability commitment.\n- Run external synthetic transactions through critical journeys at least once per minute from paths independent of the production platform.\n- Test each region and expose dependencies that defeat nominal regional redundancy.\n- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.\n- Alert on queue age and delayed payment value relative to settlement deadlines.\n- Add tenant and cohort anomaly detection for high-value customers and critical payment methods.\n- Convert credible Support, account-manager, processor, bank, and network reports into incident candidates within five minutes.\n- Replay all 31 historical incidents. Record which current detector would fire and at what minute.\n- Treat customer-first detection as a mandatory missed-detection review with a tracked action.", "dependencies": ["S5", "S12"]}, {"step_id": "S14", "title": "Implement the detection and escalation ladder", "description": "Create one time-bound path from the first credible signal to named command and the correct technical owner. Device delivery does not count as acknowledgement.\n\n- Converge automated alerts, engineer observations, support cases, account-manager reports, partner notices, and customer calls on the same declaration path.\n- Page the Duty Incident Commander and owning critical-domain primary immediately for suspected SEV1 or SEV2.\n- Escalate an unacknowledged technical page to secondary at 5 minutes, manager at 10 minutes, and director at 15 minutes.\n- Escalate unclaimed command to the backup commander at 5 minutes. The Executive Duty Officer assumes temporary command at 10 minutes until a certified transfer occurs.\n- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.\n- Keep command with the current commander when service ownership remains unclear. Assign a temporary technical lead and record the ownership gap.\n- Maintain tested external escalation routes for AWS, database support, processors, sponsor banks, networks, and critical vendors.\n- Test the full declaration, acknowledgement, fallback, conference, and status-publishing path weekly.\n- Treat missed acknowledgement, failed routing, and unowned incidents as control failures requiring review.", "dependencies": ["S7", "S8", "S11"]}, {"step_id": "S15", "title": "Standardize internal communications", "description": "Give responders one working room and stakeholders one controlled source of truth. Executives must not interrupt the technical command path.\n\n- Maintain one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.\n- Issue the initial internal notice within 10 minutes for SEV1 and 15 minutes for SEV2.\n- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.\n- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.\n- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within the same initial-notice window.\n- Route executive questions through the Executive Duty Officer or Communications Lead.\n- Record every material decision and outbound message in the incident timeline.\n- Require the executive team to sign the communication behavior rules.", "dependencies": ["S7", "S10", "S11"]}, {"step_id": "S16", "title": "Standardize customer, account, and regulatory communications", "description": "Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate observed impact before root cause is known.\n\n- Publish customer status within 15 minutes of a customer-visible SEV1 and within 30 minutes of a customer-visible SEV2.\n- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.\n- Publish a monitoring notice after mitigation and a resolution notice within 30 minutes of verified recovery.\n- Map public status components to customer capabilities rather than internal service names.\n- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, integrity investigation, security restriction, monitoring, and resolution.\n- State observed impact, affected capabilities, available workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.\n- Have account managers contact affected strategic accounts within 30 minutes for SEV1 and 60 minutes for SEV2 using the approved briefing.\n- Offer status subscriptions to all customers. Auto-enroll only where contracts, consent, and applicable communication rules permit.\n- Encode customer-specific notice deadlines and channels in the customer record.\n- Have Legal and Compliance maintain a counsel-validated matrix covering applicable NYDFS, breach, GLBA or FTC, PCI, money-transmitter, sponsor-bank, network, insurance, and customer obligations.\n- Complete and record a reportability assessment within one hour for every SEV1 and every security, privacy, or integrity-related SEV2, including decisions of not reportable.\n- Let Legal own regulatory text and submission while the Incident Commander owns operational facts. Record any legally required restriction of public detail and its alternative stakeholder plan.", "dependencies": ["S3", "S6", "S15"]}, {"step_id": "S17", "title": "Tie incidents to SLA credits and financial exposure", "description": "Link the incident record to contractual and financial outcomes. Finance should not discover outages through later credit claims.\n\n- Define the authoritative availability calculation for each contract and customer journey with Legal and Finance.\n- Calculate affected customers and minutes from incident scope and journey telemetry.\n- Produce a preliminary credit and contractual-exposure estimate within five business days of resolution.\n- Record failed payment count, failed or delayed value, settlement exposure, reconciliation breaks, support effort, and engineering effort.\n- Establish a documented approval path for proactive credits and claims-based credits.\n- Attribute credits and financial harm to recurring failure families.\n- Use the quarterly credit analysis to prioritize detection, resilience, and architectural investment.", "dependencies": ["S3", "S16"]}, {"step_id": "S18", "title": "Establish mandatory blameless postmortems", "description": "Use one learning standard with fixed deadlines. Keep learning separate from disciplinary and misconduct processes.\n\n- Require a postmortem for every SEV1 and SEV2.\n- Also require one for customer-first detection, impact lasting more than two hours, SLA credits, contractual breach, repeated contributing factors, major control failures, and ledger-integrity near misses.\n- Produce a factual draft within three business days, hold the review within five, and publish the approved version within 10.\n- Make the owning engineering director accountable for completion. The commander owns response analysis, and the scribe supplies the timeline.\n- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.\n- Require explicit answers to why detection was not earlier and why mitigation took as long as it did.\n- Use a trained facilitator who was not the commander or primary technical responder.\n- Describe decisions using the context and information available at the time. Do not name an individual as the root cause.\n- Keep HR, misconduct, and personnel matters in separate processes.\n- Publish broadly useful findings internally while restricting security, privacy, personnel, and privileged content appropriately.", "dependencies": ["S6", "S7"]}, {"step_id": "S19", "title": "Make corrective actions enforceable risk commitments", "description": "An action is not complete when its ticket is closed. It is complete when the intended risk reduction is verified.\n\n- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.\n- Use default classes: containment within 7 days, corrective work within 30 days, and strategic work within 90 days with milestones.\n- Prefer actions that remove hazards or reduce blast radius over vague actions such as retraining or adding monitoring.\n- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.\n- Reserve 10%–15% of engineering capacity for approved reliability work.\n- Escalate overdue high-risk actions to the manager after 7 days, director after 14, and CTO after 30.\n- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.\n- Permit unrepaired severe conditions to block related releases.\n- Verify effectiveness using tests, telemetry, exercises, or production evidence before closure.\n- Triage all 53 historical open actions within 30 days as complete, re-planned, superseded with evidence, or formally risk-accepted.", "dependencies": ["S18"]}, {"step_id": "S20", "title": "Set the service-readiness bar and payments playbooks", "description": "A critical service must be supportable at 3 a.m. before it enters direct overnight coverage. Existing critical detection remains active while readiness gaps are repaired.\n\n- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and integrity constraints.\n- Require responders to demonstrate access and safe execution before independent primary duty.\n- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional impairment, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.\n- Define stop-processing and read-only modes for credible financial-integrity risk.\n- Document split-brain prevention, replay protection, failover, controlled backlog recovery, and post-recovery reconciliation.\n- Require executive-approved RTO and RPO for the shared ledger cluster.\n- Exercise critical runbooks at least twice a year and after material changes.\n- Block new Tier 0 or Tier 1 releases and paging rules when readiness requirements are missing.\n- Handle existing gaps using named owners, compensating controls, executive-approved expiry dates, and remediation plans.\n- Run a parallel architecture workstream to reduce ledger blast radius, isolate non-critical readers, and strengthen regional independence.", "dependencies": ["S5", "S10", "S12", "S13"]}, {"step_id": "S21", "title": "Train and certify every response role", "description": "Command and communication are learned skills. Use paid working time and certify people before independent duty.\n\n- Give all employees a 30-minute module on recognizing impact, declaring incidents, and locating the status page.\n- Give responders a half-day module on severity, acknowledgement, escalation, evidence, runbooks, and financial-integrity precautions.\n- Train scribes for two hours on timeline quality and fact-versus-hypothesis labeling.\n- Train Incident Commanders for two days on delegation, uncertainty, severity, mitigation strategy, fatigue, handoffs, and executive management.\n- Require commander candidates to complete a simulated SEV1 and two shadowed incidents or exercises.\n- Train Communications Leads for one day on status writing, customer segmentation, legal boundaries, and contractual clocks.\n- Require domain responders to demonstrate dashboards, access, rollback, failover, escalation, and relevant playbooks.\n- Require two shadow shifts before independent primary duty.\n- Renew certification annually through simulation.\n- Maintain the training, assessment, and certification register as operational and audit evidence.\n- Nominate an incident-management champion in each of the 28 teams.", "dependencies": ["S7", "S14", "S15", "S18", "S20"]}, {"step_id": "S22", "title": "Exercise command, recovery, and tool failure", "description": "Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.\n\n- Run the first cross-company command and communications tabletop within 30 days.\n- Run monthly domain tabletops using scenarios from the 31 historical incidents.\n- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous operational and security events.\n- Test loss of chat, identity, paging, conference, and status-page providers.\n- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors when relevant.\n- Conduct at least one unannounced after-hours paging test before the audit and two annually thereafter.\n- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.\n- Measure acknowledgement, command, customer notice, mitigation decision, handoff, and recovery times.\n- Create tracked corrective actions for every material exercise finding.", "dependencies": ["S11", "S20", "S21"]}, {"step_id": "S23", "title": "Publish the signed policy and control set", "description": "Convert the design into concise documents that people can use during an incident. The actual operating process must also be the documented and audited process.\n\n- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, evidence, and exception standards.\n- Include one-page cards for severity, roles, authority, escalation, and communication timings inside the incident tool.\n- State explicitly that responders support only owned or formally accepted and trained service portfolios.\n- Include the compensation structure, fatigue rules, declaration rights, and non-retaliation commitment.\n- Obtain approval from the CTO, HR, Legal, Security, Compliance, and Internal Audit.\n- Announce the policy at an all-hands and through team briefings.\n- Create an exception register with owner, rationale, compensating control, approver, review date, and expiry.\n- Version every policy change. Do not rewrite historical records when the process changes.", "dependencies": ["S6", "S7", "S8", "S9", "S10", "S12", "S14", "S15", "S16", "S18", "S19"]}, {"step_id": "S24", "title": "Instrument the scorecard and review forums", "description": "Measure the process before the pilot so failures become visible immediately. Report medians and 90th percentiles rather than averages alone.\n\n- Measure impact-to-detection, detection-to-declaration, declaration-to-command, acknowledgement, mitigation, recovery, and resolution.\n- Split results by severity, service tier, customer journey, region, detection source, and business-hours status.\n- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.\n- Track notification episodes, actionability, duplicates, after-hours load, missed detection, routing errors, and page-budget breaches.\n- Track postmortem timeliness, action age, due-date performance, verified effectiveness, and repeated contributing factors.\n- Track journey availability, error-budget burn, failed or delayed value, reconciliation breaks, and SLA credits.\n- Track rotation size, duty frequency, overnight work, recovery days, exceptions, sentiment, and responder attrition.\n- Hold a weekly Incident Review Board chaired by the Head of Reliability with relevant directors.\n- Hold a monthly executive reliability review and a quarterly control and resilience review with Internal Audit.\n- Reconcile incident records monthly against support cases, customer complaints, status history, credits, and major operational anomalies to detect under-reporting.\n- Use team-level scorecards to direct help and investment. Never penalize an individual for good-faith declaration.", "dependencies": ["S3", "S11", "S18", "S19"]}, {"step_id": "S25", "title": "Pilot the complete process on the payment path", "description": "Run a four-to-six-week pilot across the highest-risk journey before expanding. The interim operating floor remains active for the rest of the company.\n\n- Include payment orchestration, ledger application, PostgreSQL platform, API edge, authentication, settlement, reconciliation, Kubernetes platform, and Support intake.\n- Include teams with existing on-call experience and teams new to the model.\n- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking together.\n- Run legacy and new paging paths in parallel for no more than one week, then make the new platform authoritative.\n- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.\n- Have the program team coach incidents without silently taking command.\n- Correct critical process or tooling defects within 48 hours.\n- Exit only after 95% timely command assignment, 95% communications compliance, no unpaid pages, complete required postmortems, tested fallbacks, and at least 50% lower pilot noise.\n- Publish the pilot results, defects, and policy changes company-wide.", "dependencies": ["S9", "S11", "S13", "S20", "S21", "S23", "S24"]}, {"step_id": "S26", "title": "Roll out by risk with readiness gates", "description": "Expand in controlled waves and finish early enough to accumulate operating evidence before the audit. A calendar date does not override a failed readiness gate.\n\n- Roll out remaining Tier 0 domains first, followed by Tier 1, Tier 2, and Tier 3.\n- Use four waves of six to eight teams, each lasting two to three weeks.\n- Gate each service on catalog ownership, appropriate coverage, active compensation, trained responders, access, tested escalation, alert quality, runbooks, and a passed tabletop.\n- Require at least six responders only for direct 24x7 technical rotations. Use business-hours coverage for lower tiers.\n- Give each wave a named coach and director sign-off.\n- Reschedule failed gates or use a time-limited executive exception with compensating controls. Do not create silent waivers.\n- Disable legacy human-paging routes after verified cutover for each wave.\n- Run a quota-based noise reduction sprint in every wave, starting with the highest-volume rules.\n- Pair every suppression with a compensating-detection check.\n- Publish an internal adoption dashboard by team, service tier, coverage, and control gap.\n- Complete critical coverage by approximately week 12 and all 28 teams by week 18.", "dependencies": ["S25"]}, {"step_id": "S27", "title": "Prove SOC 2 operating effectiveness", "description": "Generate evidence through normal operation rather than reconstructing it before fieldwork. Test both control design and consistent execution.\n\n- Map controls to the applicable Trust Services Criteria with Compliance and the auditor, including monitoring, incident identification, response, recovery, communications, and availability.\n- Retain approved policies, exceptions, service ownership, schedules, compensation activation, access reviews, training, incidents, communications, reportability decisions, postmortems, actions, and exercises.\n- Sample incidents monthly from first signal through verified corrective action.\n- Include customer-reported events, downgraded incidents, missed timelines, non-reportable decisions, and exercises in the testing population.\n- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.\n- Run a formal mock audit in month 7 using the evidence populations and interviews expected from the external auditor.\n- Correct deviations through tracked actions with owners and dates. Never edit history to create apparent compliance.\n- Verify that evidence retention covers the full auditor-defined observation period.\n- Brief commanders, engineers, Support, and Compliance on the actual process without scripting inaccurate answers.", "dependencies": ["S22", "S23", "S24", "S26"]}, {"step_id": "S28", "title": "Inspect, adapt, and institutionalize ownership", "description": "Prevent the process from decaying after rollout or the audit. Change it using measured operating evidence rather than opinion.\n\n- Review the policy after 90 days of live operation using severity calibration, page load, missed detection, communication compliance, action closure, fatigue, and survey results.\n- Remove steps that create work without reducing risk. Add controls only where incidents, exercises, or evidence show a gap.\n- Reassess Tier 0 and Tier 1 classification and domain boundaries every six months.\n- Assign permanent owners for policy, catalog, paging, status page, training, metrics, evidence, and exercise scheduling.\n- Review compensation bands, rotation burden, accommodations, and staffing annually.\n- Report severe incidents, credits, overdue high-risk actions, and ledger concentration risk to the board or risk committee quarterly.\n- Maintain the ledger blast-radius program as an executive risk until failover, degraded mode, reconciliation, and regional independence meet approved objectives.\n- Evaluate follow-the-sun coverage using one year of actual activation and staffing data.\n- Build year-two plans for automated mitigation, safer deployments, graceful degradation, and error-budget release controls.", "dependencies": ["S26", "S27"]}], "estimated_complexity": "high", "success_metrics": "- By day 7, every suspected major incident uses one record, one coordination channel, and a named Incident Commander within 10 minutes.\n- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes.\n- By month 3, at least 95% of SEV1 and SEV2 incidents have a named commander within 5 minutes.\n- By month 3, at least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes.\n- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.\n- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and appropriate technical coverage.\n- By week 18, all 180 services have an owner, tier, escalation path, and tested coverage model.\n- No new mandatory night rotation begins before compensation, training, access, runbooks, and minimum staffing are active.\n- Every direct 24x7 technical rotation has at least six qualified responders or an approved, expiring executive exception.\n- No responder is routinely primary more often than one week in six or assigned to two simultaneous primary rotations.\n- Median impact-to-detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.\n- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.\n- By day 90, incident replay identifies a current detector for at least 90% of the 31 historical incidents, including its expected detection minute.\n- Median time to mitigation falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.\n- At least 95% of applicable SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.\n- At least 95% of incidents meet the required internal and customer update cadence by month 3.\n- Monthly human notification episodes fall from the current 3,400 alert events to no more than 1,500 by day 90 and 500 by month 6.\n- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without reducing Tier 0 or Tier 1 replay coverage.\n- Average after-hours load remains at or below two notification episodes per responder per week; every sustained breach receives a dated remediation plan.\n- All six monitoring sources route human pages through the controlled paging platform by week 12, with direct legacy routes retired by week 18.\n- All required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 from month 3 onward.\n- All 53 historical open actions are triaged within 30 days.\n- At least 90% of high-priority corrective actions are completed by their approved due dates with effectiveness evidence by month 6.\n- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.\n- Reportability is assessed and recorded within 1 hour for 100% of SEV1 and qualifying SEV2 incidents, including not-reportable decisions.\n- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.\n- Approved ledger RTO and RPO, controlled failover, read-only mode, and post-recovery reconciliation are exercised before month 6.\n- Quarterly responder surveys reach at least 75% favorable responses for fairness, ownership boundaries, compensation, and sustainability by month 6.\n- Monthly contracted availability meets or exceeds 99.95% by month 6 using the contractually authoritative measurement method.\n- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.\n- The month-7 mock audit finds no unowned high-risk control gap and at least 95% of sampled incidents contain complete operating evidence.\n- The SOC 2 Type II incident-response controls complete external testing without an unresolved material exception."}A competent merge: it absorbed P4's audit clock and staffing math, P1's replay test and P2's accommodations, and split the readiness bar out of the execution doctrine. It adds nothing the other plans do not already have, and it created two new mega-steps.
- Step 1 now confirms the SOC 2 observation window in week 1 so the day-7 floor counts as evidence.
- Step 9 publishes the staffing arithmetic (260 engineers cannot sustain 28 rotas) and bans two-person 24x7 rotations, making the coverage design arguable rather than asserted.
- Step 12 adds the replay test and the two-region fake-redundancy validation; step 24 reports replay coverage.
- Step 4 adds a non-punitive accommodation process, the platform-does-not-inherit rule, and the hard gate that no new mandatory night rotation starts before contract, compensation and training are live.
- Step 15 pulls the service-readiness bar, playbooks and ledger blast-radius reduction out of the execution doctrine into their own step with board-visible milestones.
- Step 14 now merges the escalation ladder and the live execution doctrine into 12 bullets, undercutting its own claim to be "one short operating procedure".
- Step 16 merges internal comms, status page, account-manager outreach, the regulatory obligation matrix and the contact matrix into one step; round 2 had these as three separate steps (16, 17, 18).
- Dropped the round-2 option to evaluate a first-line overnight triage desk, leaving Layer C nights covered only by manager escalation lists.
- Metrics still claim "SOC 2 Type II incident-response controls pass with zero exceptions"; replay coverage appears in step 24 but not in the success-metric list.
- Proposal 4 : Confirm the SOC 2 Type II observation period with the auditor in week 1.
- Proposal 4 : Staffing math justifying 10–12 domains and the prohibition on two-person 24x7 rotations.
- Proposal 1 : The replay test against the 31 historical incidents.
- Proposal 2 : Confidential, non-punitive accommodations and the rule that platform on-call triages but does not inherit another team's service.
- Proposal 4 : No new mandatory night rotation starts until the fairness contract, compensation and training are live.
- Proposal 4 : Integrity/security flags attached to any severity level.
- Proposal 1 : A paid overnight Duty Triage Desk owning the first ten minutes of a page.
- Proposal 2 : Complete company rollout by week 18 with control tests in months 4 and 6.
+ Map compliance, evidence, and audit requirements from day one+ Adopt the severity scale, declaration rights, and incident lifecycle+ Codify the five-minute escalation path and live execution doctrine+ Set the service-readiness bar and write major-incident playbooks+ Enforce action ownership, reserved capacity, and tracking+ Publish Incident Management Policy v1 and the exception register+ Run change management, fairness, and pager culture from day one+ Maintain the program risk register and pre-committed contingenciesLock severity levels and what each one triggersWrite the five-minute path from signal to commandCodify live execution, readiness, and ledger playbooksTrack actions as risk commitments with reserved capacityPublish Policy v1 and the signed fairness contract
The plan produced
1. Charter the program, fund it, and start the audit clock from P4 step 1
Convert the CEO email into a company operating process with one accountable owner, a budget, and a dated timeline that starts this week.
- Name the CTO as executive sponsor and appoint a Head of Reliability & Incident Management as the single accountable owner with full-time authority over all 28 teams.
- Stand up a three-person program office: program lead, platform engineer, reliability analyst.
- Form an eight-person steering group spanning Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, and HR. It proposes; the sponsor decides within 48 hours. Never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one human-paging path, one incident record, one postmortem format, paid on-call, mandatory action tracking, and named service ownership.
- Approve budget anchored against the $1.3M in SLA credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs, and 15% reserved engineering capacity.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Publish a one-page charter company-wide on day 2. Incident response is a company process, not a per-team preference.
- Timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day interim command floor (after 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place this week so the next outage has a named commander and the evidence clock starts immediately.
- Publish a one-page interim severity card and one declaration path: a Slack command, a phone number, and the existing pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs from the 12 teams already on-call. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical action items into complete, re-plan, or formally risk-accept within 30 days, prioritising ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline of incidents, alerts, and money lost (after 1)
Rebuild the facts before locking any design. This is both the design input and the frozen 'before' picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines in one signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the written fairness contract (after 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written, signed deal.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- State explicitly that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Map compliance, evidence, and audit requirements from day one (after 1)
Design evidence as a by-product of operations, not a reconstruction before the auditor arrives. The interim process in week 1 is already evidence.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for internal and external communication; CC5 for control activities; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure with verification, training register, drill records, reportability decisions including 'not reportable'.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records. Require role-based access, MFA, and periodic access review.
- Version and approve all policy documents from day one: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard, Alert Quality Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (after 3)
You cannot route a page correctly across 180 services until every service has a named owner. A wrong owner recreates the pager objection. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Assign one accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. Tier 0: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for every Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform. They fail differently and need different hands.
- Orphan services get an owner within 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Missing ownership or runbooks blocks Tier 0/1 releases.
7. Adopt the severity scale, declaration rights, and incident lifecycle (after 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix. Anyone may declare. Nobody is penalised for over-declaring.
- SEV1 (crisis): money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; deploy freeze; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 (critical): material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 (contained): narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads; business-hours comms; postmortem if customer-detected, over two hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages.
- Payments-specific anchors: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident become at least SEV2.
- Only the Incident Commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
8. Define incident roles, authority, dual-control, and handover discipline (after 7)
Solve 'nobody in charge for an hour' by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- Incident Commander: owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- Communications Lead: the single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: maintains the timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates at SEV1/SEV2.
- Subject-Matter Responders: engineers of the owning team. They mitigate; they do not run the room.
- Executive Duty Officer (SEV1): removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Command claimed within 5 minutes and stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Staff 24x7 with a three-layer coverage model, not 28 night rotations (after 6, 8) from P4 step 8
Do not create 28 night rotations. That is precisely what engineers are rejecting. Centralise coordination in a trained command corps and keep technical ownership local.
- Layer A — Incident Command corps: approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. Paired with a Communications Lead pool of approximately 18 from Support, CS, and engineering management, and a scribe pool used as the training entry point.
- Layer B — critical-path domain rotations: consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- Layer C — everyone else: business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat a Lisbon or APAC follow-the-sun cell as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
10. Approve paid on-call, New York labor compliance, and fatigue safeguards (after 4, 9) from P1 step 9
Unpaid on-call in New York is both a retention problem and a wage-hour exposure. Pay must be live in payroll before any new mandatory night rotation starts. Publish the actual numbers, then ask people to sign up.
- Indicative scheme locked by HR, Finance, and employment counsel within 14 days: approximately $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as approximately 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
- Publish on-call load by team quarterly.
- Hard gate: no mandatory night rotation starts before its compensation is live in payroll. Interim duty is paid retroactively.
11. Set the alert-quality standard and a hard page budget (after 3, 6) from P4 step 10
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target: 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, noise below 15%, with no loss of Tier 0/1 detection coverage.
12. Detect payment and ledger failures before customers do (after 6, 11) from P4 step 11
The goal is blunt: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
- Run the replay test: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish replay coverage as a leading metric.
13. Consolidate to one paging platform, one incident record, one status page (after 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
14. Codify the five-minute escalation path and live execution doctrine (after 8, 9, 13) from P2 step 13
Write one unskippable path from 'something looks wrong' to 'someone is in charge'. The default action is never waiting. If nobody claims command within 5 minutes, the platform assigns it and announces it.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute.
- Separate diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Closure requires a stability observation window and explicit handback.
15. Set the service-readiness bar and write major-incident playbooks (after 6, 9, 12) from P2 step 19
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human at night. The shared ledger cluster is the single largest structural risk.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction: partitioning, read replicas, isolation of non-critical readers, with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
16. Run one communications clock for internals, customers, and regulators (after 7, 8, 13) from P4 step 15
Replace 'whoever is around' with one timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure. State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is 'not reportable'.
- Obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
17. Tie incidents to SLA credits and true financial cost (after 7, 16) from P4 step 16
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
18. Make blameless postmortems mandatory with one format and fixed deadlines (after 7, 8) from P4 step 17
Replace 'some incidents, various formats' with one mandatory format, fixed deadlines, and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did?
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and are never the commander of the incident under review.
- Publish a searchable library and a quarterly top-five recurring causes analysis, with restricted versions for security or privileged content.
19. Enforce action ownership, reserved capacity, and tracking (after 13, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity. Closing a ticket without evidence does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
- Target: 90% of high-priority actions closed by due date within two quarters.
20. Train and certify every role before independent duty (after 8, 14, 16, 18) from P4 step 19
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it. The certification register is an audit artefact.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days plus shadowing): command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
21. Rehearse with tabletops, game days, and unannounced drills (after 13, 15, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-mitigation.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
22. Publish Incident Management Policy v1 and the exception register (after 7, 8, 9, 10, 11, 14, 16, 17, 18, 19) from P1 step 24
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (after 10, 12, 13, 15, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse.
- Publish a one-page result to the whole company.
24. Instrument metrics, dashboards, review cadence, and anti-gaming (after 13, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
25. Roll out in risk-ordered waves with readiness gates (after 23, 24) from P4 step 24
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, never waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave. Give every team its ranked list of noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Review missed detections monthly with the same seriousness as noise.
- Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
26. Run change management, fairness, and pager culture from day one (after 4, 10)
Run this in parallel from day one. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
27. Produce SOC 2 evidence by operating, test internally, and mock-audit (after 22, 24, 25) from P4 step 25
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence. Documented exceptions beat a claim of perfection.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5, CC2.2/CC2.3, CC5, and A1.2. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including 'not reportable', postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in months 4 and 6, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
28. Maintain the program risk register and pre-committed contingencies (after 1)
Name the ways this programme fails and pre-commit the response. Review monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
29. Inspect at 90 days, lock year-two ownership, and prevent decay (after 24, 25, 27) from P4 step 26
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
- A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of 'not reportable'.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: d55e2704-ef8a-4e9c-8fd5-e6d956728a8d, Agent: opus5_refine_1, LLM: anthropic/claude-opus-5):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (32):
1. Charter the program: one owner, one mandate, funded, dated before the audit
Turn the CEO email into a chartered company program with a single accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the **burning-platform narrative** and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the written on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (depends on: 5, 7)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 8, 4)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (depends on: 10, 5)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (depends on: 7, 12, 13)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
18. SLA credit and financial impact workflow (depends on: 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Use the credit delta as the standing business case for on-call pay and reserved reliability capacity.
19. Blameless postmortem standard and Incident Review Board (depends on: 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, reserved capacity and enforcement (depends on: 19, 12)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
21. Runbooks, readiness bar and ledger blast-radius reduction (depends on: 5, 8)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (depends on: 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (depends on: 22, 12, 21)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
24. Publish Incident Management Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 11, 12, 21, 22, 9)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Alert noise burn-down campaign (depends on: 10, 12, 25)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
28. Wave rollout to all 28 teams with readiness gates (depends on: 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
29. Change management, fairness and pager culture (depends on: 4, 9, 25)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (depends on: 24, 26, 28)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 26, 28, 30)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 2 (ID: ddb335b2-98d9-46f8-afa0-88fca2e25749, Agent: gpt5.6-sol_refine_2, LLM: openai/gpt-5.6-sol):
Estimated Complexity: high
Success Metrics: - By day 7, every suspected major incident uses one incident record, one coordination channel, and a named commander.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes; at least 95% receive command within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and technical coverage.
- By week 18, all 180 services have an owner, tier, tested escalation path, and coverage appropriate to their risk.
- No mandatory night rotation starts before compensation, training, access, runbooks, and staffing controls are active.
- Every direct 24x7 technical rotation has at least six qualified responders or a documented, expiring executive exception.
- No responder is routinely assigned primary duty more often than one week in six or simultaneously assigned to two primary rotations.
- At least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of qualifying SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet their required internal and customer update cadence by month 3.
- Monthly human pages fall from 3,400 alert events to no more than 1,500 by day 90 and no more than 500 actionable pages by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, with no decline in Tier 0/1 detection coverage.
- Average after-hours load remains at or below two pages per responder per week; every sustained breach produces a remediation plan.
- All six monitoring sources route human pages through the single paging platform by week 12; direct legacy paging paths are disabled by week 18.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Monthly contracted availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- Quarterly on-call surveys reach at least 75% favorable responses on fairness, ownership boundaries, compensation, and sustainability by month 6.
- The month-7 mock audit finds no unowned high-risk control gap, and at least 95% of sampled incidents contain complete operating evidence.
Steps (26):
1. Establish the mandate, owner, funding, and schedule
Launch incident management as a **company operating program within 48 hours**. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Form a decision group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Approve the non-negotiables: one severity scale, one human-paging platform, one incident record, paid on-call, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, and incident actions.
- Fund tooling, compensation, training, exercises, and reliability work. Use the $1.3M in credits as the minimum financial comparison.
- Set milestones: interim process by day 7, policy and critical-path pilot by week 6, Tier 0/1 coverage by week 10, company rollout by week 18, internal audit in month 6, and mock audit in month 7.
2. Install a seven-day incident-response floor (depends on: 1)
Do not wait for policy design or tooling migration. Put a minimum process into operation immediately and begin retaining evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communication schedule.
- Create one monitored declaration path through chat, telephone, and the current paging environment.
- Use one incident channel, bridge, timeline document, and naming convention for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced on-call engineers and engineering leaders.
- Require a named commander within 10 minutes. The duty engineering director assumes command if nobody else does.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Compensate interim duty retroactively under the permanent policy.
- Hold a 15-minute daily operational review until the permanent process is live.
3. Create the factual, legal, and audit baseline (depends on: 1)
Build one defensible baseline for process design, executive decisions, and SOC 2 testing. Confirm the auditor's expected Type II observation period immediately.
- Reconstruct all 31 incidents from first customer impact through resolution, communications, credits, postmortem, and corrective action.
- Analyze the two incidents with unclear command minute by minute.
- Identify the missing detection signal for every customer-first incident.
- Inventory the six alert sources, 3,400 monthly alert events, noisy rules, duplicates, missing owners, and missing runbooks.
- Record the current rotations, unpaid work, after-hours load, and teams without coverage.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M credits, 85% noise, and 11 of 64 actions closed.
- Map incident-response controls to applicable SOC 2 criteria with Compliance and the auditor.
- Establish retention, confidentiality, legal-hold, and access requirements for incident evidence.
4. Co-design the fairness contract with engineers (depends on: 1)
Treat pager resistance as a valid design constraint. Make the employment and ownership bargain explicit before expanding on-call.
- Interview representatives from all 28 teams and the Support, Customer Success, Security, and Operations groups.
- Distinguish objections involving unpaid work, unfamiliar code, bad alerts, weak runbooks, sleep disruption, or blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for services they own or have formally accepted, supported by trained command, and given capacity to remove recurring defects.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Measure baseline trust, fairness, fatigue, and psychological safety. Repeat the survey at days 60 and 120, then quarterly.
5. Build the service catalog and criticality model (depends on: 3)
Make a machine-readable catalog the source of truth for routing, escalation, customer impact, and audit evidence. Every production service must have one accountable owner.
- Record the owning team, manager, product capability, escalation policy, communication channel, dashboard, runbook, dependencies, regions, and data stores for all 180 services.
- Map payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, and reporting to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing services whose failure can create material contractual impact.
- Classify Tier 2 as internal or deferrable services, and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional mode, recovery method, and contractual obligations for Tier 0 and Tier 1.
- Give orphan services an owner or approved decommission date within 30 days.
- Maintain separate but coordinated ownership for the ledger application and PostgreSQL platform.
6. Adopt one severity scale and incident lifecycle (depends on: 3, 5)
Use four impact-based levels. Classify on actual or credible customer, financial, security, regulatory, and contractual harm rather than organizational seniority.
- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; both regions impaired; or a core payment journey broadly unavailable. Page every role immediately, open a bridge, notify executives, publish customer status within 15 minutes when applicable, begin legal assessment, and require a postmortem.
- **SEV2 — major:** material payment degradation; settlement deadline at risk; a critical customer or material cohort unavailable; regional impairment with reduced resilience; or likely SLA breach. Page command and technical roles, publish customer status within 30 minutes when customer-visible, and require a postmortem.
- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads; page only when immediate action is necessary.
- **SEV4 — operational event:** no current customer impact and no urgent risk. Create a ticket and handle during normal operations.
- Treat percentages as supporting guardrails, never reasons to under-classify integrity, settlement, security, or contractual risk.
- Anyone may declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Automatically use at least SEV2 posture for suspected ledger-integrity events, cross-team incidents with unknown ownership, and materially unknown impact lasting 15 minutes.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed. Resolution requires stability, backlog recovery, and necessary reconciliation.
7. Define roles, authority, and handoffs (depends on: 6)
Separate command, communications, recordkeeping, and technical repair. One named person must hold command throughout every SEV1 and SEV2.
- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not serve as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- **Subject-Matter Responders:** diagnose and mitigate only services for which they have ownership, access, training, or formally accepted support responsibility.
- **Executive Duty Officer:** removes organizational barriers and makes exceptional business decisions without displacing the commander.
- Security, Legal, Compliance, Vendor Management, and Finance join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Require distinct commander, communications, scribe, and primary technical lead for SEV1. Communications and scribe may combine temporarily for bounded SEV2 incidents.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, and reconciliation requirements for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, with the exact transfer time.
8. Create sustainable 24x7 coverage across 28 teams (depends on: 4, 5, 7)
Use central command coverage and risk-based technical coverage rather than creating 28 fragile night rotations. All services receive a response path, but only critical domains maintain direct overnight technical rotations.
- Create a 24x7 Incident Command corps of 24–30 certified people, with primary and secondary coverage at all times.
- Create 18–24 trained Communications Leads and a similarly sized scribe pool using Support, Customer Operations, Engineering Operations, and qualified managers.
- Maintain a surge roster for simultaneous incidents and a 24x7 Executive Duty Officer schedule.
- Group Tier 0 and Tier 1 ownership into approximately 8–12 coherent domains only where responders have training, access, runbooks, and explicit acceptance.
- Staff each critical domain with primary and secondary responders and at least six qualified people. Target no more than one primary week in six.
- Give Tier 2 and Tier 3 services business-hours coverage plus a maintained manager and director escalation path.
- When lower-tier impact becomes SEV1 or SEV2, central command activates the manager escalation path and obtains the necessary owner.
- Route unknown-owner pages to Duty Command and platform triage temporarily. Log each one as a catalog control failure.
- Do not merge small teams merely for schedule convenience. Provide training, reassign services, add staffing, or decommission unsupported services.
9. Approve compensation and fatigue protections (depends on: 4, 8)
End unpaid on-call before expanding mandatory night coverage. HR, Finance, Payroll, and employment counsel should approve the policy within 14 days.
- Use market-validated weekly bands, initially budgeting approximately $800–$1,200 for Tier 0/1 primary duty and $300–$500 for secondary duty.
- Budget approximately $900–$1,300 for Duty Incident Commander weeks and $400–$800 for Communications Lead or scribe duty, adjusted for actual burden.
- Pay holiday premiums and compensate active after-hours work according to exempt or non-exempt status and applicable federal and New York rules.
- Provide a protected paid recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal sprint commitments by approximately 15% during a primary on-call week.
- Prohibit simultaneous primary assignments, consecutive primary weeks, on-call during leave, and invisible schedule swaps.
- Allow responders to declare themselves temporarily unfit after disruptive night work without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Budget roughly $0.9M–$1.2M annually, then refine using actual rotation count, employment classification, and activation data.
10. Establish one paging and incident system of record (depends on: 3, 5, 8)
Monitoring tools may remain specialized, but all human pages must enter one controlled platform. This removes conflicting schedules and creates one evidence trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page.
- Ingest events from the six existing monitoring tools before disabling their direct paging paths.
- Route pages through service-catalog ownership and deduplicate related events.
- Provide a single declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, severity changes, decisions, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status-page publishing, and conference facilities.
- Require role-based access, MFA, access reviews, and immutable or tamper-evident history.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, or an AWS region.
- Retire each legacy human-paging path only after ownership review, end-to-end tests, and two weeks of verified operation.
11. Enforce an alert-quality contract and page budget (depends on: 3, 5, 10)
Treat every page as a production interface with an owner and an expected action. Noise reduction must not create detection gaps.
- Require every paging rule to identify the service, owning team, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Page only when prompt human judgment or intervention can materially reduce customer, financial, security, or contractual risk.
- Route informational, capacity, and non-urgent infrastructure conditions to dashboards or ticket queues.
- Prefer payment-outcome, settlement-risk, queue-age, and error-budget-burn alerts over raw CPU, memory, pod, or log thresholds.
- Run new paging rules in shadow mode for at least seven days unless an emergency exception is approved.
- Review rules with repeated no-action acknowledgements, low actionability, or excessive firing within two business days.
- Set a load budget of no more than two after-hours pages per responder per week, measured over four weeks.
- Require an owner, compensating detection, and recorded approval before suppressing or deleting a rule.
- Burn down the top 50 noisy rules first. Review missed detections and noise together so teams cannot improve metrics by becoming blind.
12. Detect payment and ledger failures before customers (depends on: 5, 11)
Move detection from host health to customer journeys and financial outcomes. Use internal SLOs with enough headroom to protect the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, API access, webhooks, and reporting freshness.
- Run external synthetic transactions through critical payment journeys at least every minute and from paths independent of the production platform.
- Validate each AWS region and expose dependencies that undermine nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed value relative to settlement deadlines, not only queue depth.
- Add tenant and cohort anomaly detection for high-value customers and major payment methods.
- Convert high-priority Support, account-manager, processor, sponsor-bank, and network reports into incident candidates within five minutes.
- Record the detection source for every incident. Treat customer-first detection as a mandatory missed-detection review.
13. Codify detection, escalation, and live execution (depends on: 7, 10)
Create one time-bound path from the first credible signal to named command and mitigation. Notification delivery does not count as human acknowledgement.
- Page the owning critical-domain primary and Duty Incident Commander immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at five minutes, manager at 10, and director at 15.
- Escalate an unclaimed command page to backup commander at five minutes. The Executive Duty Officer assumes command at 10 minutes until a certified handoff occurs.
- Require Support and account managers to use the same declaration path as automated monitoring and engineers.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Begin with a standard statement of severity, known impact, assigned roles, current objective, workstreams, and next update time.
- Freeze unrelated changes during SEV1 and normally during SEV2. Record any exception.
- Prefer reversible mitigation such as rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require formal command handoff for long incidents, shift changes, or fatigue. Do not leave an incident unowned during transfer.
- Reconcile the ledger and safely drain backlogs before resolving payment or ledger incidents.
14. Standardize internal incident communications (depends on: 7, 10, 13)
Give responders one working room and stakeholders one controlled information source. Executives must not interrupt the technical command path.
- Create one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue an initial internal brief within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within 15 minutes for SEV1 and 30 minutes for SEV2.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record material decisions and outbound messages in the incident timeline.
15. Standardize status-page and customer communications (depends on: 6, 10, 14)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate impact before root cause is known.
- Publish customer status within 15 minutes of declaring a customer-visible SEV1 and within 30 minutes for customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice promptly after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map status-page components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, data-integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, known workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Give affected strategic accounts direct account-manager outreach within 30 minutes for SEV1 and 60 minutes for SEV2.
- Require account managers to use the approved briefing and prohibit independent technical explanations.
- Provide a customer-facing incident summary within five business days for SEV1 and qualifying SEV2 events.
- Record any legally necessary delay or restriction of public detail, its approver, and the alternative communication plan.
16. Operationalize legal, regulatory, contractual, and credit decisions (depends on: 3, 6, 15)
Some payments incidents start external notification clocks. Make assessment mandatory without assuming that every operational incident is reportable.
- Build a counsel-validated matrix covering applicable NYDFS rules, state breach laws, GLBA or FTC obligations, PCI requirements, money-transmitter obligations, sponsor-bank and network contracts, cyber insurance, and customer contracts.
- Engage Legal and Compliance immediately for every SEV1, security event, and suspected ledger-integrity event.
- Record an initial reportability assessment within one hour for SEV1 and within two hours for other potentially reportable events.
- Record non-reportable decisions as evidence, including facts considered, approver, timestamp, and reassessment trigger.
- Maintain tested 24x7 contacts for counsel, regulators, sponsor banks, networks, insurers, and critical vendors.
- Encode customer-specific notice deadlines and channels in the customer record used by the Communications Lead.
- Have Legal own regulatory text and submission. Keep technical command with the Incident Commander.
- Have Finance calculate affected minutes, delayed value, likely credits, and contractual exposure within five business days.
- Track credits by incident and recurring cause to support reliability investment decisions.
17. Make postmortems mandatory, consistent, and blameless (depends on: 6, 7, 10)
Use one learning standard with fixed deadlines. Keep postmortems separate from performance, misconduct, and disciplinary processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact over two hours, SLA credits, contractual breach, repeated contributing factors, control failures, and ledger-integrity near misses.
- Produce the factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the Incident Commander responsible for the timeline and the owning engineering director accountable for completion.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit analysis of why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the incident commander.
- Describe decisions in the context and information available at the time. Do not name an individual as the root cause.
- Publish broadly useful findings internally while restricting security, privacy, personnel, or privileged material appropriately.
18. Give corrective actions enforceable ownership (depends on: 10, 17)
Treat incident actions as risk commitments, not suggestions. A ticket is not complete until the expected risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Escalate overdue high-risk actions to the manager after seven days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit overdue actions to block related releases where the unrepaired condition could reproduce severe impact.
- Verify effectiveness through tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days. Complete, re-plan, or formally accept each risk.
19. Set the service-readiness bar and critical playbooks (depends on: 5, 8, 11, 12)
A critical service must be supportable at 3 a.m. before its team is placed on direct overnight coverage. Existing critical detection must remain active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and data-integrity constraints.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional failure, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document controlled failover, split-brain prevention, replay protection, backlog recovery, and post-recovery reconciliation.
- Exercise critical runbooks at least twice per year and after material changes.
- Block new Tier 0/1 releases and new paging rules when readiness requirements are missing.
- Handle existing gaps through a named owner, compensating control, executive-approved expiry date, and remediation plan rather than disabling detection.
20. Publish policy and certify every response role (depends on: 6, 7, 8, 9, 11, 13, 14, 15, 16, 17, 18)
Convert the operating design into concise, signed documents and practical training. Training and exercises occur during paid working time.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, and exception standards.
- Provide one-page severity, role, authority, escalation, and communication cards inside the incident tool.
- Train all employees to recognize and declare incidents.
- Train all engineers in severity, acknowledgement, evidence preservation, handoff, and financial-integrity precautions.
- Certify responders only after they demonstrate access, dashboards, runbooks, rollback, and escalation competence.
- Certify Incident Commanders through formal instruction, simulation, and at least two shadowed incidents or exercises.
- Train Communications Leads in status writing, account segmentation, contractual clocks, and legal escalation.
- Train scribes in timeline quality and fact-versus-hypothesis labeling.
- Require two shadow shifts before independent primary duty and annual recertification.
- Maintain the training and certification register as operational and audit evidence.
21. Pilot on the payment critical path (depends on: 9, 10, 12, 19, 20)
Run a four-to-six-week pilot across the highest-risk journey before expanding. Use real incidents and exercises to correct the model quickly.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, API edge, authentication, settlement, and Support intake.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently assuming command.
- Correct critical process or tool defects within 48 hours and update the standard visibly.
- Exit only after 95% timely command assignment, 95% communication compliance, no unpaid pages, complete required postmortems, and at least 50% lower noise.
22. Roll out by customer journey and risk (depends on: 21)
Expand in controlled waves while the interim process remains active company-wide. Complete rollout early enough to accumulate operating evidence before the audit.
- Roll out remaining Tier 0 domains first, then Tier 1, Tier 2, and Tier 3.
- Use two-to-three-week waves with a named program coach and director sign-off.
- Gate each team on catalog ownership, appropriate coverage, compensation, trained responders, tested escalation, alert quality, runbooks, access, and a passed tabletop.
- Require six or more responders only for direct 24x7 technical rotations. Apply business-hours coverage and manager escalation to lower tiers.
- Reschedule failed gates or approve a time-limited executive exception with a compensating control.
- Disable legacy human-paging paths after verified cutover for each wave.
- Publish an internal adoption dashboard by team, service tier, and control gap.
- Finish critical coverage by week 10 and all 28 teams by week 18.
23. Exercise command, communications, regional recovery, and fallbacks (depends on: 19, 20)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous security and operational events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors where appropriate.
- Conduct at least one unannounced after-hours paging test before the audit.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure time to acknowledgement, command, customer notice, mitigation decision, handoff, and recovery.
- Create tracked actions for every material exercise finding.
24. Measure performance and operate fixed review forums (depends on: 3, 10, 17, 18)
Use a balanced scorecard that exposes weak controls without rewarding hidden incidents or suppressed alerts. Report medians and 90th percentiles rather than averages alone.
- Measure time from first impact to detection, declaration, acknowledgement, command assignment, mitigation, recovery, and resolution.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track pages, actionability, duplicates, after-hours load, missed detections, and page-budget breaches.
- Track postmortem timeliness, action age, closure by due date, verified effectiveness, and recurring contributing factors.
- Track availability by customer journey, error-budget burn, failed or delayed payment value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight activations, recovery days, schedule exceptions, sentiment, and attrition.
- Hold a weekly Incident Review Board for incidents, postmortems, control failures, noisy alerts, and overdue actions.
- Hold a monthly executive reliability review for trends, funding, contractual exposure, and accepted risks.
- Hold a quarterly control and resilience review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile incident records monthly against support cases, status history, credits, and customer complaints to detect under-reporting.
25. Prove SOC 2 operating effectiveness before fieldwork (depends on: 20, 22, 23, 24)
Generate audit evidence through normal operation rather than reconstructing it later. Test both control design and consistent execution.
- Maintain approved and versioned policies, exceptions, catalog records, schedules, compensation activation, access reviews, training, incidents, communications, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, non-reportable legal decisions, missed timelines, and exercises in the population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the populations, evidence requests, and interviews expected from the external auditor.
- Record deviations honestly with owners, remediation dates, and compensating controls. Never rewrite historical records.
- Confirm that evidence retention covers the full auditor-defined observation period.
- Brief commanders, responders, Support, and Compliance on the real process without scripting inaccurate answers.
26. Institutionalize improvement and reduce structural risk (depends on: 22, 24, 25)
Prevent the program from decaying after the audit. Use incident evidence to drive permanent ownership and architectural investment.
- Assign permanent owners for policy, service catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review severity thresholds, staffing, compensation, and communication timings annually and after material process failures.
- Recertify command and communications personnel annually.
- Review recurring failure families quarterly and require executive decisions where corrective work repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fairness, fatigue, and tool friction.
- Report severe incidents, credits, overdue high-risk actions, and resilience investment to the board or risk committee quarterly.
- Treat the shared ledger cluster as a strategic concentration risk. Fund failover assurance, blast-radius reduction, isolation of non-critical readers, and stronger regional independence.
- Evaluate follow-the-sun command or technical coverage using one year of page-load and staffing data.
- Build year-two plans for automated mitigation, deployment safety, graceful degradation, and error-budget release controls.
Previous Proposal 3 (ID: f7ee763b-3b76-4f07-9e7d-5164701bef1c, Agent: qwen3.8-max_refine_3, LLM: alibaba/qwen3.8-max):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (32):
1. Establish executive mandate, program office, and funding
Turn the CEO email into a **chartered company program** with one accountable owner and authority over all 28 teams.
- Appoint the CTO as executive sponsor and a Head of Reliability as program owner with full-time authority.
- Create a permanent program office: one program lead, one platform engineer, one analyst.
- Form an eight-person steering group: Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours.
- Lock the non-negotiables: one severity scale, one paging platform, one postmortem format, paid on-call, mandatory action tracking, named service ownership.
- Approve budget anchored against the $1.3M in credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs.
- Reserve 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Publish a one-page charter stating incident response is a company operating process, not a per-team choice.
2. Install a seven-day interim command bridge (depends on: 1)
Do not leave the company unprotected while the permanent process is designed. Put a **crude but real** command structure in place within one week.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Direct Support to escalate credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical actions: complete, re-plan, or formally risk-accept, prioritizing ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Build the forensic baseline of incidents, alerts, and money lost (depends on: 1)
Rebuild the facts before designing anything. This becomes both the **design input** and the before picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific missing signal. This list drives the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Freeze baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the on-call fairness deal (depends on: 1)
Engineer pushback is the **largest delivery risk**. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS, and Sales in two weeks. Separate the real objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey covering fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
5. Map compliance, evidence, and audit requirements from day one (depends on: 1)
Design evidence as a **by-product of operations**, not a reconstruction before the auditor arrives. Confirm the SOC 2 observation window immediately.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for communications; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure, training register, drill records, reportability decisions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve all policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has a **named owner**. Build a machine-readable catalog as the single source of truth.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact: Tier 0 (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for each Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
7. Adopt the severity scale, declaration rules, and lifecycle (depends on: 3, 6)
Replace judgment calls with a **lookup table**. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter.
- SEV1 (crisis): money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of all roles and executive; bridge in 5 minutes; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 (critical): material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- SEV3 (major, contained): narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected, Declared, Triaged, Mitigated (customer impact ends), Monitoring, Resolved (backlog processed and ledger reconciled), Reviewed.
- Publish a decision tree with 12 worked examples from the real 31 incidents.
8. Define incident roles, authority, and handover discipline (depends on: 7)
Solve nobody-in-charge-for-an-hour by making command **explicit, single-holder, transferable, and logged**. Separate coordination from debugging.
- Incident Commander: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates for SEV1/SEV2.
- Subject-Matter Responders: engineers of the owning team; they mitigate, they do not run the room.
- Executive Duty Officer (SEV1): removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Design the three-layer 24x7 coverage model (depends on: 6, 8)
Do not create 28 night rotations. **Centralise coordination** in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- Layer A, Incident Command corps: approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of approximately 18 from Support, CS, and engineering management. A scribe pool used as the training entry point.
- Layer B, critical-path domain rotations: consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- Layer C, everyone else: business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing: a US-only paid night rotation now, a follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
10. Approve on-call compensation, labor compliance, and fatigue safeguards (depends on: 9, 4)
Unpaid on-call in New York is both a **retention problem and a legal exposure**. Pay must be live in payroll before any mandatory night rotation starts.
- Indicative scheme: approximately $1,000 per primary 24x7 week, $400 secondary, $250 for business-hours rotations, a separate $1,200 Duty Commander stipend, holiday premiums, approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- HR, Finance, and employment counsel publish amounts, eligibility, tax treatment, and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a 15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
11. Set the alert quality standard, page budget, and noise burn-down (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes. Make alert quality a **condition of being allowed to page a human**.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly.
- Target 3,400 to under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
12. Build payment-outcome detection and ledger assurance (depends on: 11, 6)
Stop customers telling you first. Detection must be driven by **payment outcomes and ledger truth**, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer detected first becomes a named defect class with a mandatory tracked action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Consolidate to one paging platform, one incident record, one status page (depends on: 7, 9, 11)
Collapse six alerting tools into a **single operational system of record** so there is one queue, one timeline, and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is unavailable. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
14. Codify the escalation ladder and five-minute command rule (depends on: 8, 9, 13)
Write one unskippable path from something looks wrong to **someone is in charge**. The default action is never waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately, secondary at 5 minutes, domain manager and Duty Commander at 10, Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension, and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
15. Write the live incident execution doctrine and major-incident playbooks (depends on: 8, 13, 14)
Give responders one short operating procedure from the first minutes through closure. Priority is **limiting customer and financial harm**, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the single largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window, an explicit handback to the owning team, Support, and Customer Success, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
16. Standardize internal communications protocol (depends on: 8, 13)
Standardise the internal picture so executives, Support, and Sales are informed **without interrupting** the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
17. Build customer communications, status page, and account-manager outreach (depends on: 7, 16)
Replace whoever is around with a **timed, owned, pre-approved process**. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook, and RSS. All 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity, or blame.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
18. Create the regulatory, partner, and legal notification playbook (depends on: 7, 17)
In payments, some incidents start a **legal clock at detection**. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded even when the answer is not reportable, with evidence, decision-maker, and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
19. Operationalize SLA credit and financial impact workflow (depends on: 7, 17)
Tie incidents to money so severity, credits, and investment decisions **stay honest**, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one. Use the delta as the standing business case for on-call pay and reliability capacity.
20. Establish the blameless postmortem standard and Incident Review Board (depends on: 7, 8)
Replace some incidents, various formats with **one mandatory format, fixed deadlines, and a forum with teeth**.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly top-five recurring causes analysis.
21. Enforce action ownership, capacity reservation, and tracking (depends on: 20, 13)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of **customer commitments** and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
22. Build the training and certification academy (depends on: 8, 14, 16, 20)
Command is a skill, not a title. **Certify before assigning duty**, and use paid working time for all of it.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days + shadowing): command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Comms Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum; never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Run the exercise programme: tabletops, game days, and unannounced drills (depends on: 22, 13, 15)
The process must meet a **simulated SEV1** before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger. Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (depends on: 7, 8, 9, 10, 11, 14, 16, 17, 18, 20, 21)
Collapse the design into a document people will **actually open mid-outage**, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 12, 13, 15, 22, 10)
Prove the process on the **highest-risk surface** with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment.
- Publish a one-page result to the whole company.
26. Run the alert noise burn-down campaign (depends on: 11, 13, 25)
Run noise reduction as a **visible, quota-driven campaign** in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness. Review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 to 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
27. Establish metrics, dashboards, review cadence, and anti-gaming (depends on: 13, 20, 25)
Instrument the process itself so improvement is visible and the auditor sees **evidence of monitoring and review**. Report median and 90th percentile, never averages alone.
- Response: time to detect, declare, commander, acknowledgement, mitigate, resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
28. Wave rollout to all 28 teams with readiness gates (depends on: 25, 27)
Roll out in four waves of six to eight teams every two to three weeks, ordered by **customer risk**. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. Failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Publish a live adoption scoreboard.
29. Run change management, fairness, and pager culture programme (depends on: 4, 10, 25)
Run this from day one in parallel. Engineers judge the process on **fairness**; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. Build SOC 2 evidence by design, internal testing, and mock audit (depends on: 24, 27, 28)
Design the evidence as a **by-product of doing the work**. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any change as an exception.
31. Maintain the program risk register and contingencies (depends on: 1)
Name the ways this programme fails and **pre-commit the response**. Review it monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 27, 28, 30)
Guard against the classic failure: the process **decays once the audit is signed**. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 4 (ID: 101b9773-a36a-47bf-8af6-14fe9a48abb2, Agent: grok4.6_refine_4, LLM: xai/grok-4.6):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (26):
1. Charter the program and start the audit clock
Convert the CEO email into a company operating process with one owner, a budget, and an observation window that starts this week.
Incident response is no longer a per-team choice.
- Name the CTO as sponsor and a Head of Reliability as the single accountable owner, with a three-person program office.
- Form a small decision group: Engineering, SRE, Support, Customer Success, Security, Legal, Finance, and HR. The sponsor decides within 48 hours.
- Lock non-negotiables: one severity scale, one paging path, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, and reserved engineering capacity against the $1.3M in SLA credits.
- Timeline: operating floor in 7 days, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, mock audit in month 7, audit in month 8.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for new tools or the final policy. Put a crude but real process in place this week so the next outage has a named commander.
This week is the first audit evidence.
- Publish a one-page interim severity card and one declaration path: Slack command, phone number, and existing pagers, all reaching the same duty person.
- Staff interim primary and backup Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in the incident channel.
- Tell Support to declare from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline (depends on: 1)
Rebuild the facts before locking design. This is both the design input and the before picture for the CEO and the auditor.
Freeze the numbers so they cannot drift during design.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisy rules and every rule with no owner or runbook.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11 of 64 actions closed, on-call in 12 of 28 teams.
4. Map resistance and publish the fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
Publish the deal in writing before any new mandatory pager is assigned.
- Interview all 28 teams plus Support, Customer Success, and Sales in two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the **fairness contract**: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and money-path tiers (depends on: 3)
You cannot page the right person across 180 services until every service has one owner. A wrong owner recreates the pager objection.
The catalog is the single source of truth for paging, impact, status-page components, and audit.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to services, data stores, regions, and third parties, including sponsor banks and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
6. Lock severity levels and what each one triggers (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by who reported it or how hard the fix looks.
Anyone may declare. Nobody is punished for over-declaring. Only the Incident Commander may downgrade, with the evidence recorded.
- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 major, contained: narrow or single-customer impact with a workaround. Owning team leads; business-hours comms; postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages.
- Attach a financial-integrity or security flag to any severity. The flag forces dual-control, Legal, and the regulatory checkpoint without inventing a fifth level.
- Auto-escalate: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer impact ends) → Monitoring → Resolved (backlog processed and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged.
Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. A human validates for SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1 and SEV2. Every handover is announced verbally and in writing with the exact time.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
8. Staff 24x7 with a command corps, not 28 night rotations (depends on: 5, 7)
Do not create 28 night rotations. That is what engineers are rejecting.
Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified volunteers from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to seven months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate 28 teams into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat Lisbon or APAC follow-the-sun as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
9. Pay on-call, meet New York labor rules, and cap fatigue (depends on: 4, 8)
Unpaid on-call in New York is a retention problem and a legal exposure. Pay must be in payroll before any new mandatory night rotation starts.
Publish the numbers. Then ask people to sign up.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
10. Set the alert-quality bar and a hard page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them.
Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers do (depends on: 5, 10)
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail.
Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One Slack command creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is down. Test weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Write the five-minute path from signal to command (depends on: 7, 8, 12)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting.
If nobody claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Codify live execution, readiness, and ledger playbooks (depends on: 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving a root cause.
A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down. No command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers.
15. Run one communications clock for internals, customers, and regulators (depends on: 6, 7, 12)
Replace whoever is around with one timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is not reportable.
- Obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
16. Tie incidents to SLA credits and true financial cost (depends on: 6, 15)
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
Credit calculation is an output of the incident record, not a negotiation.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
17. Make blameless postmortems mandatory and consistent (depends on: 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth.
The discipline lives in the deadlines and the review, not the template.
- Mandatory for every SEV1 and SEV2, any incident detected by a customer first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed.
18. Track actions as risk commitments with reserved capacity (depends on: 12, 17)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
Closing a ticket without evidence of effectiveness does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Target 90% of high-priority actions closed by due date within two quarters.
19. Train and certify every role before independent duty (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
The certification register is an audit artefact.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
20. Rehearse with tabletops, game days, and night drills (depends on: 12, 14, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
21. Publish Policy v1 and the signed fairness contract (depends on: 6, 7, 8, 9, 10, 13, 15, 17, 18)
Collapse the design into a document people will actually open mid-outage, and make it official.
Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
22. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 19, 21)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse. Publish a one-page result to the whole company.
23. Instrument metrics, reviews, and anti-gaming (depends on: 12, 17, 22)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review.
Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
24. Roll out in risk-ordered waves with readiness gates (depends on: 22, 23)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave, not as a background hope. Give every team its noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
25. Produce SOC 2 evidence by operating, then mock-audit (depends on: 21, 23, 24)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions. Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (depends on: 23, 24, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
Name the ways this programme fails and pre-commit the response.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 5 (ID: 289def1b-8f78-402b-85b3-0d9e18661842, Agent: deepseek-v4-pro_refine_5, LLM: deepseek/deepseek-v4-pro):
Estimated Complexity: high
Success Metrics: - Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, no loss of Tier 0/1 detection.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% Tier 0/1 services have named owning team, tier, escalation policy, dashboard and runbook by day 30; all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10-12 domain rotations of at least six trained responders.
- 30 certified incident commanders and 18 certified communications leads active.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% required postmortems drafted in 3 business days, reviewed in 5, published in 10 from month 2.
- Postmortem action closure rises from 17% to 90% closed by due date within two quarters; all 53 legacy actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k annualised within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% SEV1/security SEV2 including not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two cross-company exercises completed before audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% sampled incidents; SOC 2 Type II incident response controls pass zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days; no increase in attrition among responders.
Steps (26):
1. Executive mandate, budget, and governance
Secure a written CTO/CEO charter making incident management a company operating process, not a team option. Name one accountable Head of Reliability and a small steering group with authority to decide within 48 hours.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking, and protected engineering capacity.
- Approve budget against the $1.3M annual credits: tooling $150-250k/yr, on-call compensation, training, and 2-3 program FTEs.
- Set timeline: interim process in week 1, design weeks 1-6, pilot weeks 7-12, rollout weeks 13-24, mock audit month 7, SOC 2 at month 8.
- Freeze current baselines: 31 incidents, 22-min MTTD, 40% customer-first detection, 3h10 MTTM, $1.3M credits, 3,400 alerts/month at 85% noise, 11/64 actions closed.
2. Interim 7-day command (depends on: 1)
Put a crude but real process in place immediately so no incident remains unowned while the permanent design proceeds.
- Publish a one-page interim severity card, one declaration path (phone, Slack command, pager), and one incident channel/bridge/timeline naming convention.
- Staff an interim 24x7 duty commander with primary and backup from engineering managers and senior SREs; compensate retroactively under final policy.
- Require a named incident commander within 10 minutes of any suspected major incident, announced in channel.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions in 30 days, prioritizing ledger integrity, duplicate payment, regional failover, security and detection gaps.
- Hold a daily 15-minute operations review until the permanent process is live.
3. Forensic baseline and alert estate analysis (depends on: 1)
Reconstruct the true before picture from the last 12 months; it drives design and serves as audit baseline.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact start/detect/declare/commander/mitigate/resolve, who led, credits paid, contributing factors.
- For every customer-first detection, name the missing signal; this becomes the detection backlog.
- Reconstruct the two `nobody in charge` incidents minute by minute; use them as the burning-platform narrative and test case.
- Profile the 3,400 monthly alerts by tool, team, rule, outcome; identify top 50 noisy rules and every rule with no owner or runbook.
- Freeze all baseline metrics in one signed document for the executive and the auditor.
4. Listening tour and fairness contract (depends on: 1, 3)
Treat engineer pushback as the main delivery risk and convert it into a written deal that removes the objection to carrying a pager for other teams' code.
- Interview all 28 teams plus Support, CS and Sales in two weeks to separate objections: unpaid work, nights, unfamiliar code, missing runbooks, or fear of blame.
- Harvest practices from the 12 teams already on-call; they supply pilot teams and first commanders.
- Publish the deal: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Recruit 10-15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey and repeat at 60, 120 and 365 days.
5. Service ownership catalog and criticality tiers (depends on: 3)
Create a machine-readable catalog as the single source of truth for paging, routing, status page components, and audit evidence.
- Assign one accountable team, engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list per service.
- Tier by business impact: Tier 0 money movement/ledger/auth/shared PostgreSQL; Tier 1 customer-facing degradable; Tier 2 internal/batch; Tier 3 non-critical.
- Map customer journeys (initiate, authorize, settle, reconcile, report, onboard) to services, data stores, regions, third parties.
- Give orphan services an owner within 30 days or a decommission date approved by the sponsor; Tier 0 without an owner is an executive escalation.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
6. Severity scale, declaration rules, and automatic triggers (depends on: 3, 5)
Adopt one severity scale as a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade with recorded rationale. Start high when uncertain.
- SEV1: money moved wrongly/duplicated/lost, ledger integrity in doubt, confirmed security/data exposure, both regions impaired, payment processing halted. Full role page, bridge in 5 min, status page in 15 min, regulator assessment within 1 hour, mandatory postmortem.
- SEV2: material degradation of a payment journey, settlement window at risk, large/strategic customer fully down, SLA breach likely. Commander and SMEs paged, status page in 30 min, mandatory postmortem.
- SEV3: narrow or single-customer impact with workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any ledger cluster incident, any SEV3 open >2h, any unknown impact after 15 min, any cross-team incident become at least SEV2.
- Publish a decision tree with 12 worked examples from the 31 real incidents.
7. Incident roles, authority, and handoffs (depends on: 6)
Codify roles so coordination never depends on seniority or heroics. The Incident Commander coordinates; responders fix only services they own.
- IC: owns severity, priorities, roles, cadence, closure; pre-authorized to freeze deploys, rollback, disable features, shift traffic, invoke failover, put ledger read-only, commit spend; does not type in terminals; keeps command when VP joins.
- Communications Lead: sole author for status page, account managers, executives, and hand-off to Legal for regulators; speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, state changes; human validates for SEV1/SEV2.
- Subject-Matter Responders: diagnose and mitigate only their owned services.
- Executive Duty Officer (SEV1): removes obstacles, shields IC from exec questions, owns board/regulator escalation; does not take command unless formally transferred.
- Rule: IC claimed within 5 min and announced; distinct people for IC, Comms, technical lead at SEV1/SEV2; every handover announced verbally and in writing with time.
- Financial controls survive: IC coordinates ledger recovery but cannot bypass dual control, reconciliation or privileged access.
8. 24x7 command and responder coverage model (depends on: 5, 7)
Do not create 28 fragile night rotations. Centralize coordination in a trained command corps, keep technical ownership local.
- Layer A: incident command corps of ~30 certified volunteers with primary/secondary 24x7; paired Comms Lead pool ~18 and scribe pool as entry.
- Layer B: consolidate 28 teams into 10-12 critical-path domains (ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes, data/reporting, integrations) each with 24x7 primary+secondary and minimum six trained responders.
- Layer C: all other teams business-hours on-call with written after-hours escalation lists held by managers.
- Platform on-call is safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate follow-the-sun only as a 12-month option, not year-1 dependency.
- Acknowledgement discipline: 5 min at SEV1/SEV2 with automatic failover.
9. Paid on-call and fatigue safeguards (depends on: 8, 4)
Unpaid on-call in New York is a retention and legal risk. Pay must be in payroll before any mandatory night rotation starts.
- Indicative: ~$1,000 per primary 24x7 week, secondary ~$400, business-hours ~$250, duty commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond one hour.
- Mandatory recovery: paid recovery day after >2h overnight work, SEV1, or qualifying SEV2; managers arrange coverage.
- HR/Legal/Finance publish amounts, tax treatment, FLSA/NY wage-hour status within 14 days.
- Load rules: no primary more often than 1 in 6, no consecutive primary/secondary weeks, no primary on two rotations, no on-call during PTO, annual cap.
- If a rotation averages >2 out-of-hours pages per person per week for 4 weeks, trigger staffing or alert-remediation review.
- Count on-call as ~15% delivery load; document exemption path for health/caring without career penalty.
10. Service readiness bar and runbooks (depends on: 5, 8)
A service must earn the right to page a human at 3 a.m. Define a minimum readiness bar and major incident playbooks before any Tier 0/1 service goes live with on-call.
- Require for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switch, escalation contacts, data-loss/latency impact.
- Write playbooks for top failures: shared PostgreSQL failure/corruption, cross-region failover, Kubernetes control-plane loss, processor/sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, duplicate payment.
- Treat ledger cluster as largest structural risk: documented failover, read-only degraded mode, reconciliation after recovery, executive-signed RPO/RTO.
- Start a parallel workstream on blast-radius reduction: tenant/function partitioning, read replicas, isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging at night unless manager accepts a dated exception in writing.
- Runbooks are peer-reviewed, version-controlled, marked stale if not exercised twice a year.
11. Alert quality standard and page budget (depends on: 3, 5)
Make alert quality a condition of paging a human. Burn down noise deliberately rather than by mass silencing.
- Every paging alert must declare owning team, affected service, customer/SLO impact, expected action, dashboard, runbook, dedup key, severity mapping, escalation policy. Anything failing becomes a ticket.
- Page on symptoms of customer harm (SLO burn rate, payment success rate, queue age vs settlement deadlines), not raw CPU/memory.
- Run new alerts in shadow mode 7 days; test firing and recovery.
- Set page budget of 2 out-of-hours pages per person per week; breach blocks new alert creation and triggers tuning sprint.
- Auto-quarantine rules firing >5 times/month without action or >70% no-action acks; return to owner with correction deadline.
- Never silently disable: verify compensating detection, record decision, name owner, review missed detections monthly.
- Target 3,400 to <500 pages/month, actionability >75% within six months, no loss of Tier 0/1 detection.
12. Detection uplift on money path (depends on: 5, 11)
Stop customers telling you first by detecting on payment outcomes and ledger truth.
- Define SLOs and business SLIs per customer journey: initiation success, authorization latency, settlement timeliness, reconciliation break rate, API availability, reporting freshness; set internal targets stricter than 99.95%.
- Run external synthetic end-to-end payments every 60 seconds from both regions, covering all critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate detection, replication lag, failover-readiness, settlement-window countdown.
- Add per-customer anomaly detection for top 100 accounts; five similar support tickets in 10 minutes auto-creates triage incident.
- Route partner and processor notifications into declaration path within 5 minutes.
- Record detection source on every incident; treat `customer detected first` as a named defect class with mandatory action.
13. Consolidate to one paging and incident platform (depends on: 6, 8, 11, 12)
Collapse six alerting tools into one operational system of record: one queue, one timeline, one audit trail, without creating a monitoring gap.
- Select one paging/scheduling platform, one incident record, one hosted status page; time-box selection to 2 weeks.
- Ingest all six sources first, deduplicate/correlate; retire legacy path only after named owners, successful end-to-end test, and 2 weeks verified operation.
- One-command declaration in Slack creates channel/bridge, pages commander, sets severity, opens timeline, starts clock.
- Route every page through service catalog: service label -> owning domain schedule -> escalation policy.
- Capture evidence automatically: declaration, acks, role assignments, severity changes, decisions, comms, mitigated/resolved times, postmortem link; retain 12 months.
- Test out-of-band SMS/phone paging, mobile fallback, offline runbook weekly; ensure works during one AWS region/chat/identity provider failure.
- Set hard date after which pages outside this tool create no on-call obligation.
14. Escalation paths and acknowledgement SLAs (depends on: 7, 8, 13)
Write one unskippable path from signal to named commander in under 5 minutes; default action is never waiting.
- Converge all entry points on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately -> secondary at 5 min -> domain manager + Duty Commander at 10 -> Executive Duty Officer at 15. SEV2 requires command assigned within 15 min.
- If no one claims command in 5 min, platform assigns and announces it; assignee may hand over but not decline.
- Commander may page any domain's on-call directly with 10-min ack obligation; this reciprocity makes single-team ownership viable.
- If ownership unclear after 10 min, commander keeps incident and names temporary owner; missing catalog entry logged as control defect.
- Pre-authorize regional failover, ledger read-only, payment suspension, partner-bank notification so IC never waits for an executive.
15. Standardize internal and customer communications (depends on: 6, 7, 13)
Replace whoever is around with a timed, owned, pre-approved process; the Communications Lead is single author and never writes from scratch under pressure.
- Internal: one working channel/bridge plus one read-only broadcast for executives/Support/Sales; cadence SEV1 15-30 min, SEV2 60 min even if nothing changed; fixed template with impact, action, next update, IC/Comms names.
- Executives ask questions only of Executive Duty Officer; publish this as signed behavior rule.
- Status page: SEV1 within 15 min, SEV2 within 30 min; updates every 30/60 min; resolution notice within 30 min of verified recovery; customer-facing summary within 5 business days for SEV1/qualifying SEV2.
- Pre-approve 12-15 templates with Legal for degradation, settlement delay, API errors, regional failure, data-integrity investigation, security event.
- Top 100 accounts get named account manager call/email within 30 min of SEV1 with briefing pack; all 2,100 subscribed by default.
- Language rules: state impact and next update; never speculate on cause, recovery time, data integrity or blame.
16. Regulatory, partner, and financial impact workflows (depends on: 6, 15)
In payments some incidents start a legal clock at detection. Build obligation assessment into the process and tie incidents to money.
- Legal/Compliance produce obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC, PCI, money-transmitter, sponsor-bank/card-network windows, cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and security-related SEV2 within 1 hour, recorded even when `not reportable`, with decision-maker and evidence.
- Maintain 24x7 contact matrix for regulators, sponsor banks, networks, outside counsel, insurer.
- Agree availability measurement method per contract; compute affected minutes per customer from incident record and journey telemetry; propose credit schedule within 5 business days.
- Attribute credits to root-cause families; track failed payment volume, delayed value, reconciliation breaks; target annual credits from $1.3M to under $400k.
17. Blameless postmortem standard and action tracking (depends on: 6, 7)
Replace various formats with one mandatory format, fixed deadlines, and enforceable actions. Closing a ticket without evidence does not close the action.
- Mandatory for every SEV1/SEV2, customer-first detection, incident >2h, repeat of known cause, credit-generating/contract breach, ledger near-miss.
- Draft within 3 business days, review within 5, publish within 10; IC owns delivery, owning manager accountable.
- One template: summary, customer/financial impact, detection source/gap, timeline, response analysis, contributing conditions, what worked, actions.
- Two mandatory questions: why did a customer see this first? why did mitigation take as long as it did?
- Blameless in writing: systems and context, no individual named as cause, never used in performance reviews; HR handles personnel/misconduct separately.
- Every action gets owner, priority, due date, verification method, ticket auto-created; classes containment 7 days, corrective 30, strategic 90; reserve 15% engineering capacity.
- Escalation: manager at +7 days, director at +14, CTO dashboard at +30; overdue P0 can block related releases.
- Weekly Incident Review Board ratifies severity, challenges quality, monitors actions; target 90% closure on time within two quarters.
18. Train and certify incident roles (depends on: 7, 13, 15)
Command is a skill, not a title. Certify before duty; use paid working time.
- All employees: 30-min module on recognizing impact, declaring, finding channel/status page.
- Responder: half-day on severity, escalation, runbook, financial integrity; mandatory before joining rotation.
- Scribe: 2 hours timeline discipline; entry point.
- Incident Commander: 2 days + shadowing; command presence, delegation, severity calls, running room, handover; certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead: 1 day on status writing, customer tiering, legal boundaries, regulator triggers.
- Domain responders demonstrate dashboard, runbook, rollback, failover, access before primary; two shadow shifts; no primary in first 90 days.
- Certification valid 12 months, renewed by simulation; register is audit artifact.
19. Simulations and game days (depends on: 10, 14, 15, 18)
Rehearse process before real SEV1; exercise tooling failure and ledger scenarios.
- Monthly tabletop per group reusing an incident from baseline, rotating commander.
- Quarterly game day in staging or tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes degradation.
- Twice-yearly unannounced paging drill to measure real overnight ack times.
- Annually one combined operational/security exercise and one regulator-notification exercise with Legal/CISO.
- Exercise status page down, chat/paging provider unavailable.
- Never inject uncontrolled change into production ledger; validate backups/restore/RPO/RTO on replicas.
- Every exercise produces tracked actions; publish time-to-commander, time-to-first-update, time-to-mitigation.
20. Pilot on critical path (depends on: 9, 11, 12, 13, 14, 15, 17, 18, 19)
Prove full process on highest-risk surface with willing teams for six weeks before full rollout.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, Support intake, including two of the 12 already on-call teams.
- Activate severity scale, command corps, single paging tool, page budget, status page policy, mandatory postmortems, paid rotations.
- Parallel old paths for one week then cut over; real incidents use new process only.
- Program lead attends every SEV2+ as coach, never shadow commander; review every pilot page within one business day.
- Exit gate: commander named within 5 min in 95% cases, first status update on time, MTTD under 10 min for pilot services, pages halved, no unpaid page, all required postmortems on time, positive sentiment.
- Publish one-page result company-wide.
21. Metrics, dashboards, and review cadence (depends on: 13, 17)
Instrument the process so improvement is visible and audit sees monitoring/review evidence. Report median and p90, never averages alone.
- Response: time to detect/declare/commander/ack/mitigate/resolve split by severity, tier, journey, region, detection source.
- Quality: customer-first rate, status page timeliness, update cadence, missed escalations, role conflicts, alert actionability, out-of-hours pages per person.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat factors.
- Business: availability per journey, error budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, frequency, recovery days, sentiment, attrition.
- Weekly Incident Review Board; monthly reliability review per team; monthly executive review to CEO; quarterly control review with Security/Compliance/Risk; quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, status page history; team scorecards direct help, never individual penalties.
22. Wave rollout to all 28 teams with readiness gates (depends on: 20, 21)
Roll out in four waves by customer risk, each passing an explicit gate rather than a date. Complete all teams by week 24 to leave ~3 months of evidence before audit fieldwork.
- Wave 1 remaining Tier 0, Wave 2 Tier 1, Wave 3 Tier 2, Wave 4 Tier 3/internal.
- Per-team onboarding kit: catalog entry complete, alerts migrated within budget, runbooks at readiness bar, rotation staffed with 6 trained responders or time-limited exception, one IC candidate nominated, one tabletop passed, payroll active.
- Named coach per wave for three weeks; director signs gate.
- Legacy paging paths disabled per wave, not kept as comfort fallback.
- Publish live adoption scoreboard.
- Any team unable to staff fair rotation gets headcount, service reassignment, or explicit executive risk acceptance.
23. Culture, fairness, and continuous feedback (depends on: 4, 9)
Run this in parallel from day one. Engineers judge fairness; executives judge results.
- Repeat the deal in every forum: paid on-call, paged only for owned services, trained commander, real sprint capacity.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
- Weekly office hours first 8 weeks, Slack support channel 4-hour SLA, per-team champion, weekly newsletter with metrics and bad news.
- Incident response contribution in promotion criteria; quarterly awards for best postmortem and biggest noise reduction; public thanks after SEV1.
- Pulse survey at 60, 120, 365 days; if fairness or load red, pause expansion until fixed.
24. SOC 2 evidence design and internal dry-run (depends on: 13, 17, 21, 22)
Make evidence a by-product of operations; test before external auditor.
- Map with Compliance to Trust Services Criteria CC7.2-CC7.5, CC2.2/CC2.3, A1.2; confirm observation window early.
- Evidence set automated and indexed: versioned policies/exception, rotation schedules, compensation activation, incident records with timestamps/roles, paging/ack logs, status history, reportability decisions, postmortems, action closure with verification, training/certification register, drill records, access reviews.
- Internal Audit tests process in months 4 and 6, sampling incidents end-to-end.
- Run formal mock audit month 7 using same evidence populations; include interviews with IC, engineer, Support, Compliance.
- Correct failures through tracked actions, never by editing history.
- Freeze process wording after month 7; any change logged as exception.
25. Risk register and contingency planning (depends on: 1)
Name likely failure modes and pre-commit responses; review monthly with sponsor.
- Too few commander volunteers: make command rostered duty for managers and staff engineers until pool reaches 30.
- Compensation not approved in time: fallback to time-off-in-lieu plus phased stipend, never launch mandatory night on-call uncompensated.
- Tool migration slips: cut scope on incident record layer, never on paging consolidation.
- Noise pruning hides real failure: demote to ticket first, observe 30 days, keep recovery path, monthly missed-detection review.
- Burnout in experienced teams: weekly load monitoring with page caps.
- Major SEV1 mid-rollout: program lead becomes full-time responder, wave schedule slips one wave, sponsor told same day.
- Shared ledger concentration: if blast-radius workstream slips, escalate to board as accepted risk with dated plan.
26. Inspect and adapt; year-two sustainability (depends on: 22, 24)
Prevent decay after audit by revising on data and assigning permanent owners.
- At 90 days live, revise policy using measurements: severity calibration, Layer B/C membership from page data, uncovered shifts, commander burn, missed updates, action closure, survey results.
- Assign permanent owners for policy, paging platform, status page, service catalog, training, metrics.
- Re-baseline targets every six months; shift from lagging to leading metrics: error budget burn, near-miss rate, drill performance.
- Year-two candidates: follow-the-sun coverage, automated mitigation for top recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, completion of ledger blast-radius reduction.
- Keep annual policy review, certification renewal, exercise calendar, board reporting as permanent commitments.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Charter the program, fund it, and start the audit clock", "description": "Convert the CEO email into a company operating process with one accountable owner, a budget, and a dated timeline that starts this week.\n\n- Name the CTO as executive sponsor and appoint a Head of Reliability & Incident Management as the single accountable owner with full-time authority over all 28 teams.\n- Stand up a three-person program office: program lead, platform engineer, reliability analyst.\n- Form an eight-person steering group spanning Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, and HR. It proposes; the sponsor decides within 48 hours. Never a 28-person committee.\n- Lock the non-negotiables now: one severity scale, one human-paging path, one incident record, one postmortem format, paid on-call, mandatory action tracking, and named service ownership.\n- Approve budget anchored against the $1.3M in SLA credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs, and 15% reserved engineering capacity.\n- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.\n- Publish a one-page charter company-wide on day 2. Incident response is a company process, not a per-team preference.\n- Timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.", "dependencies": []}, {"step_id": "S2", "title": "Install a seven-day interim command floor", "description": "Do not leave the company unprotected for six weeks of design. Put a crude but real process in place this week so the next outage has a named commander and the evidence clock starts immediately.\n\n- Publish a one-page interim severity card and one declaration path: a Slack command, a phone number, and the existing pagers, all reaching the same duty person.\n- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs from the 12 teams already on-call. Compensate retroactively under the final policy.\n- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.\n- Instruct Support and Customer Success to declare from credible customer reports immediately without waiting for engineering confirmation.\n- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.\n- Triage all 53 open historical action items into complete, re-plan, or formally risk-accept within 30 days, prioritising ledger integrity, duplicate payments, regional failover, and detection gaps.\n- Hold a 15-minute daily operations review until the permanent process is live.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Rebuild the forensic baseline of incidents, alerts, and money lost", "description": "Rebuild the facts before locking any design. This is both the design input and the frozen 'before' picture for the executive team and the auditor.\n\n- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, and contributing factors.\n- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the detection backlog.\n- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.\n- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.\n- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.\n- Freeze the baselines in one signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11/64 actions closed, on-call in 12 of 28 teams.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Run the listening tour and publish the written fairness contract", "description": "Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written, signed deal.\n\n- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.\n- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.\n- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.\n- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.\n- State explicitly that platform on-call may triage unknown ownership but does not permanently inherit another team's service.\n- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.\n- Run a baseline sentiment survey on fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.\n- No new mandatory night rotation starts until this contract, compensation, and training are live.", "dependencies": ["S1"]}, {"step_id": "S5", "title": "Map compliance, evidence, and audit requirements from day one", "description": "Design evidence as a by-product of operations, not a reconstruction before the auditor arrives. The interim process in week 1 is already evidence.\n\n- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for internal and external communication; CC5 for control activities; A1.2 for availability.\n- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure with verification, training register, drill records, reportability decisions including 'not reportable'.\n- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records. Require role-based access, MFA, and periodic access review.\n- Version and approve all policy documents from day one: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard, Alert Quality Standard.\n- Record every control exception with an owner, compensating control, approval, and expiry date.\n- Start monthly evidence sampling immediately rather than reconstructing before the audit.", "dependencies": ["S1"]}, {"step_id": "S6", "title": "Build the service ownership catalog and criticality tiers", "description": "You cannot route a page correctly across 180 services until every service has a named owner. A wrong owner recreates the pager objection. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.\n\n- Assign one accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link, and dependency list.\n- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. **Tier 1**: customer-facing but degradable. **Tier 2**: internal or batch. **Tier 3**: non-critical.\n- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, and third parties including sponsor banks and processors.\n- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for every Tier 0/1 service.\n- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform. They fail differently and need different hands.\n- Orphan services get an owner within 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.\n- Missing ownership or runbooks blocks Tier 0/1 releases.", "dependencies": ["S3"]}, {"step_id": "S7", "title": "Adopt the severity scale, declaration rights, and incident lifecycle", "description": "Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix. Anyone may declare. Nobody is penalised for over-declaring.\n\n- **SEV1 (crisis):** money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; deploy freeze; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.\n- **SEV2 (critical):** material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.\n- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads; business-hours comms; postmortem if customer-detected, over two hours, a repeat, or credit-generating.\n- **SEV4:** no customer impact. Ticket only. Never pages.\n- Payments-specific anchors: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.\n- Auto-escalation: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident become at least SEV2.\n- Only the Incident Commander may downgrade, with recorded evidence.\n- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.\n- Publish a decision tree with 12 worked examples taken from the real 31 incidents.", "dependencies": ["S3", "S6"]}, {"step_id": "S8", "title": "Define incident roles, authority, dual-control, and handover discipline", "description": "Solve 'nobody in charge for an hour' by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.\n\n- **Incident Commander:** owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.\n- **Communications Lead:** the single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.\n- **Scribe:** maintains the timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates at SEV1/SEV2.\n- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.\n- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.\n- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.\n- Command claimed within 5 minutes and stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.\n- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.", "dependencies": ["S7"]}, {"step_id": "S9", "title": "Staff 24x7 with a three-layer coverage model, not 28 night rotations", "description": "Do not create 28 night rotations. That is precisely what engineers are rejecting. Centralise coordination in a trained command corps and keep technical ownership local.\n\n- **Layer A — Incident Command corps:** approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. Paired with a Communications Lead pool of approximately 18 from Support, CS, and engineering management, and a scribe pool used as the training entry point.\n- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.\n- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.\n- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.\n- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.\n- Evaluate a US-only paid night rotation now. Treat a Lisbon or APAC follow-the-sun cell as a 12-month option, not a year-one dependency.\n- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.", "dependencies": ["S6", "S8"]}, {"step_id": "S10", "title": "Approve paid on-call, New York labor compliance, and fatigue safeguards", "description": "Unpaid on-call in New York is both a retention problem and a wage-hour exposure. Pay must be live in payroll before any new mandatory night rotation starts. Publish the actual numbers, then ask people to sign up.\n\n- Indicative scheme locked by HR, Finance, and employment counsel within 14 days: approximately $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and approximately $150 per out-of-hours page plus hourly beyond the first hour.\n- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.\n- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.\n- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.\n- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.\n- Count on-call as approximately 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.\n- Publish on-call load by team quarterly.\n- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.", "dependencies": ["S4", "S9"]}, {"step_id": "S11", "title": "Set the alert-quality standard and a hard page budget", "description": "3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.\n\n- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.\n- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.\n- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.\n- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation for that team and triggers a tuning sprint.\n- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.\n- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.\n- Target: 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, noise below 15%, with no loss of Tier 0/1 detection coverage.", "dependencies": ["S3", "S6"]}, {"step_id": "S12", "title": "Detect payment and ledger failures before customers do", "description": "The goal is blunt: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. Every customer-first incident becomes a named defect with a tracked action.\n\n- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.\n- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.\n- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.\n- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.\n- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.\n- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.\n- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.\n- Run the replay test: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish replay coverage as a leading metric.", "dependencies": ["S6", "S11"]}, {"step_id": "S13", "title": "Consolidate to one paging platform, one incident record, one status page", "description": "Collapse six alerting tools into a single operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.\n\n- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.\n- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.\n- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.\n- Route every page through the service catalog: service label → owning domain schedule → escalation policy.\n- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access.\n- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test paging, escalation, status publication, and bridge access weekly.\n- Set a hard date after which a page outside this platform creates no on-call obligation.", "dependencies": ["S7", "S9", "S11"]}, {"step_id": "S14", "title": "Codify the five-minute escalation path and live execution doctrine", "description": "Write one unskippable path from 'something looks wrong' to 'someone is in charge'. The default action is never waiting. If nobody claims command within 5 minutes, the platform assigns it and announces it.\n\n- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.\n- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.\n- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.\n- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.\n- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.\n- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.\n- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.\n- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.\n- Freeze unrelated production changes during SEV1 and SEV2. Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute.\n- Separate diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down. No command by direct message.\n- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.\n- Formal commander handover beyond four hours, with a second shift staffed. Closure requires a stability observation window and explicit handback.", "dependencies": ["S8", "S9", "S13"]}, {"step_id": "S15", "title": "Set the service-readiness bar and write major-incident playbooks", "description": "Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human at night. The shared ledger cluster is the single largest structural risk.\n\n- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.\n- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.\n- Treat the shared ledger cluster as the largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.\n- Open a parallel architecture workstream on blast-radius reduction: partitioning, read replicas, isolation of non-critical readers, with board-visible milestones.\n- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.\n- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.", "dependencies": ["S6", "S9", "S12"]}, {"step_id": "S16", "title": "Run one communications clock for internals, customers, and regulators", "description": "Replace 'whoever is around' with one timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure. State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.\n\n- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.\n- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.\n- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.\n- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.\n- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.\n- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.\n- Component-level status page mapped to customer journeys. Subscribe all 2,100 customers by default.\n- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.\n- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is 'not reportable'.\n- Obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.\n- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.\n- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.", "dependencies": ["S7", "S8", "S13"]}, {"step_id": "S17", "title": "Tie incidents to SLA credits and true financial cost", "description": "Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.\n\n- Agree with Legal and Finance the availability measurement method per contract and per component.\n- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.\n- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.\n- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.\n- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.\n- Target: credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.", "dependencies": ["S7", "S16"]}, {"step_id": "S18", "title": "Make blameless postmortems mandatory with one format and fixed deadlines", "description": "Replace 'some incidents, various formats' with one mandatory format, fixed deadlines, and a forum with teeth. The discipline lives in the deadlines and the review, not the template.\n\n- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.\n- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.\n- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.\n- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did?\n- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.\n- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.\n- Facilitators are trained and are never the commander of the incident under review.\n- Publish a searchable library and a quarterly top-five recurring causes analysis, with restricted versions for security or privileged content.", "dependencies": ["S7", "S8"]}, {"step_id": "S19", "title": "Enforce action ownership, reserved capacity, and tracking", "description": "Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity. Closing a ticket without evidence does not close the action.\n\n- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.\n- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.\n- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.\n- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.\n- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.\n- Triage all 53 currently open historical actions within 30 days. Complete, re-plan, or formally risk-accept.\n- Target: 90% of high-priority actions closed by due date within two quarters.", "dependencies": ["S13", "S18"]}, {"step_id": "S20", "title": "Train and certify every role before independent duty", "description": "Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it. The certification register is an audit artefact.\n\n- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.\n- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.\n- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.\n- Incident Commander (2 days plus shadowing): command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.\n- Communications Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.\n- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.\n- Certification valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.", "dependencies": ["S8", "S14", "S16", "S18"]}, {"step_id": "S21", "title": "Rehearse with tabletops, game days, and unannounced drills", "description": "The process must meet a simulated SEV1 before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.\n\n- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.\n- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.\n- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.\n- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.\n- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.\n- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.\n- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-mitigation.\n- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.", "dependencies": ["S13", "S15", "S20"]}, {"step_id": "S22", "title": "Publish Incident Management Policy v1 and the exception register", "description": "Collapse the design into a document people will actually open mid-outage, and make it official.\n\n- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.\n- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.\n- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.\n- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.\n- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.", "dependencies": ["S7", "S8", "S9", "S10", "S11", "S14", "S16", "S17", "S18", "S19"]}, {"step_id": "S23", "title": "Pilot on the payments critical path", "description": "Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.\n\n- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.\n- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.\n- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.\n- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.\n- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.\n- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse.\n- Publish a one-page result to the whole company.", "dependencies": ["S10", "S12", "S13", "S15", "S20", "S22"]}, {"step_id": "S24", "title": "Instrument metrics, dashboards, review cadence, and anti-gaming", "description": "Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.\n\n- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.\n- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.\n- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.\n- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.\n- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.\n- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.\n- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.\n- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.", "dependencies": ["S13", "S18", "S23"]}, {"step_id": "S25", "title": "Roll out in risk-ordered waves with readiness gates", "description": "Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.\n\n- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.\n- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.\n- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, never waived.\n- Legacy alert paths are disabled per wave, not kept as a comfort fallback.\n- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.\n- Run noise burn-down as a quota inside each wave. Give every team its ranked list of noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.\n- Pair every suppression with a compensating-detection check. Review missed detections monthly with the same seriousness as noise.\n- Publish a live adoption scoreboard.\n- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.", "dependencies": ["S23", "S24"]}, {"step_id": "S26", "title": "Run change management, fairness, and pager culture from day one", "description": "Run this in parallel from day one. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.\n\n- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.\n- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.\n- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.\n- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.\n- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.\n- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.", "dependencies": ["S4", "S10"]}, {"step_id": "S27", "title": "Produce SOC 2 evidence by operating, test internally, and mock-audit", "description": "Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence. Documented exceptions beat a claim of perfection.\n\n- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5, CC2.2/CC2.3, CC5, and A1.2. Confirm the observation window early.\n- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including 'not reportable', postmortems, action closure with verification, training and certification register, drill records, access reviews.\n- Internal Audit or an independent control owner tests design and operation in months 4 and 6, sampling incidents end to end from first signal to verified action closure.\n- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.\n- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.\n- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.", "dependencies": ["S22", "S24", "S25"]}, {"step_id": "S28", "title": "Maintain the program risk register and pre-committed contingencies", "description": "Name the ways this programme fails and pre-commit the response. Review monthly with the sponsor.\n\n- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.\n- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.\n- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.\n- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.\n- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.\n- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.\n- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.", "dependencies": ["S1"]}, {"step_id": "S29", "title": "Inspect at 90 days, lock year-two ownership, and prevent decay", "description": "Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.\n\n- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.\n- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.\n- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.\n- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.\n- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.\n- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.", "dependencies": ["S24", "S25", "S27"]}], "estimated_complexity": "high", "success_metrics": "- A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.\n- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.\n- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.\n- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.\n- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.\n- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.\n- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.\n- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.\n- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.\n- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.\n- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.\n- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.\n- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.\n- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.\n- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.\n- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.\n- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.\n- Monthly availability meets or exceeds 99.95% per customer journey by month 6.\n- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of 'not reportable'.\n- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.\n- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.\n- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.\n- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly."}It imported P1's Duty Triage Desk and replay test and then improved both, and added two things nobody else has: a vendor-incident class and a rule for who writes the status page overnight. It stays the most compact serious plan at 26 steps.
- Step 8 adopts the Duty Triage Desk but fixes its latency risk: the desk never sits on suspected ledger, payment-halt or security events, which page commander and likely domains in parallel; step 18 adds a half-day triage-desk training module.
- Step 8 answers a gap every other plan leaves: SEV1 pages a Communications Lead 24x7, while for customer-visible SEV2 the commander publishes the first status from a template and Comms is paged at 30 minutes or when a top-100 account is affected.
- Step 13 adds a vendor-incident class (processor, sponsor bank, card network, cloud control plane) measured by time-to-customer-notice, time-to-failover-decision and queue management rather than root cause at the vendor.
- Step 3 builds a replay catalog with the gap owner, and step 11 makes closing replay gaps a precondition for claiming detection improved; replay coverage enters the metrics.
- Step 6 widens the flag set to integrity, security, settlement and regulatory, with the explicit worked case that one-customer ledger corruption is still a flagged crisis.
- Step 14 adds "existing detection stays on while gaps are repaired", closing the readiness-bar loophole that could silence detectors.
- Contingencies remain four bullets inside step 26 rather than a standing risk register reviewed monthly with the sponsor, as in P1, P3 and P5.
- Step 23 still bundles the noise burn-down quota with the wave rollout, so noise reduction has no independent owner or schedule.
- Typo in step 11 ("replay test against the S3 catalog") and metrics still claim SOC 2 passes with zero exceptions.
- Proposal 1 : A paid overnight Duty Triage Desk owning the first ten minutes of ambiguous or unowned pages.
- Proposal 1 : Replay the 31 incidents to name the detector that would now fire and at which minute.
- Proposal 1 : A month-5 internal control test before the month-7 mock audit.
- Proposal 2 : Keep existing critical detection active while readiness gaps are repaired.
- Proposal 2 : A pay band for Communications Lead and scribe duty.
- Proposal 2 : Do not merge small teams merely for schedule convenience; reclassify or add staffing instead.
- Proposal 1 : A standing programme risk register reviewed monthly with the sponsor.
- Proposal 2 : Complete all 28 teams by week 18.
+ Lock severity levels, integrity flags, and the incident lifecycle+ Staff 24x7 with command, domains, and a Duty Triage Desk+ Codify escalation, the five-minute command rule, and vendor incidents+ Operationalize regulatory notice, partner clocks, and SLA credits+ Publish Policy v1 and the exception register+ Burn down alert noise and roll out by risk with gates+ Operate the fairness and culture program in parallel+ Prove SOC 2 operating effectiveness before fieldworkMap resistance and publish the fairness contractLock severity levels and what each one triggersStaff 24x7 with a command corps, not 28 night rotationsWrite the five-minute path from signal to commandTie incidents to SLA credits and true financial costTrack actions as risk commitments with reserved capacityRoll out in risk-ordered waves with readiness gatesProduce SOC 2 evidence by operating, then mock-audit
The plan produced
1. Charter the program, fund it, and start the audit clock
Turn the CEO email into a chartered company program within 48 hours. Incident response becomes a company operating process, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability as full-time owner, with a three-person office: program lead, platform engineer, and analyst.
- Form a small decision group of Engineering, SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit. It proposes. The sponsor decides within 48 hours.
- Lock the non-negotiables now: one severity scale, one human-paging platform, one incident record, one postmortem format, mandatory action tracking, paid on-call, and named service ownership.
- Confirm the SOC 2 Type II observation window with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, exercises, and reserved engineering capacity against the $1.3M in SLA credits.
- Reserve 15% of engineering capacity for detection, runbooks, and incident actions, protected by the sponsor.
- Publish a one-page charter on day 2. Clock: floor by day 7, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day operating floor (after 1)
Do not wait for policy, tooling, or the audit. Put a crude but real process in place this week so the next outage already has an owner.
- Publish a one-page interim severity card and one declaration path: chat command, phone number, and current pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline, and one naming convention for every suspected major incident.
- Triage the 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
- Replay one of the two nobody-in-charge incidents as a tabletop within 14 days using this floor process.
3. Rebuild the forensic baseline and incident-replay catalog (after 1)
Rebuild the facts before locking design. This is the design input, the frozen before-picture for the CEO, and the test set for detection work.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact missing signal. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Name the top 50 noisy rules and every rule with no owner or runbook.
- Quantify true cost: credits, failed payment volume, delayed value, reconciliation breaks, and engineering hours lost.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11 of 64 actions closed, on-call in 12 of 28 teams.
- Build a replay catalog: for each historical incident, the detector that should now fire, at which minute, and the owner of the gap.
4. Publish the on-call fairness contract (after 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders. Do not punish them by making them the company's permanent night watch.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the fairness contract in writing: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive exemption path for health, disability, or caregiving.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and journey tiers (after 3)
You cannot route a page correctly across 180 services until every service has one owner. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Record owning team, manager, chat channel, escalation policy, dashboards, runbook, dependencies, regions, and data stores for all 180 services.
- Tier by business impact, not technology. Tier 0: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, refund, report, onboard — to services, data stores, regions, sponsor banks, and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform. They fail differently.
- Orphan services get an owner in 30 days or an approved decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Coverage follows the journey, not the org chart. Small teams that own Tier 0 pieces get headcount, service reassignment, or membership in a domain rotation. Never a two-person 24x7 rota.
6. Lock severity levels, integrity flags, and the incident lifecycle (after 3, 5) from P2 step 6
Replace judgement calls with a lookup table. Classify on actual or credible customer, financial, security, and contractual harm, never on who reported it or how hard the fix looks.
- SEV1 crisis: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 contained: narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages a human.
- Attach an integrity, security, settlement, or regulatory flag to any severity. The flag forces dual-control, Legal, and the reportability checkpoint without inventing a fifth level. A one-customer ledger corruption is still a flagged crisis.
- Auto-escalate to at least SEV2: any ledger-cluster event, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident.
- Anyone may declare. Nobody is punished for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (after 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- Incident Commander: owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- Communications Lead: single voice for the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists. A human validates at SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory text. The commander owns the facts.
- Distinct people for command, comms, and technical lead at SEV1. Comms and scribe may combine only for bounded SEV2.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
- Every role assignment and handover is announced verbally and in writing with the exact time and open risks.
8. Staff 24x7 with command, domains, and a Duty Triage Desk (after 4, 5, 7) new
Do not create 28 night rotations. That is what engineers are rejecting. Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified people from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to eight months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate ownership into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- Duty Triage Desk: a small paid overnight first-line rotation that owns the first ten minutes of ambiguous, unowned, or low-confidence pages. It verifies, enriches, and applies only the runbook's safe steps, then wakes the owning domain. It never sits on a suspected ledger, payment-halt, or security event. Those page commander and likely domains immediately, in parallel.
- Overnight comms: SEV1 pages a Communications Lead 24x7. Customer-visible SEV2 lets the commander publish the first status from a template; Comms is paged if the incident is still open at 30 minutes or a top-100 account is affected.
- Staffing math: 260 engineers can sustain 10–12 domain rotations, one command corps, and one triage desk. They cannot sustain 28 night rotas. Role exclusivity: nobody is primary on two rotations in the same week. Commanders may also be domain responders in different weeks.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
- Evaluate follow-the-sun as a 12-month option, not a year-one dependency. Seed Layer B from the 12 teams already on-call.
9. Pay on-call, meet New York labor rules, and cap fatigue (after 4, 8)
Unpaid on-call in New York is a retention problem and a wage-hour exposure. Pay must be in payroll before any new mandatory night rotation starts.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, $400–$800 for Comms or triage-desk duty, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- Hard gate: no mandatory night rotation starts before compensation is live in payroll. Interim duty is paid retroactively. Budget roughly $0.9M–$1.2M a year, then refine with actual rotation count.
10. Enforce an alert-quality contract and page budget (after 3, 5) from P2 step 11
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery, unless an emergency exception is recorded.
- Set a page budget of two out-of-hours pages per responder per week. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers, then replay the past (after 5, 10)
Stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. A detector is not done until it would have caught the last 12 months.
- Define SLIs and SLOs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner, processor, and sponsor-bank notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a named defect class with a mandatory tracked action.
- Run the replay test against the S3 catalog. For each of the 31 incidents, name the detector that would now fire and at which minute. Close gaps the replay exposes before calling detection improved.
12. Consolidate to one pager, one incident record, one status page (after 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One chat command creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy. Unowned pages go to the Duty Triage Desk and Duty Command, and log a catalog defect.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Codify escalation, the five-minute command rule, and vendor incidents (after 7, 8, 12) from P3 step 14
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting. Payments also fail at processors and banks, which you cannot patch.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Vendor-incident class: processor, sponsor-bank, card-network, or cloud-control-plane failure. Command is still required. Success is time-to-customer-notice, time-to-failover-decision, and queue management, not root cause at the vendor.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Write live execution doctrine, readiness bars, and ledger playbooks (after 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving root cause. A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis. Split diagnosis and mitigation workstreams once staffing allows.
- Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Drain backlogs under control. Complete reconciliation before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date. Existing detection stays on while gaps are repaired.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers, with board-visible milestones and executive-signed RPO/RTO.
15. Run one communications clock for internals, customers, and account managers (after 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for customer-visible SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys, not internal service names. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer record.
- Language rules: state impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
16. Operationalize regulatory notice, partner clocks, and SLA credits (after 6, 15) from P2 step 16
In payments, some incidents start a legal clock at detection. Tie incidents to money so Finance does not learn about outages from invoices.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded even when not reportable, with facts, decision-maker, and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families. Track failed payment volume, delayed value, and reconciliation breaks as the true incident cost.
17. Make postmortems mandatory and actions enforceable (after 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth. Eleven of 64 closed is a process nobody enforces.
- Mandatory for every SEV1 and SEV2, any incident a customer detected first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and gap, timeline, response analysis, contributing conditions, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items. Facilitators are trained and are never the commander of the incident under review.
- Every action gets one named owner, priority, due date, expected risk reduction, verification method, and a ticket created automatically. If it is not on the board, it does not exist.
- Classes: containment in 7 days, corrective in 30, strategic in 90. SEV1 recurrence-prevention enters the next sprint ahead of roadmap work. Reserve 15% capacity.
- Overdue ladder: manager at +7 days, director at +14, CTO at +30. Overdue high-risk items need written residual-risk acceptance and can block related releases. Verify effectiveness before closing.
18. Train and certify every response role before independent duty (after 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Duty Triage Desk, half day: enrichment, safe-step limits, when to wake a domain immediately, when not to delay.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
19. Publish Policy v1 and the exception register (after 6, 7, 8, 9, 10, 13, 15, 16, 17) from P1 step 24
Collapse the design into a document people will actually open mid-outage, and make it official. Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
20. Pilot on the payments critical path (after 9, 11, 12, 14, 18, 19)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, Duty Triage Desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
21. Instrument metrics, reviews, and anti-gaming (after 12, 17, 20)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- Response: time from first impact to detect, declare, commander, acknowledge, mitigate, and resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
22. Rehearse with tabletops, game days, and night drills (after 12, 14, 18, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced overnight paging drill to measure real acknowledgement times.
- One combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO, before the audit.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
23. Burn down alert noise and roll out by risk with gates (after 20, 21) new
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Noise reduction is a quota inside each wave, not a background hope.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy paging paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Give every team its noisiest rules from the baseline. Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed. Pair every suppression with a compensating-detection check.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
- If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
24. Operate the fairness and culture program in parallel (after 4, 9) from P3 step 29
Run this from the moment the deal is published. Engineers judge the process on fairness. Executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
25. Prove SOC 2 operating effectiveness before fieldwork (after 19, 21, 22, 23) from P2 step 25
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (after 21, 23, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, metrics, and the ledger blast-radius workstream, independent of the audit cycle.
- Pre-committed contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
- A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- By day 7, every suspected major incident uses one record, one channel, and a named commander; Support can declare without engineering confirmation.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of customer-visible SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: d55e2704-ef8a-4e9c-8fd5-e6d956728a8d, Agent: opus5_refine_1, LLM: anthropic/claude-opus-5):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (32):
1. Charter the program: one owner, one mandate, funded, dated before the audit
Turn the CEO email into a chartered company program with a single accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the **burning-platform narrative** and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the written on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (depends on: 5, 7)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 8, 4)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (depends on: 10, 5)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (depends on: 7, 12, 13)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
18. SLA credit and financial impact workflow (depends on: 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Use the credit delta as the standing business case for on-call pay and reserved reliability capacity.
19. Blameless postmortem standard and Incident Review Board (depends on: 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, reserved capacity and enforcement (depends on: 19, 12)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
21. Runbooks, readiness bar and ledger blast-radius reduction (depends on: 5, 8)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (depends on: 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (depends on: 22, 12, 21)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
24. Publish Incident Management Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 11, 12, 21, 22, 9)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Alert noise burn-down campaign (depends on: 10, 12, 25)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
28. Wave rollout to all 28 teams with readiness gates (depends on: 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
29. Change management, fairness and pager culture (depends on: 4, 9, 25)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (depends on: 24, 26, 28)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 26, 28, 30)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 2 (ID: ddb335b2-98d9-46f8-afa0-88fca2e25749, Agent: gpt5.6-sol_refine_2, LLM: openai/gpt-5.6-sol):
Estimated Complexity: high
Success Metrics: - By day 7, every suspected major incident uses one incident record, one coordination channel, and a named commander.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes; at least 95% receive command within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and technical coverage.
- By week 18, all 180 services have an owner, tier, tested escalation path, and coverage appropriate to their risk.
- No mandatory night rotation starts before compensation, training, access, runbooks, and staffing controls are active.
- Every direct 24x7 technical rotation has at least six qualified responders or a documented, expiring executive exception.
- No responder is routinely assigned primary duty more often than one week in six or simultaneously assigned to two primary rotations.
- At least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of qualifying SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet their required internal and customer update cadence by month 3.
- Monthly human pages fall from 3,400 alert events to no more than 1,500 by day 90 and no more than 500 actionable pages by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, with no decline in Tier 0/1 detection coverage.
- Average after-hours load remains at or below two pages per responder per week; every sustained breach produces a remediation plan.
- All six monitoring sources route human pages through the single paging platform by week 12; direct legacy paging paths are disabled by week 18.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Monthly contracted availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- Quarterly on-call surveys reach at least 75% favorable responses on fairness, ownership boundaries, compensation, and sustainability by month 6.
- The month-7 mock audit finds no unowned high-risk control gap, and at least 95% of sampled incidents contain complete operating evidence.
Steps (26):
1. Establish the mandate, owner, funding, and schedule
Launch incident management as a **company operating program within 48 hours**. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Form a decision group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Approve the non-negotiables: one severity scale, one human-paging platform, one incident record, paid on-call, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, and incident actions.
- Fund tooling, compensation, training, exercises, and reliability work. Use the $1.3M in credits as the minimum financial comparison.
- Set milestones: interim process by day 7, policy and critical-path pilot by week 6, Tier 0/1 coverage by week 10, company rollout by week 18, internal audit in month 6, and mock audit in month 7.
2. Install a seven-day incident-response floor (depends on: 1)
Do not wait for policy design or tooling migration. Put a minimum process into operation immediately and begin retaining evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communication schedule.
- Create one monitored declaration path through chat, telephone, and the current paging environment.
- Use one incident channel, bridge, timeline document, and naming convention for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced on-call engineers and engineering leaders.
- Require a named commander within 10 minutes. The duty engineering director assumes command if nobody else does.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Compensate interim duty retroactively under the permanent policy.
- Hold a 15-minute daily operational review until the permanent process is live.
3. Create the factual, legal, and audit baseline (depends on: 1)
Build one defensible baseline for process design, executive decisions, and SOC 2 testing. Confirm the auditor's expected Type II observation period immediately.
- Reconstruct all 31 incidents from first customer impact through resolution, communications, credits, postmortem, and corrective action.
- Analyze the two incidents with unclear command minute by minute.
- Identify the missing detection signal for every customer-first incident.
- Inventory the six alert sources, 3,400 monthly alert events, noisy rules, duplicates, missing owners, and missing runbooks.
- Record the current rotations, unpaid work, after-hours load, and teams without coverage.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M credits, 85% noise, and 11 of 64 actions closed.
- Map incident-response controls to applicable SOC 2 criteria with Compliance and the auditor.
- Establish retention, confidentiality, legal-hold, and access requirements for incident evidence.
4. Co-design the fairness contract with engineers (depends on: 1)
Treat pager resistance as a valid design constraint. Make the employment and ownership bargain explicit before expanding on-call.
- Interview representatives from all 28 teams and the Support, Customer Success, Security, and Operations groups.
- Distinguish objections involving unpaid work, unfamiliar code, bad alerts, weak runbooks, sleep disruption, or blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for services they own or have formally accepted, supported by trained command, and given capacity to remove recurring defects.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Measure baseline trust, fairness, fatigue, and psychological safety. Repeat the survey at days 60 and 120, then quarterly.
5. Build the service catalog and criticality model (depends on: 3)
Make a machine-readable catalog the source of truth for routing, escalation, customer impact, and audit evidence. Every production service must have one accountable owner.
- Record the owning team, manager, product capability, escalation policy, communication channel, dashboard, runbook, dependencies, regions, and data stores for all 180 services.
- Map payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, and reporting to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing services whose failure can create material contractual impact.
- Classify Tier 2 as internal or deferrable services, and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional mode, recovery method, and contractual obligations for Tier 0 and Tier 1.
- Give orphan services an owner or approved decommission date within 30 days.
- Maintain separate but coordinated ownership for the ledger application and PostgreSQL platform.
6. Adopt one severity scale and incident lifecycle (depends on: 3, 5)
Use four impact-based levels. Classify on actual or credible customer, financial, security, regulatory, and contractual harm rather than organizational seniority.
- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; both regions impaired; or a core payment journey broadly unavailable. Page every role immediately, open a bridge, notify executives, publish customer status within 15 minutes when applicable, begin legal assessment, and require a postmortem.
- **SEV2 — major:** material payment degradation; settlement deadline at risk; a critical customer or material cohort unavailable; regional impairment with reduced resilience; or likely SLA breach. Page command and technical roles, publish customer status within 30 minutes when customer-visible, and require a postmortem.
- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads; page only when immediate action is necessary.
- **SEV4 — operational event:** no current customer impact and no urgent risk. Create a ticket and handle during normal operations.
- Treat percentages as supporting guardrails, never reasons to under-classify integrity, settlement, security, or contractual risk.
- Anyone may declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Automatically use at least SEV2 posture for suspected ledger-integrity events, cross-team incidents with unknown ownership, and materially unknown impact lasting 15 minutes.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed. Resolution requires stability, backlog recovery, and necessary reconciliation.
7. Define roles, authority, and handoffs (depends on: 6)
Separate command, communications, recordkeeping, and technical repair. One named person must hold command throughout every SEV1 and SEV2.
- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not serve as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- **Subject-Matter Responders:** diagnose and mitigate only services for which they have ownership, access, training, or formally accepted support responsibility.
- **Executive Duty Officer:** removes organizational barriers and makes exceptional business decisions without displacing the commander.
- Security, Legal, Compliance, Vendor Management, and Finance join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Require distinct commander, communications, scribe, and primary technical lead for SEV1. Communications and scribe may combine temporarily for bounded SEV2 incidents.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, and reconciliation requirements for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, with the exact transfer time.
8. Create sustainable 24x7 coverage across 28 teams (depends on: 4, 5, 7)
Use central command coverage and risk-based technical coverage rather than creating 28 fragile night rotations. All services receive a response path, but only critical domains maintain direct overnight technical rotations.
- Create a 24x7 Incident Command corps of 24–30 certified people, with primary and secondary coverage at all times.
- Create 18–24 trained Communications Leads and a similarly sized scribe pool using Support, Customer Operations, Engineering Operations, and qualified managers.
- Maintain a surge roster for simultaneous incidents and a 24x7 Executive Duty Officer schedule.
- Group Tier 0 and Tier 1 ownership into approximately 8–12 coherent domains only where responders have training, access, runbooks, and explicit acceptance.
- Staff each critical domain with primary and secondary responders and at least six qualified people. Target no more than one primary week in six.
- Give Tier 2 and Tier 3 services business-hours coverage plus a maintained manager and director escalation path.
- When lower-tier impact becomes SEV1 or SEV2, central command activates the manager escalation path and obtains the necessary owner.
- Route unknown-owner pages to Duty Command and platform triage temporarily. Log each one as a catalog control failure.
- Do not merge small teams merely for schedule convenience. Provide training, reassign services, add staffing, or decommission unsupported services.
9. Approve compensation and fatigue protections (depends on: 4, 8)
End unpaid on-call before expanding mandatory night coverage. HR, Finance, Payroll, and employment counsel should approve the policy within 14 days.
- Use market-validated weekly bands, initially budgeting approximately $800–$1,200 for Tier 0/1 primary duty and $300–$500 for secondary duty.
- Budget approximately $900–$1,300 for Duty Incident Commander weeks and $400–$800 for Communications Lead or scribe duty, adjusted for actual burden.
- Pay holiday premiums and compensate active after-hours work according to exempt or non-exempt status and applicable federal and New York rules.
- Provide a protected paid recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal sprint commitments by approximately 15% during a primary on-call week.
- Prohibit simultaneous primary assignments, consecutive primary weeks, on-call during leave, and invisible schedule swaps.
- Allow responders to declare themselves temporarily unfit after disruptive night work without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Budget roughly $0.9M–$1.2M annually, then refine using actual rotation count, employment classification, and activation data.
10. Establish one paging and incident system of record (depends on: 3, 5, 8)
Monitoring tools may remain specialized, but all human pages must enter one controlled platform. This removes conflicting schedules and creates one evidence trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page.
- Ingest events from the six existing monitoring tools before disabling their direct paging paths.
- Route pages through service-catalog ownership and deduplicate related events.
- Provide a single declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, severity changes, decisions, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status-page publishing, and conference facilities.
- Require role-based access, MFA, access reviews, and immutable or tamper-evident history.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, or an AWS region.
- Retire each legacy human-paging path only after ownership review, end-to-end tests, and two weeks of verified operation.
11. Enforce an alert-quality contract and page budget (depends on: 3, 5, 10)
Treat every page as a production interface with an owner and an expected action. Noise reduction must not create detection gaps.
- Require every paging rule to identify the service, owning team, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Page only when prompt human judgment or intervention can materially reduce customer, financial, security, or contractual risk.
- Route informational, capacity, and non-urgent infrastructure conditions to dashboards or ticket queues.
- Prefer payment-outcome, settlement-risk, queue-age, and error-budget-burn alerts over raw CPU, memory, pod, or log thresholds.
- Run new paging rules in shadow mode for at least seven days unless an emergency exception is approved.
- Review rules with repeated no-action acknowledgements, low actionability, or excessive firing within two business days.
- Set a load budget of no more than two after-hours pages per responder per week, measured over four weeks.
- Require an owner, compensating detection, and recorded approval before suppressing or deleting a rule.
- Burn down the top 50 noisy rules first. Review missed detections and noise together so teams cannot improve metrics by becoming blind.
12. Detect payment and ledger failures before customers (depends on: 5, 11)
Move detection from host health to customer journeys and financial outcomes. Use internal SLOs with enough headroom to protect the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, API access, webhooks, and reporting freshness.
- Run external synthetic transactions through critical payment journeys at least every minute and from paths independent of the production platform.
- Validate each AWS region and expose dependencies that undermine nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed value relative to settlement deadlines, not only queue depth.
- Add tenant and cohort anomaly detection for high-value customers and major payment methods.
- Convert high-priority Support, account-manager, processor, sponsor-bank, and network reports into incident candidates within five minutes.
- Record the detection source for every incident. Treat customer-first detection as a mandatory missed-detection review.
13. Codify detection, escalation, and live execution (depends on: 7, 10)
Create one time-bound path from the first credible signal to named command and mitigation. Notification delivery does not count as human acknowledgement.
- Page the owning critical-domain primary and Duty Incident Commander immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at five minutes, manager at 10, and director at 15.
- Escalate an unclaimed command page to backup commander at five minutes. The Executive Duty Officer assumes command at 10 minutes until a certified handoff occurs.
- Require Support and account managers to use the same declaration path as automated monitoring and engineers.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Begin with a standard statement of severity, known impact, assigned roles, current objective, workstreams, and next update time.
- Freeze unrelated changes during SEV1 and normally during SEV2. Record any exception.
- Prefer reversible mitigation such as rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require formal command handoff for long incidents, shift changes, or fatigue. Do not leave an incident unowned during transfer.
- Reconcile the ledger and safely drain backlogs before resolving payment or ledger incidents.
14. Standardize internal incident communications (depends on: 7, 10, 13)
Give responders one working room and stakeholders one controlled information source. Executives must not interrupt the technical command path.
- Create one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue an initial internal brief within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within 15 minutes for SEV1 and 30 minutes for SEV2.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record material decisions and outbound messages in the incident timeline.
15. Standardize status-page and customer communications (depends on: 6, 10, 14)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate impact before root cause is known.
- Publish customer status within 15 minutes of declaring a customer-visible SEV1 and within 30 minutes for customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice promptly after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map status-page components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, data-integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, known workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Give affected strategic accounts direct account-manager outreach within 30 minutes for SEV1 and 60 minutes for SEV2.
- Require account managers to use the approved briefing and prohibit independent technical explanations.
- Provide a customer-facing incident summary within five business days for SEV1 and qualifying SEV2 events.
- Record any legally necessary delay or restriction of public detail, its approver, and the alternative communication plan.
16. Operationalize legal, regulatory, contractual, and credit decisions (depends on: 3, 6, 15)
Some payments incidents start external notification clocks. Make assessment mandatory without assuming that every operational incident is reportable.
- Build a counsel-validated matrix covering applicable NYDFS rules, state breach laws, GLBA or FTC obligations, PCI requirements, money-transmitter obligations, sponsor-bank and network contracts, cyber insurance, and customer contracts.
- Engage Legal and Compliance immediately for every SEV1, security event, and suspected ledger-integrity event.
- Record an initial reportability assessment within one hour for SEV1 and within two hours for other potentially reportable events.
- Record non-reportable decisions as evidence, including facts considered, approver, timestamp, and reassessment trigger.
- Maintain tested 24x7 contacts for counsel, regulators, sponsor banks, networks, insurers, and critical vendors.
- Encode customer-specific notice deadlines and channels in the customer record used by the Communications Lead.
- Have Legal own regulatory text and submission. Keep technical command with the Incident Commander.
- Have Finance calculate affected minutes, delayed value, likely credits, and contractual exposure within five business days.
- Track credits by incident and recurring cause to support reliability investment decisions.
17. Make postmortems mandatory, consistent, and blameless (depends on: 6, 7, 10)
Use one learning standard with fixed deadlines. Keep postmortems separate from performance, misconduct, and disciplinary processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact over two hours, SLA credits, contractual breach, repeated contributing factors, control failures, and ledger-integrity near misses.
- Produce the factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the Incident Commander responsible for the timeline and the owning engineering director accountable for completion.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit analysis of why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the incident commander.
- Describe decisions in the context and information available at the time. Do not name an individual as the root cause.
- Publish broadly useful findings internally while restricting security, privacy, personnel, or privileged material appropriately.
18. Give corrective actions enforceable ownership (depends on: 10, 17)
Treat incident actions as risk commitments, not suggestions. A ticket is not complete until the expected risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Escalate overdue high-risk actions to the manager after seven days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit overdue actions to block related releases where the unrepaired condition could reproduce severe impact.
- Verify effectiveness through tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days. Complete, re-plan, or formally accept each risk.
19. Set the service-readiness bar and critical playbooks (depends on: 5, 8, 11, 12)
A critical service must be supportable at 3 a.m. before its team is placed on direct overnight coverage. Existing critical detection must remain active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and data-integrity constraints.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional failure, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document controlled failover, split-brain prevention, replay protection, backlog recovery, and post-recovery reconciliation.
- Exercise critical runbooks at least twice per year and after material changes.
- Block new Tier 0/1 releases and new paging rules when readiness requirements are missing.
- Handle existing gaps through a named owner, compensating control, executive-approved expiry date, and remediation plan rather than disabling detection.
20. Publish policy and certify every response role (depends on: 6, 7, 8, 9, 11, 13, 14, 15, 16, 17, 18)
Convert the operating design into concise, signed documents and practical training. Training and exercises occur during paid working time.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, and exception standards.
- Provide one-page severity, role, authority, escalation, and communication cards inside the incident tool.
- Train all employees to recognize and declare incidents.
- Train all engineers in severity, acknowledgement, evidence preservation, handoff, and financial-integrity precautions.
- Certify responders only after they demonstrate access, dashboards, runbooks, rollback, and escalation competence.
- Certify Incident Commanders through formal instruction, simulation, and at least two shadowed incidents or exercises.
- Train Communications Leads in status writing, account segmentation, contractual clocks, and legal escalation.
- Train scribes in timeline quality and fact-versus-hypothesis labeling.
- Require two shadow shifts before independent primary duty and annual recertification.
- Maintain the training and certification register as operational and audit evidence.
21. Pilot on the payment critical path (depends on: 9, 10, 12, 19, 20)
Run a four-to-six-week pilot across the highest-risk journey before expanding. Use real incidents and exercises to correct the model quickly.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, API edge, authentication, settlement, and Support intake.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently assuming command.
- Correct critical process or tool defects within 48 hours and update the standard visibly.
- Exit only after 95% timely command assignment, 95% communication compliance, no unpaid pages, complete required postmortems, and at least 50% lower noise.
22. Roll out by customer journey and risk (depends on: 21)
Expand in controlled waves while the interim process remains active company-wide. Complete rollout early enough to accumulate operating evidence before the audit.
- Roll out remaining Tier 0 domains first, then Tier 1, Tier 2, and Tier 3.
- Use two-to-three-week waves with a named program coach and director sign-off.
- Gate each team on catalog ownership, appropriate coverage, compensation, trained responders, tested escalation, alert quality, runbooks, access, and a passed tabletop.
- Require six or more responders only for direct 24x7 technical rotations. Apply business-hours coverage and manager escalation to lower tiers.
- Reschedule failed gates or approve a time-limited executive exception with a compensating control.
- Disable legacy human-paging paths after verified cutover for each wave.
- Publish an internal adoption dashboard by team, service tier, and control gap.
- Finish critical coverage by week 10 and all 28 teams by week 18.
23. Exercise command, communications, regional recovery, and fallbacks (depends on: 19, 20)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous security and operational events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors where appropriate.
- Conduct at least one unannounced after-hours paging test before the audit.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure time to acknowledgement, command, customer notice, mitigation decision, handoff, and recovery.
- Create tracked actions for every material exercise finding.
24. Measure performance and operate fixed review forums (depends on: 3, 10, 17, 18)
Use a balanced scorecard that exposes weak controls without rewarding hidden incidents or suppressed alerts. Report medians and 90th percentiles rather than averages alone.
- Measure time from first impact to detection, declaration, acknowledgement, command assignment, mitigation, recovery, and resolution.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track pages, actionability, duplicates, after-hours load, missed detections, and page-budget breaches.
- Track postmortem timeliness, action age, closure by due date, verified effectiveness, and recurring contributing factors.
- Track availability by customer journey, error-budget burn, failed or delayed payment value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight activations, recovery days, schedule exceptions, sentiment, and attrition.
- Hold a weekly Incident Review Board for incidents, postmortems, control failures, noisy alerts, and overdue actions.
- Hold a monthly executive reliability review for trends, funding, contractual exposure, and accepted risks.
- Hold a quarterly control and resilience review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile incident records monthly against support cases, status history, credits, and customer complaints to detect under-reporting.
25. Prove SOC 2 operating effectiveness before fieldwork (depends on: 20, 22, 23, 24)
Generate audit evidence through normal operation rather than reconstructing it later. Test both control design and consistent execution.
- Maintain approved and versioned policies, exceptions, catalog records, schedules, compensation activation, access reviews, training, incidents, communications, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, non-reportable legal decisions, missed timelines, and exercises in the population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the populations, evidence requests, and interviews expected from the external auditor.
- Record deviations honestly with owners, remediation dates, and compensating controls. Never rewrite historical records.
- Confirm that evidence retention covers the full auditor-defined observation period.
- Brief commanders, responders, Support, and Compliance on the real process without scripting inaccurate answers.
26. Institutionalize improvement and reduce structural risk (depends on: 22, 24, 25)
Prevent the program from decaying after the audit. Use incident evidence to drive permanent ownership and architectural investment.
- Assign permanent owners for policy, service catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review severity thresholds, staffing, compensation, and communication timings annually and after material process failures.
- Recertify command and communications personnel annually.
- Review recurring failure families quarterly and require executive decisions where corrective work repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fairness, fatigue, and tool friction.
- Report severe incidents, credits, overdue high-risk actions, and resilience investment to the board or risk committee quarterly.
- Treat the shared ledger cluster as a strategic concentration risk. Fund failover assurance, blast-radius reduction, isolation of non-critical readers, and stronger regional independence.
- Evaluate follow-the-sun command or technical coverage using one year of page-load and staffing data.
- Build year-two plans for automated mitigation, deployment safety, graceful degradation, and error-budget release controls.
Previous Proposal 3 (ID: f7ee763b-3b76-4f07-9e7d-5164701bef1c, Agent: qwen3.8-max_refine_3, LLM: alibaba/qwen3.8-max):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (32):
1. Establish executive mandate, program office, and funding
Turn the CEO email into a **chartered company program** with one accountable owner and authority over all 28 teams.
- Appoint the CTO as executive sponsor and a Head of Reliability as program owner with full-time authority.
- Create a permanent program office: one program lead, one platform engineer, one analyst.
- Form an eight-person steering group: Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours.
- Lock the non-negotiables: one severity scale, one paging platform, one postmortem format, paid on-call, mandatory action tracking, named service ownership.
- Approve budget anchored against the $1.3M in credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs.
- Reserve 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Publish a one-page charter stating incident response is a company operating process, not a per-team choice.
2. Install a seven-day interim command bridge (depends on: 1)
Do not leave the company unprotected while the permanent process is designed. Put a **crude but real** command structure in place within one week.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Direct Support to escalate credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical actions: complete, re-plan, or formally risk-accept, prioritizing ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Build the forensic baseline of incidents, alerts, and money lost (depends on: 1)
Rebuild the facts before designing anything. This becomes both the **design input** and the before picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific missing signal. This list drives the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Freeze baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the on-call fairness deal (depends on: 1)
Engineer pushback is the **largest delivery risk**. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS, and Sales in two weeks. Separate the real objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey covering fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
5. Map compliance, evidence, and audit requirements from day one (depends on: 1)
Design evidence as a **by-product of operations**, not a reconstruction before the auditor arrives. Confirm the SOC 2 observation window immediately.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for communications; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure, training register, drill records, reportability decisions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve all policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has a **named owner**. Build a machine-readable catalog as the single source of truth.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact: Tier 0 (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for each Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
7. Adopt the severity scale, declaration rules, and lifecycle (depends on: 3, 6)
Replace judgment calls with a **lookup table**. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter.
- SEV1 (crisis): money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of all roles and executive; bridge in 5 minutes; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 (critical): material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- SEV3 (major, contained): narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected, Declared, Triaged, Mitigated (customer impact ends), Monitoring, Resolved (backlog processed and ledger reconciled), Reviewed.
- Publish a decision tree with 12 worked examples from the real 31 incidents.
8. Define incident roles, authority, and handover discipline (depends on: 7)
Solve nobody-in-charge-for-an-hour by making command **explicit, single-holder, transferable, and logged**. Separate coordination from debugging.
- Incident Commander: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates for SEV1/SEV2.
- Subject-Matter Responders: engineers of the owning team; they mitigate, they do not run the room.
- Executive Duty Officer (SEV1): removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Design the three-layer 24x7 coverage model (depends on: 6, 8)
Do not create 28 night rotations. **Centralise coordination** in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- Layer A, Incident Command corps: approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of approximately 18 from Support, CS, and engineering management. A scribe pool used as the training entry point.
- Layer B, critical-path domain rotations: consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- Layer C, everyone else: business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing: a US-only paid night rotation now, a follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
10. Approve on-call compensation, labor compliance, and fatigue safeguards (depends on: 9, 4)
Unpaid on-call in New York is both a **retention problem and a legal exposure**. Pay must be live in payroll before any mandatory night rotation starts.
- Indicative scheme: approximately $1,000 per primary 24x7 week, $400 secondary, $250 for business-hours rotations, a separate $1,200 Duty Commander stipend, holiday premiums, approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- HR, Finance, and employment counsel publish amounts, eligibility, tax treatment, and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a 15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
11. Set the alert quality standard, page budget, and noise burn-down (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes. Make alert quality a **condition of being allowed to page a human**.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly.
- Target 3,400 to under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
12. Build payment-outcome detection and ledger assurance (depends on: 11, 6)
Stop customers telling you first. Detection must be driven by **payment outcomes and ledger truth**, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer detected first becomes a named defect class with a mandatory tracked action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Consolidate to one paging platform, one incident record, one status page (depends on: 7, 9, 11)
Collapse six alerting tools into a **single operational system of record** so there is one queue, one timeline, and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is unavailable. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
14. Codify the escalation ladder and five-minute command rule (depends on: 8, 9, 13)
Write one unskippable path from something looks wrong to **someone is in charge**. The default action is never waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately, secondary at 5 minutes, domain manager and Duty Commander at 10, Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension, and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
15. Write the live incident execution doctrine and major-incident playbooks (depends on: 8, 13, 14)
Give responders one short operating procedure from the first minutes through closure. Priority is **limiting customer and financial harm**, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the single largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window, an explicit handback to the owning team, Support, and Customer Success, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
16. Standardize internal communications protocol (depends on: 8, 13)
Standardise the internal picture so executives, Support, and Sales are informed **without interrupting** the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
17. Build customer communications, status page, and account-manager outreach (depends on: 7, 16)
Replace whoever is around with a **timed, owned, pre-approved process**. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook, and RSS. All 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity, or blame.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
18. Create the regulatory, partner, and legal notification playbook (depends on: 7, 17)
In payments, some incidents start a **legal clock at detection**. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded even when the answer is not reportable, with evidence, decision-maker, and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
19. Operationalize SLA credit and financial impact workflow (depends on: 7, 17)
Tie incidents to money so severity, credits, and investment decisions **stay honest**, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one. Use the delta as the standing business case for on-call pay and reliability capacity.
20. Establish the blameless postmortem standard and Incident Review Board (depends on: 7, 8)
Replace some incidents, various formats with **one mandatory format, fixed deadlines, and a forum with teeth**.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly top-five recurring causes analysis.
21. Enforce action ownership, capacity reservation, and tracking (depends on: 20, 13)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of **customer commitments** and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
22. Build the training and certification academy (depends on: 8, 14, 16, 20)
Command is a skill, not a title. **Certify before assigning duty**, and use paid working time for all of it.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days + shadowing): command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Comms Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum; never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Run the exercise programme: tabletops, game days, and unannounced drills (depends on: 22, 13, 15)
The process must meet a **simulated SEV1** before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger. Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (depends on: 7, 8, 9, 10, 11, 14, 16, 17, 18, 20, 21)
Collapse the design into a document people will **actually open mid-outage**, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 12, 13, 15, 22, 10)
Prove the process on the **highest-risk surface** with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment.
- Publish a one-page result to the whole company.
26. Run the alert noise burn-down campaign (depends on: 11, 13, 25)
Run noise reduction as a **visible, quota-driven campaign** in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness. Review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 to 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
27. Establish metrics, dashboards, review cadence, and anti-gaming (depends on: 13, 20, 25)
Instrument the process itself so improvement is visible and the auditor sees **evidence of monitoring and review**. Report median and 90th percentile, never averages alone.
- Response: time to detect, declare, commander, acknowledgement, mitigate, resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
28. Wave rollout to all 28 teams with readiness gates (depends on: 25, 27)
Roll out in four waves of six to eight teams every two to three weeks, ordered by **customer risk**. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. Failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Publish a live adoption scoreboard.
29. Run change management, fairness, and pager culture programme (depends on: 4, 10, 25)
Run this from day one in parallel. Engineers judge the process on **fairness**; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. Build SOC 2 evidence by design, internal testing, and mock audit (depends on: 24, 27, 28)
Design the evidence as a **by-product of doing the work**. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any change as an exception.
31. Maintain the program risk register and contingencies (depends on: 1)
Name the ways this programme fails and **pre-commit the response**. Review it monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 27, 28, 30)
Guard against the classic failure: the process **decays once the audit is signed**. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 4 (ID: 101b9773-a36a-47bf-8af6-14fe9a48abb2, Agent: grok4.6_refine_4, LLM: xai/grok-4.6):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (26):
1. Charter the program and start the audit clock
Convert the CEO email into a company operating process with one owner, a budget, and an observation window that starts this week.
Incident response is no longer a per-team choice.
- Name the CTO as sponsor and a Head of Reliability as the single accountable owner, with a three-person program office.
- Form a small decision group: Engineering, SRE, Support, Customer Success, Security, Legal, Finance, and HR. The sponsor decides within 48 hours.
- Lock non-negotiables: one severity scale, one paging path, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, and reserved engineering capacity against the $1.3M in SLA credits.
- Timeline: operating floor in 7 days, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, mock audit in month 7, audit in month 8.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for new tools or the final policy. Put a crude but real process in place this week so the next outage has a named commander.
This week is the first audit evidence.
- Publish a one-page interim severity card and one declaration path: Slack command, phone number, and existing pagers, all reaching the same duty person.
- Staff interim primary and backup Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in the incident channel.
- Tell Support to declare from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline (depends on: 1)
Rebuild the facts before locking design. This is both the design input and the before picture for the CEO and the auditor.
Freeze the numbers so they cannot drift during design.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisy rules and every rule with no owner or runbook.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11 of 64 actions closed, on-call in 12 of 28 teams.
4. Map resistance and publish the fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
Publish the deal in writing before any new mandatory pager is assigned.
- Interview all 28 teams plus Support, Customer Success, and Sales in two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the **fairness contract**: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and money-path tiers (depends on: 3)
You cannot page the right person across 180 services until every service has one owner. A wrong owner recreates the pager objection.
The catalog is the single source of truth for paging, impact, status-page components, and audit.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to services, data stores, regions, and third parties, including sponsor banks and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
6. Lock severity levels and what each one triggers (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by who reported it or how hard the fix looks.
Anyone may declare. Nobody is punished for over-declaring. Only the Incident Commander may downgrade, with the evidence recorded.
- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 major, contained: narrow or single-customer impact with a workaround. Owning team leads; business-hours comms; postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages.
- Attach a financial-integrity or security flag to any severity. The flag forces dual-control, Legal, and the regulatory checkpoint without inventing a fifth level.
- Auto-escalate: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer impact ends) → Monitoring → Resolved (backlog processed and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged.
Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. A human validates for SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1 and SEV2. Every handover is announced verbally and in writing with the exact time.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
8. Staff 24x7 with a command corps, not 28 night rotations (depends on: 5, 7)
Do not create 28 night rotations. That is what engineers are rejecting.
Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified volunteers from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to seven months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate 28 teams into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat Lisbon or APAC follow-the-sun as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
9. Pay on-call, meet New York labor rules, and cap fatigue (depends on: 4, 8)
Unpaid on-call in New York is a retention problem and a legal exposure. Pay must be in payroll before any new mandatory night rotation starts.
Publish the numbers. Then ask people to sign up.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
10. Set the alert-quality bar and a hard page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them.
Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers do (depends on: 5, 10)
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail.
Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One Slack command creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is down. Test weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Write the five-minute path from signal to command (depends on: 7, 8, 12)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting.
If nobody claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Codify live execution, readiness, and ledger playbooks (depends on: 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving a root cause.
A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down. No command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers.
15. Run one communications clock for internals, customers, and regulators (depends on: 6, 7, 12)
Replace whoever is around with one timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is not reportable.
- Obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
16. Tie incidents to SLA credits and true financial cost (depends on: 6, 15)
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
Credit calculation is an output of the incident record, not a negotiation.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
17. Make blameless postmortems mandatory and consistent (depends on: 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth.
The discipline lives in the deadlines and the review, not the template.
- Mandatory for every SEV1 and SEV2, any incident detected by a customer first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed.
18. Track actions as risk commitments with reserved capacity (depends on: 12, 17)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
Closing a ticket without evidence of effectiveness does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Target 90% of high-priority actions closed by due date within two quarters.
19. Train and certify every role before independent duty (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
The certification register is an audit artefact.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
20. Rehearse with tabletops, game days, and night drills (depends on: 12, 14, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
21. Publish Policy v1 and the signed fairness contract (depends on: 6, 7, 8, 9, 10, 13, 15, 17, 18)
Collapse the design into a document people will actually open mid-outage, and make it official.
Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
22. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 19, 21)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse. Publish a one-page result to the whole company.
23. Instrument metrics, reviews, and anti-gaming (depends on: 12, 17, 22)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review.
Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
24. Roll out in risk-ordered waves with readiness gates (depends on: 22, 23)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave, not as a background hope. Give every team its noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
25. Produce SOC 2 evidence by operating, then mock-audit (depends on: 21, 23, 24)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions. Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (depends on: 23, 24, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
Name the ways this programme fails and pre-commit the response.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 5 (ID: 289def1b-8f78-402b-85b3-0d9e18661842, Agent: deepseek-v4-pro_refine_5, LLM: deepseek/deepseek-v4-pro):
Estimated Complexity: high
Success Metrics: - Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, no loss of Tier 0/1 detection.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% Tier 0/1 services have named owning team, tier, escalation policy, dashboard and runbook by day 30; all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10-12 domain rotations of at least six trained responders.
- 30 certified incident commanders and 18 certified communications leads active.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% required postmortems drafted in 3 business days, reviewed in 5, published in 10 from month 2.
- Postmortem action closure rises from 17% to 90% closed by due date within two quarters; all 53 legacy actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k annualised within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% SEV1/security SEV2 including not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two cross-company exercises completed before audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% sampled incidents; SOC 2 Type II incident response controls pass zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days; no increase in attrition among responders.
Steps (26):
1. Executive mandate, budget, and governance
Secure a written CTO/CEO charter making incident management a company operating process, not a team option. Name one accountable Head of Reliability and a small steering group with authority to decide within 48 hours.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking, and protected engineering capacity.
- Approve budget against the $1.3M annual credits: tooling $150-250k/yr, on-call compensation, training, and 2-3 program FTEs.
- Set timeline: interim process in week 1, design weeks 1-6, pilot weeks 7-12, rollout weeks 13-24, mock audit month 7, SOC 2 at month 8.
- Freeze current baselines: 31 incidents, 22-min MTTD, 40% customer-first detection, 3h10 MTTM, $1.3M credits, 3,400 alerts/month at 85% noise, 11/64 actions closed.
2. Interim 7-day command (depends on: 1)
Put a crude but real process in place immediately so no incident remains unowned while the permanent design proceeds.
- Publish a one-page interim severity card, one declaration path (phone, Slack command, pager), and one incident channel/bridge/timeline naming convention.
- Staff an interim 24x7 duty commander with primary and backup from engineering managers and senior SREs; compensate retroactively under final policy.
- Require a named incident commander within 10 minutes of any suspected major incident, announced in channel.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions in 30 days, prioritizing ledger integrity, duplicate payment, regional failover, security and detection gaps.
- Hold a daily 15-minute operations review until the permanent process is live.
3. Forensic baseline and alert estate analysis (depends on: 1)
Reconstruct the true before picture from the last 12 months; it drives design and serves as audit baseline.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact start/detect/declare/commander/mitigate/resolve, who led, credits paid, contributing factors.
- For every customer-first detection, name the missing signal; this becomes the detection backlog.
- Reconstruct the two `nobody in charge` incidents minute by minute; use them as the burning-platform narrative and test case.
- Profile the 3,400 monthly alerts by tool, team, rule, outcome; identify top 50 noisy rules and every rule with no owner or runbook.
- Freeze all baseline metrics in one signed document for the executive and the auditor.
4. Listening tour and fairness contract (depends on: 1, 3)
Treat engineer pushback as the main delivery risk and convert it into a written deal that removes the objection to carrying a pager for other teams' code.
- Interview all 28 teams plus Support, CS and Sales in two weeks to separate objections: unpaid work, nights, unfamiliar code, missing runbooks, or fear of blame.
- Harvest practices from the 12 teams already on-call; they supply pilot teams and first commanders.
- Publish the deal: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Recruit 10-15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey and repeat at 60, 120 and 365 days.
5. Service ownership catalog and criticality tiers (depends on: 3)
Create a machine-readable catalog as the single source of truth for paging, routing, status page components, and audit evidence.
- Assign one accountable team, engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list per service.
- Tier by business impact: Tier 0 money movement/ledger/auth/shared PostgreSQL; Tier 1 customer-facing degradable; Tier 2 internal/batch; Tier 3 non-critical.
- Map customer journeys (initiate, authorize, settle, reconcile, report, onboard) to services, data stores, regions, third parties.
- Give orphan services an owner within 30 days or a decommission date approved by the sponsor; Tier 0 without an owner is an executive escalation.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
6. Severity scale, declaration rules, and automatic triggers (depends on: 3, 5)
Adopt one severity scale as a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade with recorded rationale. Start high when uncertain.
- SEV1: money moved wrongly/duplicated/lost, ledger integrity in doubt, confirmed security/data exposure, both regions impaired, payment processing halted. Full role page, bridge in 5 min, status page in 15 min, regulator assessment within 1 hour, mandatory postmortem.
- SEV2: material degradation of a payment journey, settlement window at risk, large/strategic customer fully down, SLA breach likely. Commander and SMEs paged, status page in 30 min, mandatory postmortem.
- SEV3: narrow or single-customer impact with workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any ledger cluster incident, any SEV3 open >2h, any unknown impact after 15 min, any cross-team incident become at least SEV2.
- Publish a decision tree with 12 worked examples from the 31 real incidents.
7. Incident roles, authority, and handoffs (depends on: 6)
Codify roles so coordination never depends on seniority or heroics. The Incident Commander coordinates; responders fix only services they own.
- IC: owns severity, priorities, roles, cadence, closure; pre-authorized to freeze deploys, rollback, disable features, shift traffic, invoke failover, put ledger read-only, commit spend; does not type in terminals; keeps command when VP joins.
- Communications Lead: sole author for status page, account managers, executives, and hand-off to Legal for regulators; speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, state changes; human validates for SEV1/SEV2.
- Subject-Matter Responders: diagnose and mitigate only their owned services.
- Executive Duty Officer (SEV1): removes obstacles, shields IC from exec questions, owns board/regulator escalation; does not take command unless formally transferred.
- Rule: IC claimed within 5 min and announced; distinct people for IC, Comms, technical lead at SEV1/SEV2; every handover announced verbally and in writing with time.
- Financial controls survive: IC coordinates ledger recovery but cannot bypass dual control, reconciliation or privileged access.
8. 24x7 command and responder coverage model (depends on: 5, 7)
Do not create 28 fragile night rotations. Centralize coordination in a trained command corps, keep technical ownership local.
- Layer A: incident command corps of ~30 certified volunteers with primary/secondary 24x7; paired Comms Lead pool ~18 and scribe pool as entry.
- Layer B: consolidate 28 teams into 10-12 critical-path domains (ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes, data/reporting, integrations) each with 24x7 primary+secondary and minimum six trained responders.
- Layer C: all other teams business-hours on-call with written after-hours escalation lists held by managers.
- Platform on-call is safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate follow-the-sun only as a 12-month option, not year-1 dependency.
- Acknowledgement discipline: 5 min at SEV1/SEV2 with automatic failover.
9. Paid on-call and fatigue safeguards (depends on: 8, 4)
Unpaid on-call in New York is a retention and legal risk. Pay must be in payroll before any mandatory night rotation starts.
- Indicative: ~$1,000 per primary 24x7 week, secondary ~$400, business-hours ~$250, duty commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond one hour.
- Mandatory recovery: paid recovery day after >2h overnight work, SEV1, or qualifying SEV2; managers arrange coverage.
- HR/Legal/Finance publish amounts, tax treatment, FLSA/NY wage-hour status within 14 days.
- Load rules: no primary more often than 1 in 6, no consecutive primary/secondary weeks, no primary on two rotations, no on-call during PTO, annual cap.
- If a rotation averages >2 out-of-hours pages per person per week for 4 weeks, trigger staffing or alert-remediation review.
- Count on-call as ~15% delivery load; document exemption path for health/caring without career penalty.
10. Service readiness bar and runbooks (depends on: 5, 8)
A service must earn the right to page a human at 3 a.m. Define a minimum readiness bar and major incident playbooks before any Tier 0/1 service goes live with on-call.
- Require for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switch, escalation contacts, data-loss/latency impact.
- Write playbooks for top failures: shared PostgreSQL failure/corruption, cross-region failover, Kubernetes control-plane loss, processor/sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, duplicate payment.
- Treat ledger cluster as largest structural risk: documented failover, read-only degraded mode, reconciliation after recovery, executive-signed RPO/RTO.
- Start a parallel workstream on blast-radius reduction: tenant/function partitioning, read replicas, isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging at night unless manager accepts a dated exception in writing.
- Runbooks are peer-reviewed, version-controlled, marked stale if not exercised twice a year.
11. Alert quality standard and page budget (depends on: 3, 5)
Make alert quality a condition of paging a human. Burn down noise deliberately rather than by mass silencing.
- Every paging alert must declare owning team, affected service, customer/SLO impact, expected action, dashboard, runbook, dedup key, severity mapping, escalation policy. Anything failing becomes a ticket.
- Page on symptoms of customer harm (SLO burn rate, payment success rate, queue age vs settlement deadlines), not raw CPU/memory.
- Run new alerts in shadow mode 7 days; test firing and recovery.
- Set page budget of 2 out-of-hours pages per person per week; breach blocks new alert creation and triggers tuning sprint.
- Auto-quarantine rules firing >5 times/month without action or >70% no-action acks; return to owner with correction deadline.
- Never silently disable: verify compensating detection, record decision, name owner, review missed detections monthly.
- Target 3,400 to <500 pages/month, actionability >75% within six months, no loss of Tier 0/1 detection.
12. Detection uplift on money path (depends on: 5, 11)
Stop customers telling you first by detecting on payment outcomes and ledger truth.
- Define SLOs and business SLIs per customer journey: initiation success, authorization latency, settlement timeliness, reconciliation break rate, API availability, reporting freshness; set internal targets stricter than 99.95%.
- Run external synthetic end-to-end payments every 60 seconds from both regions, covering all critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate detection, replication lag, failover-readiness, settlement-window countdown.
- Add per-customer anomaly detection for top 100 accounts; five similar support tickets in 10 minutes auto-creates triage incident.
- Route partner and processor notifications into declaration path within 5 minutes.
- Record detection source on every incident; treat `customer detected first` as a named defect class with mandatory action.
13. Consolidate to one paging and incident platform (depends on: 6, 8, 11, 12)
Collapse six alerting tools into one operational system of record: one queue, one timeline, one audit trail, without creating a monitoring gap.
- Select one paging/scheduling platform, one incident record, one hosted status page; time-box selection to 2 weeks.
- Ingest all six sources first, deduplicate/correlate; retire legacy path only after named owners, successful end-to-end test, and 2 weeks verified operation.
- One-command declaration in Slack creates channel/bridge, pages commander, sets severity, opens timeline, starts clock.
- Route every page through service catalog: service label -> owning domain schedule -> escalation policy.
- Capture evidence automatically: declaration, acks, role assignments, severity changes, decisions, comms, mitigated/resolved times, postmortem link; retain 12 months.
- Test out-of-band SMS/phone paging, mobile fallback, offline runbook weekly; ensure works during one AWS region/chat/identity provider failure.
- Set hard date after which pages outside this tool create no on-call obligation.
14. Escalation paths and acknowledgement SLAs (depends on: 7, 8, 13)
Write one unskippable path from signal to named commander in under 5 minutes; default action is never waiting.
- Converge all entry points on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately -> secondary at 5 min -> domain manager + Duty Commander at 10 -> Executive Duty Officer at 15. SEV2 requires command assigned within 15 min.
- If no one claims command in 5 min, platform assigns and announces it; assignee may hand over but not decline.
- Commander may page any domain's on-call directly with 10-min ack obligation; this reciprocity makes single-team ownership viable.
- If ownership unclear after 10 min, commander keeps incident and names temporary owner; missing catalog entry logged as control defect.
- Pre-authorize regional failover, ledger read-only, payment suspension, partner-bank notification so IC never waits for an executive.
15. Standardize internal and customer communications (depends on: 6, 7, 13)
Replace whoever is around with a timed, owned, pre-approved process; the Communications Lead is single author and never writes from scratch under pressure.
- Internal: one working channel/bridge plus one read-only broadcast for executives/Support/Sales; cadence SEV1 15-30 min, SEV2 60 min even if nothing changed; fixed template with impact, action, next update, IC/Comms names.
- Executives ask questions only of Executive Duty Officer; publish this as signed behavior rule.
- Status page: SEV1 within 15 min, SEV2 within 30 min; updates every 30/60 min; resolution notice within 30 min of verified recovery; customer-facing summary within 5 business days for SEV1/qualifying SEV2.
- Pre-approve 12-15 templates with Legal for degradation, settlement delay, API errors, regional failure, data-integrity investigation, security event.
- Top 100 accounts get named account manager call/email within 30 min of SEV1 with briefing pack; all 2,100 subscribed by default.
- Language rules: state impact and next update; never speculate on cause, recovery time, data integrity or blame.
16. Regulatory, partner, and financial impact workflows (depends on: 6, 15)
In payments some incidents start a legal clock at detection. Build obligation assessment into the process and tie incidents to money.
- Legal/Compliance produce obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC, PCI, money-transmitter, sponsor-bank/card-network windows, cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and security-related SEV2 within 1 hour, recorded even when `not reportable`, with decision-maker and evidence.
- Maintain 24x7 contact matrix for regulators, sponsor banks, networks, outside counsel, insurer.
- Agree availability measurement method per contract; compute affected minutes per customer from incident record and journey telemetry; propose credit schedule within 5 business days.
- Attribute credits to root-cause families; track failed payment volume, delayed value, reconciliation breaks; target annual credits from $1.3M to under $400k.
17. Blameless postmortem standard and action tracking (depends on: 6, 7)
Replace various formats with one mandatory format, fixed deadlines, and enforceable actions. Closing a ticket without evidence does not close the action.
- Mandatory for every SEV1/SEV2, customer-first detection, incident >2h, repeat of known cause, credit-generating/contract breach, ledger near-miss.
- Draft within 3 business days, review within 5, publish within 10; IC owns delivery, owning manager accountable.
- One template: summary, customer/financial impact, detection source/gap, timeline, response analysis, contributing conditions, what worked, actions.
- Two mandatory questions: why did a customer see this first? why did mitigation take as long as it did?
- Blameless in writing: systems and context, no individual named as cause, never used in performance reviews; HR handles personnel/misconduct separately.
- Every action gets owner, priority, due date, verification method, ticket auto-created; classes containment 7 days, corrective 30, strategic 90; reserve 15% engineering capacity.
- Escalation: manager at +7 days, director at +14, CTO dashboard at +30; overdue P0 can block related releases.
- Weekly Incident Review Board ratifies severity, challenges quality, monitors actions; target 90% closure on time within two quarters.
18. Train and certify incident roles (depends on: 7, 13, 15)
Command is a skill, not a title. Certify before duty; use paid working time.
- All employees: 30-min module on recognizing impact, declaring, finding channel/status page.
- Responder: half-day on severity, escalation, runbook, financial integrity; mandatory before joining rotation.
- Scribe: 2 hours timeline discipline; entry point.
- Incident Commander: 2 days + shadowing; command presence, delegation, severity calls, running room, handover; certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead: 1 day on status writing, customer tiering, legal boundaries, regulator triggers.
- Domain responders demonstrate dashboard, runbook, rollback, failover, access before primary; two shadow shifts; no primary in first 90 days.
- Certification valid 12 months, renewed by simulation; register is audit artifact.
19. Simulations and game days (depends on: 10, 14, 15, 18)
Rehearse process before real SEV1; exercise tooling failure and ledger scenarios.
- Monthly tabletop per group reusing an incident from baseline, rotating commander.
- Quarterly game day in staging or tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes degradation.
- Twice-yearly unannounced paging drill to measure real overnight ack times.
- Annually one combined operational/security exercise and one regulator-notification exercise with Legal/CISO.
- Exercise status page down, chat/paging provider unavailable.
- Never inject uncontrolled change into production ledger; validate backups/restore/RPO/RTO on replicas.
- Every exercise produces tracked actions; publish time-to-commander, time-to-first-update, time-to-mitigation.
20. Pilot on critical path (depends on: 9, 11, 12, 13, 14, 15, 17, 18, 19)
Prove full process on highest-risk surface with willing teams for six weeks before full rollout.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, Support intake, including two of the 12 already on-call teams.
- Activate severity scale, command corps, single paging tool, page budget, status page policy, mandatory postmortems, paid rotations.
- Parallel old paths for one week then cut over; real incidents use new process only.
- Program lead attends every SEV2+ as coach, never shadow commander; review every pilot page within one business day.
- Exit gate: commander named within 5 min in 95% cases, first status update on time, MTTD under 10 min for pilot services, pages halved, no unpaid page, all required postmortems on time, positive sentiment.
- Publish one-page result company-wide.
21. Metrics, dashboards, and review cadence (depends on: 13, 17)
Instrument the process so improvement is visible and audit sees monitoring/review evidence. Report median and p90, never averages alone.
- Response: time to detect/declare/commander/ack/mitigate/resolve split by severity, tier, journey, region, detection source.
- Quality: customer-first rate, status page timeliness, update cadence, missed escalations, role conflicts, alert actionability, out-of-hours pages per person.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat factors.
- Business: availability per journey, error budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, frequency, recovery days, sentiment, attrition.
- Weekly Incident Review Board; monthly reliability review per team; monthly executive review to CEO; quarterly control review with Security/Compliance/Risk; quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, status page history; team scorecards direct help, never individual penalties.
22. Wave rollout to all 28 teams with readiness gates (depends on: 20, 21)
Roll out in four waves by customer risk, each passing an explicit gate rather than a date. Complete all teams by week 24 to leave ~3 months of evidence before audit fieldwork.
- Wave 1 remaining Tier 0, Wave 2 Tier 1, Wave 3 Tier 2, Wave 4 Tier 3/internal.
- Per-team onboarding kit: catalog entry complete, alerts migrated within budget, runbooks at readiness bar, rotation staffed with 6 trained responders or time-limited exception, one IC candidate nominated, one tabletop passed, payroll active.
- Named coach per wave for three weeks; director signs gate.
- Legacy paging paths disabled per wave, not kept as comfort fallback.
- Publish live adoption scoreboard.
- Any team unable to staff fair rotation gets headcount, service reassignment, or explicit executive risk acceptance.
23. Culture, fairness, and continuous feedback (depends on: 4, 9)
Run this in parallel from day one. Engineers judge fairness; executives judge results.
- Repeat the deal in every forum: paid on-call, paged only for owned services, trained commander, real sprint capacity.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
- Weekly office hours first 8 weeks, Slack support channel 4-hour SLA, per-team champion, weekly newsletter with metrics and bad news.
- Incident response contribution in promotion criteria; quarterly awards for best postmortem and biggest noise reduction; public thanks after SEV1.
- Pulse survey at 60, 120, 365 days; if fairness or load red, pause expansion until fixed.
24. SOC 2 evidence design and internal dry-run (depends on: 13, 17, 21, 22)
Make evidence a by-product of operations; test before external auditor.
- Map with Compliance to Trust Services Criteria CC7.2-CC7.5, CC2.2/CC2.3, A1.2; confirm observation window early.
- Evidence set automated and indexed: versioned policies/exception, rotation schedules, compensation activation, incident records with timestamps/roles, paging/ack logs, status history, reportability decisions, postmortems, action closure with verification, training/certification register, drill records, access reviews.
- Internal Audit tests process in months 4 and 6, sampling incidents end-to-end.
- Run formal mock audit month 7 using same evidence populations; include interviews with IC, engineer, Support, Compliance.
- Correct failures through tracked actions, never by editing history.
- Freeze process wording after month 7; any change logged as exception.
25. Risk register and contingency planning (depends on: 1)
Name likely failure modes and pre-commit responses; review monthly with sponsor.
- Too few commander volunteers: make command rostered duty for managers and staff engineers until pool reaches 30.
- Compensation not approved in time: fallback to time-off-in-lieu plus phased stipend, never launch mandatory night on-call uncompensated.
- Tool migration slips: cut scope on incident record layer, never on paging consolidation.
- Noise pruning hides real failure: demote to ticket first, observe 30 days, keep recovery path, monthly missed-detection review.
- Burnout in experienced teams: weekly load monitoring with page caps.
- Major SEV1 mid-rollout: program lead becomes full-time responder, wave schedule slips one wave, sponsor told same day.
- Shared ledger concentration: if blast-radius workstream slips, escalate to board as accepted risk with dated plan.
26. Inspect and adapt; year-two sustainability (depends on: 22, 24)
Prevent decay after audit by revising on data and assigning permanent owners.
- At 90 days live, revise policy using measurements: severity calibration, Layer B/C membership from page data, uncovered shifts, commander burn, missed updates, action closure, survey results.
- Assign permanent owners for policy, paging platform, status page, service catalog, training, metrics.
- Re-baseline targets every six months; shift from lagging to leading metrics: error budget burn, near-miss rate, drill performance.
- Year-two candidates: follow-the-sun coverage, automated mitigation for top recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, completion of ledger blast-radius reduction.
- Keep annual policy review, certification renewal, exercise calendar, board reporting as permanent commitments.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Charter the program, fund it, and start the audit clock", "description": "Turn the CEO email into a chartered company program within 48 hours. Incident response becomes a **company operating process**, not a per-team preference.\n\n- Name the CTO as executive sponsor and a Head of Reliability as full-time owner, with a three-person office: program lead, platform engineer, and analyst.\n- Form a small decision group of Engineering, SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit. It proposes. The sponsor decides within 48 hours.\n- Lock the non-negotiables now: one severity scale, one human-paging platform, one incident record, one postmortem format, mandatory action tracking, paid on-call, and named service ownership.\n- Confirm the SOC 2 Type II observation window with the auditor in week 1 so the interim process counts as evidence.\n- Fund tooling, compensation, training, exercises, and reserved engineering capacity against the $1.3M in SLA credits.\n- Reserve 15% of engineering capacity for detection, runbooks, and incident actions, protected by the sponsor.\n- Publish a one-page charter on day 2. Clock: floor by day 7, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.", "dependencies": []}, {"step_id": "S2", "title": "Install a seven-day operating floor", "description": "Do not wait for policy, tooling, or the audit. Put a crude but real process in place this week so the next outage already has an owner.\n\n- Publish a one-page interim severity card and one declaration path: chat command, phone number, and current pagers, all reaching the same duty person.\n- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.\n- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.\n- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.\n- Use one channel, one bridge, one timeline, and one naming convention for every suspected major incident.\n- Triage the 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.\n- Hold a 15-minute daily operations review until the permanent process is live.\n- Replay one of the two nobody-in-charge incidents as a tabletop within 14 days using this floor process.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Rebuild the forensic baseline and incident-replay catalog", "description": "Rebuild the facts before locking design. This is the design input, the frozen before-picture for the CEO, and the test set for detection work.\n\n- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, and resolve, who led, credits paid, and contributing factors.\n- For each of the 40% customer-first detections, name the exact missing signal. That list is the detection backlog.\n- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.\n- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Name the top 50 noisy rules and every rule with no owner or runbook.\n- Quantify true cost: credits, failed payment volume, delayed value, reconciliation breaks, and engineering hours lost.\n- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11 of 64 actions closed, on-call in 12 of 28 teams.\n- Build a **replay catalog**: for each historical incident, the detector that should now fire, at which minute, and the owner of the gap.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Publish the on-call fairness contract", "description": "Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.\n\n- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, and fear of blame. Each needs a different remedy.\n- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders. Do not punish them by making them the company's permanent night watch.\n- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.\n- Publish the **fairness contract** in writing: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.\n- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.\n- Provide a non-punitive exemption path for health, disability, or caregiving.\n- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.\n- No new mandatory night rotation starts until this contract, compensation, and training are live.", "dependencies": ["S1"]}, {"step_id": "S5", "title": "Build the service catalog, ownership, and journey tiers", "description": "You cannot route a page correctly across 180 services until every service has one owner. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.\n\n- Record owning team, manager, chat channel, escalation policy, dashboards, runbook, dependencies, regions, and data stores for all 180 services.\n- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.\n- Map every customer journey — initiate, authorise, settle, reconcile, refund, report, onboard — to services, data stores, regions, sponsor banks, and processors.\n- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.\n- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform. They fail differently.\n- Orphan services get an owner in 30 days or an approved decommission date. Unowned Tier 0 is an executive escalation and a release blocker.\n- Coverage follows the journey, not the org chart. Small teams that own Tier 0 pieces get headcount, service reassignment, or membership in a domain rotation. Never a two-person 24x7 rota.", "dependencies": ["S3"]}, {"step_id": "S6", "title": "Lock severity levels, integrity flags, and the incident lifecycle", "description": "Replace judgement calls with a lookup table. Classify on actual or credible customer, financial, security, and contractual harm, never on who reported it or how hard the fix looks.\n\n- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.\n- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms if customer-visible; status page in 30 minutes; mandatory postmortem.\n- SEV3 contained: narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.\n- SEV4: no customer impact. Ticket only. Never pages a human.\n- Attach an **integrity, security, settlement, or regulatory flag** to any severity. The flag forces dual-control, Legal, and the reportability checkpoint without inventing a fifth level. A one-customer ledger corruption is still a flagged crisis.\n- Auto-escalate to at least SEV2: any ledger-cluster event, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident.\n- Anyone may declare. Nobody is punished for over-declaring. Only the commander may downgrade, with evidence recorded.\n- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed.\n- Publish a decision tree with 12 worked examples taken from the real 31 incidents.", "dependencies": ["S3", "S5"]}, {"step_id": "S7", "title": "Define roles, authority, dual-control, and handover", "description": "Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.\n\n- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.\n- Communications Lead: single voice for the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.\n- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists. A human validates at SEV1 and SEV2.\n- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.\n- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.\n- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory text. The commander owns the facts.\n- Distinct people for command, comms, and technical lead at SEV1. Comms and scribe may combine only for bounded SEV2.\n- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.\n- Every role assignment and handover is announced verbally and in writing with the exact time and open risks.", "dependencies": ["S6"]}, {"step_id": "S8", "title": "Staff 24x7 with command, domains, and a Duty Triage Desk", "description": "Do not create 28 night rotations. That is what engineers are rejecting. Centralise coordination. Keep technical ownership local. Page people only for services they own.\n\n- Layer A, Incident Command corps: about 30 certified people from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to eight months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.\n- Layer B, critical-path domains: consolidate ownership into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.\n- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.\n- **Duty Triage Desk**: a small paid overnight first-line rotation that owns the first ten minutes of ambiguous, unowned, or low-confidence pages. It verifies, enriches, and applies only the runbook's safe steps, then wakes the owning domain. It never sits on a suspected ledger, payment-halt, or security event. Those page commander and likely domains immediately, in parallel.\n- Overnight comms: SEV1 pages a Communications Lead 24x7. Customer-visible SEV2 lets the commander publish the first status from a template; Comms is paged if the incident is still open at 30 minutes or a top-100 account is affected.\n- Staffing math: 260 engineers can sustain 10–12 domain rotations, one command corps, and one triage desk. They cannot sustain 28 night rotas. Role exclusivity: nobody is primary on two rotations in the same week. Commanders may also be domain responders in different weeks.\n- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.\n- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.\n- Evaluate follow-the-sun as a 12-month option, not a year-one dependency. Seed Layer B from the 12 teams already on-call.", "dependencies": ["S4", "S5", "S7"]}, {"step_id": "S9", "title": "Pay on-call, meet New York labor rules, and cap fatigue", "description": "Unpaid on-call in New York is a retention problem and a wage-hour exposure. Pay must be in payroll before any new mandatory night rotation starts.\n\n- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, $400–$800 for Comms or triage-desk duty, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.\n- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.\n- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.\n- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.\n- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.\n- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.\n- **Hard gate: no mandatory night rotation starts before compensation is live in payroll.** Interim duty is paid retroactively. Budget roughly $0.9M–$1.2M a year, then refine with actual rotation count.", "dependencies": ["S4", "S8"]}, {"step_id": "S10", "title": "Enforce an alert-quality contract and page budget", "description": "3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them. Make alert quality a condition of being allowed to wake a human.\n\n- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.\n- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.\n- Run new alerts in shadow mode for seven days. Test both firing and recovery, unless an emergency exception is recorded.\n- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.\n- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.\n- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.\n- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.", "dependencies": ["S3", "S5"]}, {"step_id": "S11", "title": "Detect payment and ledger failures before customers, then replay the past", "description": "Stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. A detector is not done until it would have caught the last 12 months.\n\n- Define SLIs and SLOs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.\n- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.\n- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.\n- Add per-customer anomaly detection for the top 100 accounts. Five similar support tickets in ten minutes auto-creates a triage incident.\n- Route credible partner, processor, and sponsor-bank notifications into the same declaration path within five minutes.\n- Record detection source on every incident. Customer-detected-first is a named defect class with a mandatory tracked action.\n- Run the **replay test** against the S3 catalog. For each of the 31 incidents, name the detector that would now fire and at which minute. Close gaps the replay exposes before calling detection improved.", "dependencies": ["S5", "S10"]}, {"step_id": "S12", "title": "Consolidate to one pager, one incident record, one status page", "description": "Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.\n\n- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.\n- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.\n- One chat command creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.\n- Route every page through the service catalog: service label to owning domain schedule to escalation policy. Unowned pages go to the Duty Triage Desk and Duty Command, and log a catalog defect.\n- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.\n- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test weekly.\n- Set a hard date after which a page outside this platform creates no on-call obligation.", "dependencies": ["S6", "S8", "S10"]}, {"step_id": "S13", "title": "Codify escalation, the five-minute command rule, and vendor incidents", "description": "Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting. Payments also fail at processors and banks, which you cannot patch.\n\n- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.\n- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.\n- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.\n- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.\n- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.\n- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.\n- Vendor-incident class: processor, sponsor-bank, card-network, or cloud-control-plane failure. Command is still required. Success is time-to-customer-notice, time-to-failover-decision, and queue management, not root cause at the vendor.\n- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.", "dependencies": ["S7", "S8", "S12"]}, {"step_id": "S14", "title": "Write live execution doctrine, readiness bars, and ledger playbooks", "description": "Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving root cause. A service must earn the right to page a human at night.\n\n- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.\n- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.\n- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis. Split diagnosis and mitigation workstreams once staffing allows.\n- Decisions are stated aloud and written down. No command by direct message.\n- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Drain backlogs under control. Complete reconciliation before any payment or ledger incident is declared resolved.\n- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.\n- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date. Existing detection stays on while gaps are repaired.\n- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.\n- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers, with board-visible milestones and executive-signed RPO/RTO.", "dependencies": ["S5", "S7", "S13"]}, {"step_id": "S15", "title": "Run one communications clock for internals, customers, and account managers", "description": "Replace whoever is around with a timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.\n\n- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.\n- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.\n- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.\n- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.\n- Status page from declaration: within 15 minutes for customer-visible SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.\n- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys, not internal service names. Subscribe all 2,100 customers by default.\n- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer record.\n- Language rules: state impact and the next update time. **Never speculate** on cause, recovery time, data integrity, or blame.", "dependencies": ["S6", "S7", "S12"]}, {"step_id": "S16", "title": "Operationalize regulatory notice, partner clocks, and SLA credits", "description": "In payments, some incidents start a legal clock at detection. Tie incidents to money so Finance does not learn about outages from invoices.\n\n- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, and cyber-insurer notice.\n- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when not reportable**, with facts, decision-maker, and timestamp.\n- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.\n- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.\n- Agree with Legal and Finance the availability measurement method per contract and per component.\n- Compute affected minutes per customer from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.\n- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.\n- Attribute credits to root-cause families. Track failed payment volume, delayed value, and reconciliation breaks as the true incident cost.", "dependencies": ["S6", "S15"]}, {"step_id": "S17", "title": "Make postmortems mandatory and actions enforceable", "description": "Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth. Eleven of 64 closed is a process nobody enforces.\n\n- Mandatory for every SEV1 and SEV2, any incident a customer detected first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.\n- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.\n- One template: executive summary, customer and financial impact, detection source and gap, timeline, response analysis, contributing conditions, control performance, what worked, actions.\n- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.\n- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately.\n- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items. Facilitators are trained and are never the commander of the incident under review.\n- Every action gets one named owner, priority, due date, expected risk reduction, verification method, and a ticket created automatically. If it is not on the board, it does not exist.\n- Classes: containment in 7 days, corrective in 30, strategic in 90. SEV1 recurrence-prevention enters the next sprint ahead of roadmap work. Reserve 15% capacity.\n- Overdue ladder: manager at +7 days, director at +14, CTO at +30. Overdue high-risk items need written residual-risk acceptance and can block related releases. Verify effectiveness before closing.", "dependencies": ["S6", "S7"]}, {"step_id": "S18", "title": "Train and certify every response role before independent duty", "description": "Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.\n\n- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.\n- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.\n- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.\n- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.\n- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.\n- Duty Triage Desk, half day: enrichment, safe-step limits, when to wake a domain immediately, when not to delay.\n- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.\n- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.", "dependencies": ["S7", "S13", "S15", "S17"]}, {"step_id": "S19", "title": "Publish Policy v1 and the exception register", "description": "Collapse the design into a document people will actually open mid-outage, and make it official. Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings.\n\n- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.\n- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.\n- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.\n- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.\n- Link the policy directly from the incident tool.", "dependencies": ["S6", "S7", "S8", "S9", "S10", "S13", "S15", "S16", "S17"]}, {"step_id": "S20", "title": "Pilot on the payments critical path", "description": "Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.\n\n- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.\n- Activate everything at once for them: severity scale, command corps, Duty Triage Desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.\n- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.\n- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.\n- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.\n- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.", "dependencies": ["S9", "S11", "S12", "S14", "S18", "S19"]}, {"step_id": "S21", "title": "Instrument metrics, reviews, and anti-gaming", "description": "Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.\n\n- Response: time from first impact to detect, declare, commander, acknowledge, mitigate, and resolve, split by severity, tier, journey, region, and detection source.\n- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.\n- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.\n- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.\n- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.\n- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.\n- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.\n- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.", "dependencies": ["S12", "S17", "S20"]}, {"step_id": "S22", "title": "Rehearse with tabletops, game days, and night drills", "description": "The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.\n\n- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.\n- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.\n- Twice-yearly unannounced overnight paging drill to measure real acknowledgement times.\n- One combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO, before the audit.\n- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.\n- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.\n- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.\n- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.", "dependencies": ["S12", "S14", "S18", "S19"]}, {"step_id": "S23", "title": "Burn down alert noise and roll out by risk with gates", "description": "Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Noise reduction is a quota inside each wave, not a background hope.\n\n- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.\n- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.\n- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.\n- Legacy paging paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.\n- Give every team its noisiest rules from the baseline. Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota.\n- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed. Pair every suppression with a compensating-detection check.\n- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.\n- If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.", "dependencies": ["S20", "S21"]}, {"step_id": "S24", "title": "Operate the fairness and culture program in parallel", "description": "Run this from the moment the deal is published. Engineers judge the process on fairness. Executives judge it on visible results. Both need constant, honest communication.\n\n- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.\n- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.\n- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.\n- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.\n- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.\n- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.", "dependencies": ["S4", "S9"]}, {"step_id": "S25", "title": "Prove SOC 2 operating effectiveness before fieldwork", "description": "Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.\n\n- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.\n- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training register, drill records, access reviews.\n- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.\n- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support, and Compliance.\n- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.\n- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.", "dependencies": ["S19", "S21", "S22", "S23"]}, {"step_id": "S26", "title": "Inspect at 90 days and lock year-two ownership", "description": "Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.\n\n- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.\n- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.\n- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, metrics, and the ledger blast-radius workstream, independent of the audit cycle.\n- Pre-committed contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.\n- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.\n- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.", "dependencies": ["S21", "S23", "S25"]}], "estimated_complexity": "high", "success_metrics": "- A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.\n- By day 7, every suspected major incident uses one record, one channel, and a named commander; Support can declare without engineering confirmation.\n- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.\n- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.\n- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.\n- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.\n- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.\n- Status page posted within 15 minutes of customer-visible SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.\n- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.\n- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.\n- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.\n- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.\n- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk. No two-person 24x7 rotation.\n- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.\n- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.\n- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.\n- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.\n- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.\n- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.\n- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.\n- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.\n- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.\n- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.\n- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.\n- Ledger blast-radius workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly."}It closed two real holes from its round-2 version — there was no policy-publication step and no separated postmortem/action or internal/customer comms steps — but it did so by copying P1's round-2 plan almost step for step. It contributes no idea of its own and picks up none of P1's round-3 advances.
- New step 22 publishes Incident Management Policy v1 plus the exception register; round 2 had no policy step at all, which was a SOC 2 gap.
- Split internal communications (step 15) from customer communications and status page (step 16), so cadence, templates and outreach each have an owner.
- Split blameless postmortems (step 18) from action ownership and enforcement (step 19), making the 11-of-64 problem its own enforcement step.
- Added the Duty Triage Desk to the coverage model (step 8) and a standalone alert noise burn-down campaign with quotas and a weekly leaderboard (step 25).
- Step 3 now confirms the SOC 2 observation window and maps controls to the Trust Services Criteria during the baseline.
- The standalone readiness-bar and runbook step disappeared into step 14, which now carries execution doctrine, readiness bar, eight playbooks and the ledger blast-radius workstream in one step.
- SLA credits lost their own step and are now two bullets at the end of the regulatory playbook (step 17), weakening the money-to-reliability feedback loop the others make explicit.
- Metrics still say paid on-call must be live "before any rotation pages a human", which contradicts its own step 2 interim 24x7 roster paid retroactively; P1 and P4 fixed this wording to "any new mandatory night rotation".
- Replay coverage is used in step 11 but never appears in the success-metric list, and metrics still claim SOC 2 passes with zero exceptions.
- Proposal 1 : Duty Triage Desk as a paid overnight first line owning the first ten minutes of a page.
- Proposal 1 : The replay test over the 31 historical incidents.
- Proposal 1 : Publish Incident Management Policy v1 with companion standards and an exception register.
- Proposal 1 : Separate internal communications protocol from customer/status-page policy.
- Proposal 1 : A quota-driven alert noise burn-down campaign run in parallel with rollout.
- Proposal 4 : Charter the programme and start the audit clock by confirming the observation window.
- Proposal 4 : Integrity/security flags on any severity level.
- Proposal 2 : Instrument the metrics scorecard before the pilot so failures show immediately.
- Proposal 1 : A standalone service-readiness bar, runbook and ledger blast-radius step.
+ Charter the program and start the audit clockCharter the program: one owner, one mandate, funded, dated before the auditSLA credit and financial impact workflowRunbooks, readiness bar and ledger blast-radius reduction
The plan produced
1. Charter the program and start the audit clock from P4 step 1
Turn the CEO email into a chartered company program with one accountable owner and authority across all 28 teams. Incident response becomes a company operating process, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (after 1) from P1 step 2
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (after 1) from P1 step 3
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
- Confirm the SOC 2 Type II observation window with the auditor; map the future response controls to applicable Trust Services Criteria; define evidence set and retention requirements.
4. Listening tour, resistance map and the written on-call deal (after 1) from P1 step 4
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (after 3) from P1 step 5
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: Tier 0 (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (after 3, 5) from P1 step 6
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- SEV1 (crisis): money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- SEV2 (critical): material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- SEV3 (contained): narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- SEV4: no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring. Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (after 6) from P1 step 7
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- Incident Commander: owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- Communications Lead: the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- Subject-Matter Responders: engineers of the owning team. They mitigate; they do not run the room.
- Executive Duty Officer (SEV1): removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation or privileged-access rules.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (after 5, 7) from P1 step 8
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- Layer A — Incident Command corps: ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- Layer B — domain responder rotations: consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- Layer C — everyone else: business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- Duty Triage Desk: a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (after 4, 8) from P1 step 9
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Hard gate: no mandatory night rotation starts before its compensation is live in payroll. Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (after 3, 5) from P1 step 10
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a page budget of two out-of-hours pages per responder per week. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (after 5, 10) from P1 step 11
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the replay test: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (after 6, 8, 10) from P1 step 12
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (after 7, 8, 12) from P1 step 13
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (after 7, 13) from P1 step 14
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- Payments guards: protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
15. Internal communications protocol (after 7, 12) from P1 step 15
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (after 6, 12, 15) from P1 step 16
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity or blame; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (after 6, 16) from P1 step 17
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded even when the answer is "not reportable", with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
- Agree with Legal and Finance the availability measurement method per contract and per component; compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide credit posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain; attribute credits to root-cause families for investment decisions.
18. Blameless postmortem standard and Incident Review Board (after 6, 7, 12) from P1 step 19
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did?
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
19. Action ownership, reserved capacity and enforcement (after 12, 18) from P1 step 20
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
20. Training and certification academy (after 7, 13, 15, 18) from P1 step 22
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days plus shadowing): command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Comms Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
21. Exercise programme: tabletops, game days and unannounced drills (after 12, 14, 20) from P1 step 23
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced overnight paging drill to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
22. Publish Incident Management Policy v1 and the exception register (after 6, 7, 8, 9, 10, 13, 15, 16, 17, 18, 19) from P1 step 24
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (after 9, 11, 12, 14, 20, 22) from P1 step 25
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
24. Metrics, dashboards, review cadence and anti-gaming (after 12, 18, 23) from P1 step 26
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- Response: time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
25. Alert noise burn-down campaign (after 10, 12, 23) from P1 step 27
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
26. Wave rollout to all 28 teams with readiness gates (after 23, 24, 25)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. A failed gate is rescheduled, never waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
27. Change management, fairness and pager culture (after 4, 9, 23) from P1 step 29
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
28. SOC 2 evidence by design, internal testing and mock audit (after 22, 24, 26) from P1 step 30
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
29. Risk register and pre-committed contingencies (after 1) from P1 step 31
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
30. Ninety-day inspect-and-adapt, then year-two sustainability (after 24, 26, 28) from P1 step 32
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to leading indicators: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
- Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: d55e2704-ef8a-4e9c-8fd5-e6d956728a8d, Agent: opus5_refine_1, LLM: anthropic/claude-opus-5):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (32):
1. Charter the program: one owner, one mandate, funded, dated before the audit
Turn the CEO email into a chartered company program with a single accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the **burning-platform narrative** and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the written on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (depends on: 5, 7)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 8, 4)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (depends on: 10, 5)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (depends on: 7, 12, 13)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
18. SLA credit and financial impact workflow (depends on: 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Use the credit delta as the standing business case for on-call pay and reserved reliability capacity.
19. Blameless postmortem standard and Incident Review Board (depends on: 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, reserved capacity and enforcement (depends on: 19, 12)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
21. Runbooks, readiness bar and ledger blast-radius reduction (depends on: 5, 8)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (depends on: 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (depends on: 22, 12, 21)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
24. Publish Incident Management Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 11, 12, 21, 22, 9)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Alert noise burn-down campaign (depends on: 10, 12, 25)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
28. Wave rollout to all 28 teams with readiness gates (depends on: 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
29. Change management, fairness and pager culture (depends on: 4, 9, 25)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (depends on: 24, 26, 28)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 26, 28, 30)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 2 (ID: ddb335b2-98d9-46f8-afa0-88fca2e25749, Agent: gpt5.6-sol_refine_2, LLM: openai/gpt-5.6-sol):
Estimated Complexity: high
Success Metrics: - By day 7, every suspected major incident uses one incident record, one coordination channel, and a named commander.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes; at least 95% receive command within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and technical coverage.
- By week 18, all 180 services have an owner, tier, tested escalation path, and coverage appropriate to their risk.
- No mandatory night rotation starts before compensation, training, access, runbooks, and staffing controls are active.
- Every direct 24x7 technical rotation has at least six qualified responders or a documented, expiring executive exception.
- No responder is routinely assigned primary duty more often than one week in six or simultaneously assigned to two primary rotations.
- At least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of qualifying SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet their required internal and customer update cadence by month 3.
- Monthly human pages fall from 3,400 alert events to no more than 1,500 by day 90 and no more than 500 actionable pages by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, with no decline in Tier 0/1 detection coverage.
- Average after-hours load remains at or below two pages per responder per week; every sustained breach produces a remediation plan.
- All six monitoring sources route human pages through the single paging platform by week 12; direct legacy paging paths are disabled by week 18.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Monthly contracted availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- Quarterly on-call surveys reach at least 75% favorable responses on fairness, ownership boundaries, compensation, and sustainability by month 6.
- The month-7 mock audit finds no unowned high-risk control gap, and at least 95% of sampled incidents contain complete operating evidence.
Steps (26):
1. Establish the mandate, owner, funding, and schedule
Launch incident management as a **company operating program within 48 hours**. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Form a decision group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Approve the non-negotiables: one severity scale, one human-paging platform, one incident record, paid on-call, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, and incident actions.
- Fund tooling, compensation, training, exercises, and reliability work. Use the $1.3M in credits as the minimum financial comparison.
- Set milestones: interim process by day 7, policy and critical-path pilot by week 6, Tier 0/1 coverage by week 10, company rollout by week 18, internal audit in month 6, and mock audit in month 7.
2. Install a seven-day incident-response floor (depends on: 1)
Do not wait for policy design or tooling migration. Put a minimum process into operation immediately and begin retaining evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communication schedule.
- Create one monitored declaration path through chat, telephone, and the current paging environment.
- Use one incident channel, bridge, timeline document, and naming convention for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced on-call engineers and engineering leaders.
- Require a named commander within 10 minutes. The duty engineering director assumes command if nobody else does.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Compensate interim duty retroactively under the permanent policy.
- Hold a 15-minute daily operational review until the permanent process is live.
3. Create the factual, legal, and audit baseline (depends on: 1)
Build one defensible baseline for process design, executive decisions, and SOC 2 testing. Confirm the auditor's expected Type II observation period immediately.
- Reconstruct all 31 incidents from first customer impact through resolution, communications, credits, postmortem, and corrective action.
- Analyze the two incidents with unclear command minute by minute.
- Identify the missing detection signal for every customer-first incident.
- Inventory the six alert sources, 3,400 monthly alert events, noisy rules, duplicates, missing owners, and missing runbooks.
- Record the current rotations, unpaid work, after-hours load, and teams without coverage.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M credits, 85% noise, and 11 of 64 actions closed.
- Map incident-response controls to applicable SOC 2 criteria with Compliance and the auditor.
- Establish retention, confidentiality, legal-hold, and access requirements for incident evidence.
4. Co-design the fairness contract with engineers (depends on: 1)
Treat pager resistance as a valid design constraint. Make the employment and ownership bargain explicit before expanding on-call.
- Interview representatives from all 28 teams and the Support, Customer Success, Security, and Operations groups.
- Distinguish objections involving unpaid work, unfamiliar code, bad alerts, weak runbooks, sleep disruption, or blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for services they own or have formally accepted, supported by trained command, and given capacity to remove recurring defects.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Measure baseline trust, fairness, fatigue, and psychological safety. Repeat the survey at days 60 and 120, then quarterly.
5. Build the service catalog and criticality model (depends on: 3)
Make a machine-readable catalog the source of truth for routing, escalation, customer impact, and audit evidence. Every production service must have one accountable owner.
- Record the owning team, manager, product capability, escalation policy, communication channel, dashboard, runbook, dependencies, regions, and data stores for all 180 services.
- Map payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, and reporting to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing services whose failure can create material contractual impact.
- Classify Tier 2 as internal or deferrable services, and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional mode, recovery method, and contractual obligations for Tier 0 and Tier 1.
- Give orphan services an owner or approved decommission date within 30 days.
- Maintain separate but coordinated ownership for the ledger application and PostgreSQL platform.
6. Adopt one severity scale and incident lifecycle (depends on: 3, 5)
Use four impact-based levels. Classify on actual or credible customer, financial, security, regulatory, and contractual harm rather than organizational seniority.
- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; both regions impaired; or a core payment journey broadly unavailable. Page every role immediately, open a bridge, notify executives, publish customer status within 15 minutes when applicable, begin legal assessment, and require a postmortem.
- **SEV2 — major:** material payment degradation; settlement deadline at risk; a critical customer or material cohort unavailable; regional impairment with reduced resilience; or likely SLA breach. Page command and technical roles, publish customer status within 30 minutes when customer-visible, and require a postmortem.
- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads; page only when immediate action is necessary.
- **SEV4 — operational event:** no current customer impact and no urgent risk. Create a ticket and handle during normal operations.
- Treat percentages as supporting guardrails, never reasons to under-classify integrity, settlement, security, or contractual risk.
- Anyone may declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Automatically use at least SEV2 posture for suspected ledger-integrity events, cross-team incidents with unknown ownership, and materially unknown impact lasting 15 minutes.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed. Resolution requires stability, backlog recovery, and necessary reconciliation.
7. Define roles, authority, and handoffs (depends on: 6)
Separate command, communications, recordkeeping, and technical repair. One named person must hold command throughout every SEV1 and SEV2.
- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not serve as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- **Subject-Matter Responders:** diagnose and mitigate only services for which they have ownership, access, training, or formally accepted support responsibility.
- **Executive Duty Officer:** removes organizational barriers and makes exceptional business decisions without displacing the commander.
- Security, Legal, Compliance, Vendor Management, and Finance join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Require distinct commander, communications, scribe, and primary technical lead for SEV1. Communications and scribe may combine temporarily for bounded SEV2 incidents.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, and reconciliation requirements for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, with the exact transfer time.
8. Create sustainable 24x7 coverage across 28 teams (depends on: 4, 5, 7)
Use central command coverage and risk-based technical coverage rather than creating 28 fragile night rotations. All services receive a response path, but only critical domains maintain direct overnight technical rotations.
- Create a 24x7 Incident Command corps of 24–30 certified people, with primary and secondary coverage at all times.
- Create 18–24 trained Communications Leads and a similarly sized scribe pool using Support, Customer Operations, Engineering Operations, and qualified managers.
- Maintain a surge roster for simultaneous incidents and a 24x7 Executive Duty Officer schedule.
- Group Tier 0 and Tier 1 ownership into approximately 8–12 coherent domains only where responders have training, access, runbooks, and explicit acceptance.
- Staff each critical domain with primary and secondary responders and at least six qualified people. Target no more than one primary week in six.
- Give Tier 2 and Tier 3 services business-hours coverage plus a maintained manager and director escalation path.
- When lower-tier impact becomes SEV1 or SEV2, central command activates the manager escalation path and obtains the necessary owner.
- Route unknown-owner pages to Duty Command and platform triage temporarily. Log each one as a catalog control failure.
- Do not merge small teams merely for schedule convenience. Provide training, reassign services, add staffing, or decommission unsupported services.
9. Approve compensation and fatigue protections (depends on: 4, 8)
End unpaid on-call before expanding mandatory night coverage. HR, Finance, Payroll, and employment counsel should approve the policy within 14 days.
- Use market-validated weekly bands, initially budgeting approximately $800–$1,200 for Tier 0/1 primary duty and $300–$500 for secondary duty.
- Budget approximately $900–$1,300 for Duty Incident Commander weeks and $400–$800 for Communications Lead or scribe duty, adjusted for actual burden.
- Pay holiday premiums and compensate active after-hours work according to exempt or non-exempt status and applicable federal and New York rules.
- Provide a protected paid recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal sprint commitments by approximately 15% during a primary on-call week.
- Prohibit simultaneous primary assignments, consecutive primary weeks, on-call during leave, and invisible schedule swaps.
- Allow responders to declare themselves temporarily unfit after disruptive night work without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Budget roughly $0.9M–$1.2M annually, then refine using actual rotation count, employment classification, and activation data.
10. Establish one paging and incident system of record (depends on: 3, 5, 8)
Monitoring tools may remain specialized, but all human pages must enter one controlled platform. This removes conflicting schedules and creates one evidence trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page.
- Ingest events from the six existing monitoring tools before disabling their direct paging paths.
- Route pages through service-catalog ownership and deduplicate related events.
- Provide a single declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, severity changes, decisions, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status-page publishing, and conference facilities.
- Require role-based access, MFA, access reviews, and immutable or tamper-evident history.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, or an AWS region.
- Retire each legacy human-paging path only after ownership review, end-to-end tests, and two weeks of verified operation.
11. Enforce an alert-quality contract and page budget (depends on: 3, 5, 10)
Treat every page as a production interface with an owner and an expected action. Noise reduction must not create detection gaps.
- Require every paging rule to identify the service, owning team, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Page only when prompt human judgment or intervention can materially reduce customer, financial, security, or contractual risk.
- Route informational, capacity, and non-urgent infrastructure conditions to dashboards or ticket queues.
- Prefer payment-outcome, settlement-risk, queue-age, and error-budget-burn alerts over raw CPU, memory, pod, or log thresholds.
- Run new paging rules in shadow mode for at least seven days unless an emergency exception is approved.
- Review rules with repeated no-action acknowledgements, low actionability, or excessive firing within two business days.
- Set a load budget of no more than two after-hours pages per responder per week, measured over four weeks.
- Require an owner, compensating detection, and recorded approval before suppressing or deleting a rule.
- Burn down the top 50 noisy rules first. Review missed detections and noise together so teams cannot improve metrics by becoming blind.
12. Detect payment and ledger failures before customers (depends on: 5, 11)
Move detection from host health to customer journeys and financial outcomes. Use internal SLOs with enough headroom to protect the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, API access, webhooks, and reporting freshness.
- Run external synthetic transactions through critical payment journeys at least every minute and from paths independent of the production platform.
- Validate each AWS region and expose dependencies that undermine nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed value relative to settlement deadlines, not only queue depth.
- Add tenant and cohort anomaly detection for high-value customers and major payment methods.
- Convert high-priority Support, account-manager, processor, sponsor-bank, and network reports into incident candidates within five minutes.
- Record the detection source for every incident. Treat customer-first detection as a mandatory missed-detection review.
13. Codify detection, escalation, and live execution (depends on: 7, 10)
Create one time-bound path from the first credible signal to named command and mitigation. Notification delivery does not count as human acknowledgement.
- Page the owning critical-domain primary and Duty Incident Commander immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at five minutes, manager at 10, and director at 15.
- Escalate an unclaimed command page to backup commander at five minutes. The Executive Duty Officer assumes command at 10 minutes until a certified handoff occurs.
- Require Support and account managers to use the same declaration path as automated monitoring and engineers.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Begin with a standard statement of severity, known impact, assigned roles, current objective, workstreams, and next update time.
- Freeze unrelated changes during SEV1 and normally during SEV2. Record any exception.
- Prefer reversible mitigation such as rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require formal command handoff for long incidents, shift changes, or fatigue. Do not leave an incident unowned during transfer.
- Reconcile the ledger and safely drain backlogs before resolving payment or ledger incidents.
14. Standardize internal incident communications (depends on: 7, 10, 13)
Give responders one working room and stakeholders one controlled information source. Executives must not interrupt the technical command path.
- Create one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue an initial internal brief within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within 15 minutes for SEV1 and 30 minutes for SEV2.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record material decisions and outbound messages in the incident timeline.
15. Standardize status-page and customer communications (depends on: 6, 10, 14)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate impact before root cause is known.
- Publish customer status within 15 minutes of declaring a customer-visible SEV1 and within 30 minutes for customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice promptly after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map status-page components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, data-integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, known workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Give affected strategic accounts direct account-manager outreach within 30 minutes for SEV1 and 60 minutes for SEV2.
- Require account managers to use the approved briefing and prohibit independent technical explanations.
- Provide a customer-facing incident summary within five business days for SEV1 and qualifying SEV2 events.
- Record any legally necessary delay or restriction of public detail, its approver, and the alternative communication plan.
16. Operationalize legal, regulatory, contractual, and credit decisions (depends on: 3, 6, 15)
Some payments incidents start external notification clocks. Make assessment mandatory without assuming that every operational incident is reportable.
- Build a counsel-validated matrix covering applicable NYDFS rules, state breach laws, GLBA or FTC obligations, PCI requirements, money-transmitter obligations, sponsor-bank and network contracts, cyber insurance, and customer contracts.
- Engage Legal and Compliance immediately for every SEV1, security event, and suspected ledger-integrity event.
- Record an initial reportability assessment within one hour for SEV1 and within two hours for other potentially reportable events.
- Record non-reportable decisions as evidence, including facts considered, approver, timestamp, and reassessment trigger.
- Maintain tested 24x7 contacts for counsel, regulators, sponsor banks, networks, insurers, and critical vendors.
- Encode customer-specific notice deadlines and channels in the customer record used by the Communications Lead.
- Have Legal own regulatory text and submission. Keep technical command with the Incident Commander.
- Have Finance calculate affected minutes, delayed value, likely credits, and contractual exposure within five business days.
- Track credits by incident and recurring cause to support reliability investment decisions.
17. Make postmortems mandatory, consistent, and blameless (depends on: 6, 7, 10)
Use one learning standard with fixed deadlines. Keep postmortems separate from performance, misconduct, and disciplinary processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact over two hours, SLA credits, contractual breach, repeated contributing factors, control failures, and ledger-integrity near misses.
- Produce the factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the Incident Commander responsible for the timeline and the owning engineering director accountable for completion.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit analysis of why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the incident commander.
- Describe decisions in the context and information available at the time. Do not name an individual as the root cause.
- Publish broadly useful findings internally while restricting security, privacy, personnel, or privileged material appropriately.
18. Give corrective actions enforceable ownership (depends on: 10, 17)
Treat incident actions as risk commitments, not suggestions. A ticket is not complete until the expected risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Escalate overdue high-risk actions to the manager after seven days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit overdue actions to block related releases where the unrepaired condition could reproduce severe impact.
- Verify effectiveness through tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days. Complete, re-plan, or formally accept each risk.
19. Set the service-readiness bar and critical playbooks (depends on: 5, 8, 11, 12)
A critical service must be supportable at 3 a.m. before its team is placed on direct overnight coverage. Existing critical detection must remain active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and data-integrity constraints.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional failure, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document controlled failover, split-brain prevention, replay protection, backlog recovery, and post-recovery reconciliation.
- Exercise critical runbooks at least twice per year and after material changes.
- Block new Tier 0/1 releases and new paging rules when readiness requirements are missing.
- Handle existing gaps through a named owner, compensating control, executive-approved expiry date, and remediation plan rather than disabling detection.
20. Publish policy and certify every response role (depends on: 6, 7, 8, 9, 11, 13, 14, 15, 16, 17, 18)
Convert the operating design into concise, signed documents and practical training. Training and exercises occur during paid working time.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, and exception standards.
- Provide one-page severity, role, authority, escalation, and communication cards inside the incident tool.
- Train all employees to recognize and declare incidents.
- Train all engineers in severity, acknowledgement, evidence preservation, handoff, and financial-integrity precautions.
- Certify responders only after they demonstrate access, dashboards, runbooks, rollback, and escalation competence.
- Certify Incident Commanders through formal instruction, simulation, and at least two shadowed incidents or exercises.
- Train Communications Leads in status writing, account segmentation, contractual clocks, and legal escalation.
- Train scribes in timeline quality and fact-versus-hypothesis labeling.
- Require two shadow shifts before independent primary duty and annual recertification.
- Maintain the training and certification register as operational and audit evidence.
21. Pilot on the payment critical path (depends on: 9, 10, 12, 19, 20)
Run a four-to-six-week pilot across the highest-risk journey before expanding. Use real incidents and exercises to correct the model quickly.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, API edge, authentication, settlement, and Support intake.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently assuming command.
- Correct critical process or tool defects within 48 hours and update the standard visibly.
- Exit only after 95% timely command assignment, 95% communication compliance, no unpaid pages, complete required postmortems, and at least 50% lower noise.
22. Roll out by customer journey and risk (depends on: 21)
Expand in controlled waves while the interim process remains active company-wide. Complete rollout early enough to accumulate operating evidence before the audit.
- Roll out remaining Tier 0 domains first, then Tier 1, Tier 2, and Tier 3.
- Use two-to-three-week waves with a named program coach and director sign-off.
- Gate each team on catalog ownership, appropriate coverage, compensation, trained responders, tested escalation, alert quality, runbooks, access, and a passed tabletop.
- Require six or more responders only for direct 24x7 technical rotations. Apply business-hours coverage and manager escalation to lower tiers.
- Reschedule failed gates or approve a time-limited executive exception with a compensating control.
- Disable legacy human-paging paths after verified cutover for each wave.
- Publish an internal adoption dashboard by team, service tier, and control gap.
- Finish critical coverage by week 10 and all 28 teams by week 18.
23. Exercise command, communications, regional recovery, and fallbacks (depends on: 19, 20)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous security and operational events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors where appropriate.
- Conduct at least one unannounced after-hours paging test before the audit.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure time to acknowledgement, command, customer notice, mitigation decision, handoff, and recovery.
- Create tracked actions for every material exercise finding.
24. Measure performance and operate fixed review forums (depends on: 3, 10, 17, 18)
Use a balanced scorecard that exposes weak controls without rewarding hidden incidents or suppressed alerts. Report medians and 90th percentiles rather than averages alone.
- Measure time from first impact to detection, declaration, acknowledgement, command assignment, mitigation, recovery, and resolution.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track pages, actionability, duplicates, after-hours load, missed detections, and page-budget breaches.
- Track postmortem timeliness, action age, closure by due date, verified effectiveness, and recurring contributing factors.
- Track availability by customer journey, error-budget burn, failed or delayed payment value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight activations, recovery days, schedule exceptions, sentiment, and attrition.
- Hold a weekly Incident Review Board for incidents, postmortems, control failures, noisy alerts, and overdue actions.
- Hold a monthly executive reliability review for trends, funding, contractual exposure, and accepted risks.
- Hold a quarterly control and resilience review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile incident records monthly against support cases, status history, credits, and customer complaints to detect under-reporting.
25. Prove SOC 2 operating effectiveness before fieldwork (depends on: 20, 22, 23, 24)
Generate audit evidence through normal operation rather than reconstructing it later. Test both control design and consistent execution.
- Maintain approved and versioned policies, exceptions, catalog records, schedules, compensation activation, access reviews, training, incidents, communications, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, non-reportable legal decisions, missed timelines, and exercises in the population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the populations, evidence requests, and interviews expected from the external auditor.
- Record deviations honestly with owners, remediation dates, and compensating controls. Never rewrite historical records.
- Confirm that evidence retention covers the full auditor-defined observation period.
- Brief commanders, responders, Support, and Compliance on the real process without scripting inaccurate answers.
26. Institutionalize improvement and reduce structural risk (depends on: 22, 24, 25)
Prevent the program from decaying after the audit. Use incident evidence to drive permanent ownership and architectural investment.
- Assign permanent owners for policy, service catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review severity thresholds, staffing, compensation, and communication timings annually and after material process failures.
- Recertify command and communications personnel annually.
- Review recurring failure families quarterly and require executive decisions where corrective work repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fairness, fatigue, and tool friction.
- Report severe incidents, credits, overdue high-risk actions, and resilience investment to the board or risk committee quarterly.
- Treat the shared ledger cluster as a strategic concentration risk. Fund failover assurance, blast-radius reduction, isolation of non-critical readers, and stronger regional independence.
- Evaluate follow-the-sun command or technical coverage using one year of page-load and staffing data.
- Build year-two plans for automated mitigation, deployment safety, graceful degradation, and error-budget release controls.
Previous Proposal 3 (ID: f7ee763b-3b76-4f07-9e7d-5164701bef1c, Agent: qwen3.8-max_refine_3, LLM: alibaba/qwen3.8-max):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (32):
1. Establish executive mandate, program office, and funding
Turn the CEO email into a **chartered company program** with one accountable owner and authority over all 28 teams.
- Appoint the CTO as executive sponsor and a Head of Reliability as program owner with full-time authority.
- Create a permanent program office: one program lead, one platform engineer, one analyst.
- Form an eight-person steering group: Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours.
- Lock the non-negotiables: one severity scale, one paging platform, one postmortem format, paid on-call, mandatory action tracking, named service ownership.
- Approve budget anchored against the $1.3M in credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs.
- Reserve 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Publish a one-page charter stating incident response is a company operating process, not a per-team choice.
2. Install a seven-day interim command bridge (depends on: 1)
Do not leave the company unprotected while the permanent process is designed. Put a **crude but real** command structure in place within one week.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Direct Support to escalate credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical actions: complete, re-plan, or formally risk-accept, prioritizing ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Build the forensic baseline of incidents, alerts, and money lost (depends on: 1)
Rebuild the facts before designing anything. This becomes both the **design input** and the before picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific missing signal. This list drives the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Freeze baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the on-call fairness deal (depends on: 1)
Engineer pushback is the **largest delivery risk**. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS, and Sales in two weeks. Separate the real objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey covering fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
5. Map compliance, evidence, and audit requirements from day one (depends on: 1)
Design evidence as a **by-product of operations**, not a reconstruction before the auditor arrives. Confirm the SOC 2 observation window immediately.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for communications; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure, training register, drill records, reportability decisions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve all policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has a **named owner**. Build a machine-readable catalog as the single source of truth.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact: Tier 0 (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for each Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
7. Adopt the severity scale, declaration rules, and lifecycle (depends on: 3, 6)
Replace judgment calls with a **lookup table**. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter.
- SEV1 (crisis): money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of all roles and executive; bridge in 5 minutes; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 (critical): material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- SEV3 (major, contained): narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected, Declared, Triaged, Mitigated (customer impact ends), Monitoring, Resolved (backlog processed and ledger reconciled), Reviewed.
- Publish a decision tree with 12 worked examples from the real 31 incidents.
8. Define incident roles, authority, and handover discipline (depends on: 7)
Solve nobody-in-charge-for-an-hour by making command **explicit, single-holder, transferable, and logged**. Separate coordination from debugging.
- Incident Commander: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates for SEV1/SEV2.
- Subject-Matter Responders: engineers of the owning team; they mitigate, they do not run the room.
- Executive Duty Officer (SEV1): removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Design the three-layer 24x7 coverage model (depends on: 6, 8)
Do not create 28 night rotations. **Centralise coordination** in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- Layer A, Incident Command corps: approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of approximately 18 from Support, CS, and engineering management. A scribe pool used as the training entry point.
- Layer B, critical-path domain rotations: consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- Layer C, everyone else: business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing: a US-only paid night rotation now, a follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
10. Approve on-call compensation, labor compliance, and fatigue safeguards (depends on: 9, 4)
Unpaid on-call in New York is both a **retention problem and a legal exposure**. Pay must be live in payroll before any mandatory night rotation starts.
- Indicative scheme: approximately $1,000 per primary 24x7 week, $400 secondary, $250 for business-hours rotations, a separate $1,200 Duty Commander stipend, holiday premiums, approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- HR, Finance, and employment counsel publish amounts, eligibility, tax treatment, and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a 15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
11. Set the alert quality standard, page budget, and noise burn-down (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes. Make alert quality a **condition of being allowed to page a human**.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly.
- Target 3,400 to under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
12. Build payment-outcome detection and ledger assurance (depends on: 11, 6)
Stop customers telling you first. Detection must be driven by **payment outcomes and ledger truth**, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer detected first becomes a named defect class with a mandatory tracked action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Consolidate to one paging platform, one incident record, one status page (depends on: 7, 9, 11)
Collapse six alerting tools into a **single operational system of record** so there is one queue, one timeline, and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is unavailable. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
14. Codify the escalation ladder and five-minute command rule (depends on: 8, 9, 13)
Write one unskippable path from something looks wrong to **someone is in charge**. The default action is never waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately, secondary at 5 minutes, domain manager and Duty Commander at 10, Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension, and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
15. Write the live incident execution doctrine and major-incident playbooks (depends on: 8, 13, 14)
Give responders one short operating procedure from the first minutes through closure. Priority is **limiting customer and financial harm**, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the single largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window, an explicit handback to the owning team, Support, and Customer Success, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
16. Standardize internal communications protocol (depends on: 8, 13)
Standardise the internal picture so executives, Support, and Sales are informed **without interrupting** the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
17. Build customer communications, status page, and account-manager outreach (depends on: 7, 16)
Replace whoever is around with a **timed, owned, pre-approved process**. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook, and RSS. All 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity, or blame.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
18. Create the regulatory, partner, and legal notification playbook (depends on: 7, 17)
In payments, some incidents start a **legal clock at detection**. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded even when the answer is not reportable, with evidence, decision-maker, and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
19. Operationalize SLA credit and financial impact workflow (depends on: 7, 17)
Tie incidents to money so severity, credits, and investment decisions **stay honest**, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one. Use the delta as the standing business case for on-call pay and reliability capacity.
20. Establish the blameless postmortem standard and Incident Review Board (depends on: 7, 8)
Replace some incidents, various formats with **one mandatory format, fixed deadlines, and a forum with teeth**.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly top-five recurring causes analysis.
21. Enforce action ownership, capacity reservation, and tracking (depends on: 20, 13)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of **customer commitments** and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
22. Build the training and certification academy (depends on: 8, 14, 16, 20)
Command is a skill, not a title. **Certify before assigning duty**, and use paid working time for all of it.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days + shadowing): command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Comms Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum; never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Run the exercise programme: tabletops, game days, and unannounced drills (depends on: 22, 13, 15)
The process must meet a **simulated SEV1** before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger. Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (depends on: 7, 8, 9, 10, 11, 14, 16, 17, 18, 20, 21)
Collapse the design into a document people will **actually open mid-outage**, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 12, 13, 15, 22, 10)
Prove the process on the **highest-risk surface** with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment.
- Publish a one-page result to the whole company.
26. Run the alert noise burn-down campaign (depends on: 11, 13, 25)
Run noise reduction as a **visible, quota-driven campaign** in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness. Review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 to 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
27. Establish metrics, dashboards, review cadence, and anti-gaming (depends on: 13, 20, 25)
Instrument the process itself so improvement is visible and the auditor sees **evidence of monitoring and review**. Report median and 90th percentile, never averages alone.
- Response: time to detect, declare, commander, acknowledgement, mitigate, resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
28. Wave rollout to all 28 teams with readiness gates (depends on: 25, 27)
Roll out in four waves of six to eight teams every two to three weeks, ordered by **customer risk**. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. Failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Publish a live adoption scoreboard.
29. Run change management, fairness, and pager culture programme (depends on: 4, 10, 25)
Run this from day one in parallel. Engineers judge the process on **fairness**; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. Build SOC 2 evidence by design, internal testing, and mock audit (depends on: 24, 27, 28)
Design the evidence as a **by-product of doing the work**. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any change as an exception.
31. Maintain the program risk register and contingencies (depends on: 1)
Name the ways this programme fails and **pre-commit the response**. Review it monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 27, 28, 30)
Guard against the classic failure: the process **decays once the audit is signed**. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 4 (ID: 101b9773-a36a-47bf-8af6-14fe9a48abb2, Agent: grok4.6_refine_4, LLM: xai/grok-4.6):
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (26):
1. Charter the program and start the audit clock
Convert the CEO email into a company operating process with one owner, a budget, and an observation window that starts this week.
Incident response is no longer a per-team choice.
- Name the CTO as sponsor and a Head of Reliability as the single accountable owner, with a three-person program office.
- Form a small decision group: Engineering, SRE, Support, Customer Success, Security, Legal, Finance, and HR. The sponsor decides within 48 hours.
- Lock non-negotiables: one severity scale, one paging path, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, and reserved engineering capacity against the $1.3M in SLA credits.
- Timeline: operating floor in 7 days, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, mock audit in month 7, audit in month 8.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for new tools or the final policy. Put a crude but real process in place this week so the next outage has a named commander.
This week is the first audit evidence.
- Publish a one-page interim severity card and one declaration path: Slack command, phone number, and existing pagers, all reaching the same duty person.
- Staff interim primary and backup Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in the incident channel.
- Tell Support to declare from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline (depends on: 1)
Rebuild the facts before locking design. This is both the design input and the before picture for the CEO and the auditor.
Freeze the numbers so they cannot drift during design.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisy rules and every rule with no owner or runbook.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11 of 64 actions closed, on-call in 12 of 28 teams.
4. Map resistance and publish the fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
Publish the deal in writing before any new mandatory pager is assigned.
- Interview all 28 teams plus Support, Customer Success, and Sales in two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the **fairness contract**: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and money-path tiers (depends on: 3)
You cannot page the right person across 180 services until every service has one owner. A wrong owner recreates the pager objection.
The catalog is the single source of truth for paging, impact, status-page components, and audit.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to services, data stores, regions, and third parties, including sponsor banks and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
6. Lock severity levels and what each one triggers (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by who reported it or how hard the fix looks.
Anyone may declare. Nobody is punished for over-declaring. Only the Incident Commander may downgrade, with the evidence recorded.
- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 major, contained: narrow or single-customer impact with a workaround. Owning team leads; business-hours comms; postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages.
- Attach a financial-integrity or security flag to any severity. The flag forces dual-control, Legal, and the regulatory checkpoint without inventing a fifth level.
- Auto-escalate: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer impact ends) → Monitoring → Resolved (backlog processed and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged.
Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. A human validates for SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1 and SEV2. Every handover is announced verbally and in writing with the exact time.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
8. Staff 24x7 with a command corps, not 28 night rotations (depends on: 5, 7)
Do not create 28 night rotations. That is what engineers are rejecting.
Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified volunteers from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to seven months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate 28 teams into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat Lisbon or APAC follow-the-sun as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
9. Pay on-call, meet New York labor rules, and cap fatigue (depends on: 4, 8)
Unpaid on-call in New York is a retention problem and a legal exposure. Pay must be in payroll before any new mandatory night rotation starts.
Publish the numbers. Then ask people to sign up.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
10. Set the alert-quality bar and a hard page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them.
Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers do (depends on: 5, 10)
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail.
Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One Slack command creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is down. Test weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Write the five-minute path from signal to command (depends on: 7, 8, 12)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting.
If nobody claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Codify live execution, readiness, and ledger playbooks (depends on: 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving a root cause.
A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down. No command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers.
15. Run one communications clock for internals, customers, and regulators (depends on: 6, 7, 12)
Replace whoever is around with one timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is not reportable.
- Obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
16. Tie incidents to SLA credits and true financial cost (depends on: 6, 15)
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
Credit calculation is an output of the incident record, not a negotiation.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
17. Make blameless postmortems mandatory and consistent (depends on: 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth.
The discipline lives in the deadlines and the review, not the template.
- Mandatory for every SEV1 and SEV2, any incident detected by a customer first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed.
18. Track actions as risk commitments with reserved capacity (depends on: 12, 17)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
Closing a ticket without evidence of effectiveness does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Target 90% of high-priority actions closed by due date within two quarters.
19. Train and certify every role before independent duty (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
The certification register is an audit artefact.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
20. Rehearse with tabletops, game days, and night drills (depends on: 12, 14, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
21. Publish Policy v1 and the signed fairness contract (depends on: 6, 7, 8, 9, 10, 13, 15, 17, 18)
Collapse the design into a document people will actually open mid-outage, and make it official.
Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
22. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 19, 21)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse. Publish a one-page result to the whole company.
23. Instrument metrics, reviews, and anti-gaming (depends on: 12, 17, 22)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review.
Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
24. Roll out in risk-ordered waves with readiness gates (depends on: 22, 23)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave, not as a background hope. Give every team its noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
25. Produce SOC 2 evidence by operating, then mock-audit (depends on: 21, 23, 24)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions. Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (depends on: 23, 24, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
Name the ways this programme fails and pre-commit the response.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
Previous Proposal 5 (ID: 289def1b-8f78-402b-85b3-0d9e18661842, Agent: deepseek-v4-pro_refine_5, LLM: deepseek/deepseek-v4-pro):
Estimated Complexity: high
Success Metrics: - Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, no loss of Tier 0/1 detection.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% Tier 0/1 services have named owning team, tier, escalation policy, dashboard and runbook by day 30; all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10-12 domain rotations of at least six trained responders.
- 30 certified incident commanders and 18 certified communications leads active.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% required postmortems drafted in 3 business days, reviewed in 5, published in 10 from month 2.
- Postmortem action closure rises from 17% to 90% closed by due date within two quarters; all 53 legacy actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k annualised within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% SEV1/security SEV2 including not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two cross-company exercises completed before audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% sampled incidents; SOC 2 Type II incident response controls pass zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days; no increase in attrition among responders.
Steps (26):
1. Executive mandate, budget, and governance
Secure a written CTO/CEO charter making incident management a company operating process, not a team option. Name one accountable Head of Reliability and a small steering group with authority to decide within 48 hours.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking, and protected engineering capacity.
- Approve budget against the $1.3M annual credits: tooling $150-250k/yr, on-call compensation, training, and 2-3 program FTEs.
- Set timeline: interim process in week 1, design weeks 1-6, pilot weeks 7-12, rollout weeks 13-24, mock audit month 7, SOC 2 at month 8.
- Freeze current baselines: 31 incidents, 22-min MTTD, 40% customer-first detection, 3h10 MTTM, $1.3M credits, 3,400 alerts/month at 85% noise, 11/64 actions closed.
2. Interim 7-day command (depends on: 1)
Put a crude but real process in place immediately so no incident remains unowned while the permanent design proceeds.
- Publish a one-page interim severity card, one declaration path (phone, Slack command, pager), and one incident channel/bridge/timeline naming convention.
- Staff an interim 24x7 duty commander with primary and backup from engineering managers and senior SREs; compensate retroactively under final policy.
- Require a named incident commander within 10 minutes of any suspected major incident, announced in channel.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions in 30 days, prioritizing ledger integrity, duplicate payment, regional failover, security and detection gaps.
- Hold a daily 15-minute operations review until the permanent process is live.
3. Forensic baseline and alert estate analysis (depends on: 1)
Reconstruct the true before picture from the last 12 months; it drives design and serves as audit baseline.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact start/detect/declare/commander/mitigate/resolve, who led, credits paid, contributing factors.
- For every customer-first detection, name the missing signal; this becomes the detection backlog.
- Reconstruct the two `nobody in charge` incidents minute by minute; use them as the burning-platform narrative and test case.
- Profile the 3,400 monthly alerts by tool, team, rule, outcome; identify top 50 noisy rules and every rule with no owner or runbook.
- Freeze all baseline metrics in one signed document for the executive and the auditor.
4. Listening tour and fairness contract (depends on: 1, 3)
Treat engineer pushback as the main delivery risk and convert it into a written deal that removes the objection to carrying a pager for other teams' code.
- Interview all 28 teams plus Support, CS and Sales in two weeks to separate objections: unpaid work, nights, unfamiliar code, missing runbooks, or fear of blame.
- Harvest practices from the 12 teams already on-call; they supply pilot teams and first commanders.
- Publish the deal: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Recruit 10-15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey and repeat at 60, 120 and 365 days.
5. Service ownership catalog and criticality tiers (depends on: 3)
Create a machine-readable catalog as the single source of truth for paging, routing, status page components, and audit evidence.
- Assign one accountable team, engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list per service.
- Tier by business impact: Tier 0 money movement/ledger/auth/shared PostgreSQL; Tier 1 customer-facing degradable; Tier 2 internal/batch; Tier 3 non-critical.
- Map customer journeys (initiate, authorize, settle, reconcile, report, onboard) to services, data stores, regions, third parties.
- Give orphan services an owner within 30 days or a decommission date approved by the sponsor; Tier 0 without an owner is an executive escalation.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
6. Severity scale, declaration rules, and automatic triggers (depends on: 3, 5)
Adopt one severity scale as a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade with recorded rationale. Start high when uncertain.
- SEV1: money moved wrongly/duplicated/lost, ledger integrity in doubt, confirmed security/data exposure, both regions impaired, payment processing halted. Full role page, bridge in 5 min, status page in 15 min, regulator assessment within 1 hour, mandatory postmortem.
- SEV2: material degradation of a payment journey, settlement window at risk, large/strategic customer fully down, SLA breach likely. Commander and SMEs paged, status page in 30 min, mandatory postmortem.
- SEV3: narrow or single-customer impact with workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any ledger cluster incident, any SEV3 open >2h, any unknown impact after 15 min, any cross-team incident become at least SEV2.
- Publish a decision tree with 12 worked examples from the 31 real incidents.
7. Incident roles, authority, and handoffs (depends on: 6)
Codify roles so coordination never depends on seniority or heroics. The Incident Commander coordinates; responders fix only services they own.
- IC: owns severity, priorities, roles, cadence, closure; pre-authorized to freeze deploys, rollback, disable features, shift traffic, invoke failover, put ledger read-only, commit spend; does not type in terminals; keeps command when VP joins.
- Communications Lead: sole author for status page, account managers, executives, and hand-off to Legal for regulators; speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, state changes; human validates for SEV1/SEV2.
- Subject-Matter Responders: diagnose and mitigate only their owned services.
- Executive Duty Officer (SEV1): removes obstacles, shields IC from exec questions, owns board/regulator escalation; does not take command unless formally transferred.
- Rule: IC claimed within 5 min and announced; distinct people for IC, Comms, technical lead at SEV1/SEV2; every handover announced verbally and in writing with time.
- Financial controls survive: IC coordinates ledger recovery but cannot bypass dual control, reconciliation or privileged access.
8. 24x7 command and responder coverage model (depends on: 5, 7)
Do not create 28 fragile night rotations. Centralize coordination in a trained command corps, keep technical ownership local.
- Layer A: incident command corps of ~30 certified volunteers with primary/secondary 24x7; paired Comms Lead pool ~18 and scribe pool as entry.
- Layer B: consolidate 28 teams into 10-12 critical-path domains (ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes, data/reporting, integrations) each with 24x7 primary+secondary and minimum six trained responders.
- Layer C: all other teams business-hours on-call with written after-hours escalation lists held by managers.
- Platform on-call is safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate follow-the-sun only as a 12-month option, not year-1 dependency.
- Acknowledgement discipline: 5 min at SEV1/SEV2 with automatic failover.
9. Paid on-call and fatigue safeguards (depends on: 8, 4)
Unpaid on-call in New York is a retention and legal risk. Pay must be in payroll before any mandatory night rotation starts.
- Indicative: ~$1,000 per primary 24x7 week, secondary ~$400, business-hours ~$250, duty commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond one hour.
- Mandatory recovery: paid recovery day after >2h overnight work, SEV1, or qualifying SEV2; managers arrange coverage.
- HR/Legal/Finance publish amounts, tax treatment, FLSA/NY wage-hour status within 14 days.
- Load rules: no primary more often than 1 in 6, no consecutive primary/secondary weeks, no primary on two rotations, no on-call during PTO, annual cap.
- If a rotation averages >2 out-of-hours pages per person per week for 4 weeks, trigger staffing or alert-remediation review.
- Count on-call as ~15% delivery load; document exemption path for health/caring without career penalty.
10. Service readiness bar and runbooks (depends on: 5, 8)
A service must earn the right to page a human at 3 a.m. Define a minimum readiness bar and major incident playbooks before any Tier 0/1 service goes live with on-call.
- Require for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switch, escalation contacts, data-loss/latency impact.
- Write playbooks for top failures: shared PostgreSQL failure/corruption, cross-region failover, Kubernetes control-plane loss, processor/sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, duplicate payment.
- Treat ledger cluster as largest structural risk: documented failover, read-only degraded mode, reconciliation after recovery, executive-signed RPO/RTO.
- Start a parallel workstream on blast-radius reduction: tenant/function partitioning, read replicas, isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging at night unless manager accepts a dated exception in writing.
- Runbooks are peer-reviewed, version-controlled, marked stale if not exercised twice a year.
11. Alert quality standard and page budget (depends on: 3, 5)
Make alert quality a condition of paging a human. Burn down noise deliberately rather than by mass silencing.
- Every paging alert must declare owning team, affected service, customer/SLO impact, expected action, dashboard, runbook, dedup key, severity mapping, escalation policy. Anything failing becomes a ticket.
- Page on symptoms of customer harm (SLO burn rate, payment success rate, queue age vs settlement deadlines), not raw CPU/memory.
- Run new alerts in shadow mode 7 days; test firing and recovery.
- Set page budget of 2 out-of-hours pages per person per week; breach blocks new alert creation and triggers tuning sprint.
- Auto-quarantine rules firing >5 times/month without action or >70% no-action acks; return to owner with correction deadline.
- Never silently disable: verify compensating detection, record decision, name owner, review missed detections monthly.
- Target 3,400 to <500 pages/month, actionability >75% within six months, no loss of Tier 0/1 detection.
12. Detection uplift on money path (depends on: 5, 11)
Stop customers telling you first by detecting on payment outcomes and ledger truth.
- Define SLOs and business SLIs per customer journey: initiation success, authorization latency, settlement timeliness, reconciliation break rate, API availability, reporting freshness; set internal targets stricter than 99.95%.
- Run external synthetic end-to-end payments every 60 seconds from both regions, covering all critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate detection, replication lag, failover-readiness, settlement-window countdown.
- Add per-customer anomaly detection for top 100 accounts; five similar support tickets in 10 minutes auto-creates triage incident.
- Route partner and processor notifications into declaration path within 5 minutes.
- Record detection source on every incident; treat `customer detected first` as a named defect class with mandatory action.
13. Consolidate to one paging and incident platform (depends on: 6, 8, 11, 12)
Collapse six alerting tools into one operational system of record: one queue, one timeline, one audit trail, without creating a monitoring gap.
- Select one paging/scheduling platform, one incident record, one hosted status page; time-box selection to 2 weeks.
- Ingest all six sources first, deduplicate/correlate; retire legacy path only after named owners, successful end-to-end test, and 2 weeks verified operation.
- One-command declaration in Slack creates channel/bridge, pages commander, sets severity, opens timeline, starts clock.
- Route every page through service catalog: service label -> owning domain schedule -> escalation policy.
- Capture evidence automatically: declaration, acks, role assignments, severity changes, decisions, comms, mitigated/resolved times, postmortem link; retain 12 months.
- Test out-of-band SMS/phone paging, mobile fallback, offline runbook weekly; ensure works during one AWS region/chat/identity provider failure.
- Set hard date after which pages outside this tool create no on-call obligation.
14. Escalation paths and acknowledgement SLAs (depends on: 7, 8, 13)
Write one unskippable path from signal to named commander in under 5 minutes; default action is never waiting.
- Converge all entry points on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately -> secondary at 5 min -> domain manager + Duty Commander at 10 -> Executive Duty Officer at 15. SEV2 requires command assigned within 15 min.
- If no one claims command in 5 min, platform assigns and announces it; assignee may hand over but not decline.
- Commander may page any domain's on-call directly with 10-min ack obligation; this reciprocity makes single-team ownership viable.
- If ownership unclear after 10 min, commander keeps incident and names temporary owner; missing catalog entry logged as control defect.
- Pre-authorize regional failover, ledger read-only, payment suspension, partner-bank notification so IC never waits for an executive.
15. Standardize internal and customer communications (depends on: 6, 7, 13)
Replace whoever is around with a timed, owned, pre-approved process; the Communications Lead is single author and never writes from scratch under pressure.
- Internal: one working channel/bridge plus one read-only broadcast for executives/Support/Sales; cadence SEV1 15-30 min, SEV2 60 min even if nothing changed; fixed template with impact, action, next update, IC/Comms names.
- Executives ask questions only of Executive Duty Officer; publish this as signed behavior rule.
- Status page: SEV1 within 15 min, SEV2 within 30 min; updates every 30/60 min; resolution notice within 30 min of verified recovery; customer-facing summary within 5 business days for SEV1/qualifying SEV2.
- Pre-approve 12-15 templates with Legal for degradation, settlement delay, API errors, regional failure, data-integrity investigation, security event.
- Top 100 accounts get named account manager call/email within 30 min of SEV1 with briefing pack; all 2,100 subscribed by default.
- Language rules: state impact and next update; never speculate on cause, recovery time, data integrity or blame.
16. Regulatory, partner, and financial impact workflows (depends on: 6, 15)
In payments some incidents start a legal clock at detection. Build obligation assessment into the process and tie incidents to money.
- Legal/Compliance produce obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC, PCI, money-transmitter, sponsor-bank/card-network windows, cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and security-related SEV2 within 1 hour, recorded even when `not reportable`, with decision-maker and evidence.
- Maintain 24x7 contact matrix for regulators, sponsor banks, networks, outside counsel, insurer.
- Agree availability measurement method per contract; compute affected minutes per customer from incident record and journey telemetry; propose credit schedule within 5 business days.
- Attribute credits to root-cause families; track failed payment volume, delayed value, reconciliation breaks; target annual credits from $1.3M to under $400k.
17. Blameless postmortem standard and action tracking (depends on: 6, 7)
Replace various formats with one mandatory format, fixed deadlines, and enforceable actions. Closing a ticket without evidence does not close the action.
- Mandatory for every SEV1/SEV2, customer-first detection, incident >2h, repeat of known cause, credit-generating/contract breach, ledger near-miss.
- Draft within 3 business days, review within 5, publish within 10; IC owns delivery, owning manager accountable.
- One template: summary, customer/financial impact, detection source/gap, timeline, response analysis, contributing conditions, what worked, actions.
- Two mandatory questions: why did a customer see this first? why did mitigation take as long as it did?
- Blameless in writing: systems and context, no individual named as cause, never used in performance reviews; HR handles personnel/misconduct separately.
- Every action gets owner, priority, due date, verification method, ticket auto-created; classes containment 7 days, corrective 30, strategic 90; reserve 15% engineering capacity.
- Escalation: manager at +7 days, director at +14, CTO dashboard at +30; overdue P0 can block related releases.
- Weekly Incident Review Board ratifies severity, challenges quality, monitors actions; target 90% closure on time within two quarters.
18. Train and certify incident roles (depends on: 7, 13, 15)
Command is a skill, not a title. Certify before duty; use paid working time.
- All employees: 30-min module on recognizing impact, declaring, finding channel/status page.
- Responder: half-day on severity, escalation, runbook, financial integrity; mandatory before joining rotation.
- Scribe: 2 hours timeline discipline; entry point.
- Incident Commander: 2 days + shadowing; command presence, delegation, severity calls, running room, handover; certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead: 1 day on status writing, customer tiering, legal boundaries, regulator triggers.
- Domain responders demonstrate dashboard, runbook, rollback, failover, access before primary; two shadow shifts; no primary in first 90 days.
- Certification valid 12 months, renewed by simulation; register is audit artifact.
19. Simulations and game days (depends on: 10, 14, 15, 18)
Rehearse process before real SEV1; exercise tooling failure and ledger scenarios.
- Monthly tabletop per group reusing an incident from baseline, rotating commander.
- Quarterly game day in staging or tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes degradation.
- Twice-yearly unannounced paging drill to measure real overnight ack times.
- Annually one combined operational/security exercise and one regulator-notification exercise with Legal/CISO.
- Exercise status page down, chat/paging provider unavailable.
- Never inject uncontrolled change into production ledger; validate backups/restore/RPO/RTO on replicas.
- Every exercise produces tracked actions; publish time-to-commander, time-to-first-update, time-to-mitigation.
20. Pilot on critical path (depends on: 9, 11, 12, 13, 14, 15, 17, 18, 19)
Prove full process on highest-risk surface with willing teams for six weeks before full rollout.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, Support intake, including two of the 12 already on-call teams.
- Activate severity scale, command corps, single paging tool, page budget, status page policy, mandatory postmortems, paid rotations.
- Parallel old paths for one week then cut over; real incidents use new process only.
- Program lead attends every SEV2+ as coach, never shadow commander; review every pilot page within one business day.
- Exit gate: commander named within 5 min in 95% cases, first status update on time, MTTD under 10 min for pilot services, pages halved, no unpaid page, all required postmortems on time, positive sentiment.
- Publish one-page result company-wide.
21. Metrics, dashboards, and review cadence (depends on: 13, 17)
Instrument the process so improvement is visible and audit sees monitoring/review evidence. Report median and p90, never averages alone.
- Response: time to detect/declare/commander/ack/mitigate/resolve split by severity, tier, journey, region, detection source.
- Quality: customer-first rate, status page timeliness, update cadence, missed escalations, role conflicts, alert actionability, out-of-hours pages per person.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat factors.
- Business: availability per journey, error budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, frequency, recovery days, sentiment, attrition.
- Weekly Incident Review Board; monthly reliability review per team; monthly executive review to CEO; quarterly control review with Security/Compliance/Risk; quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, status page history; team scorecards direct help, never individual penalties.
22. Wave rollout to all 28 teams with readiness gates (depends on: 20, 21)
Roll out in four waves by customer risk, each passing an explicit gate rather than a date. Complete all teams by week 24 to leave ~3 months of evidence before audit fieldwork.
- Wave 1 remaining Tier 0, Wave 2 Tier 1, Wave 3 Tier 2, Wave 4 Tier 3/internal.
- Per-team onboarding kit: catalog entry complete, alerts migrated within budget, runbooks at readiness bar, rotation staffed with 6 trained responders or time-limited exception, one IC candidate nominated, one tabletop passed, payroll active.
- Named coach per wave for three weeks; director signs gate.
- Legacy paging paths disabled per wave, not kept as comfort fallback.
- Publish live adoption scoreboard.
- Any team unable to staff fair rotation gets headcount, service reassignment, or explicit executive risk acceptance.
23. Culture, fairness, and continuous feedback (depends on: 4, 9)
Run this in parallel from day one. Engineers judge fairness; executives judge results.
- Repeat the deal in every forum: paid on-call, paged only for owned services, trained commander, real sprint capacity.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
- Weekly office hours first 8 weeks, Slack support channel 4-hour SLA, per-team champion, weekly newsletter with metrics and bad news.
- Incident response contribution in promotion criteria; quarterly awards for best postmortem and biggest noise reduction; public thanks after SEV1.
- Pulse survey at 60, 120, 365 days; if fairness or load red, pause expansion until fixed.
24. SOC 2 evidence design and internal dry-run (depends on: 13, 17, 21, 22)
Make evidence a by-product of operations; test before external auditor.
- Map with Compliance to Trust Services Criteria CC7.2-CC7.5, CC2.2/CC2.3, A1.2; confirm observation window early.
- Evidence set automated and indexed: versioned policies/exception, rotation schedules, compensation activation, incident records with timestamps/roles, paging/ack logs, status history, reportability decisions, postmortems, action closure with verification, training/certification register, drill records, access reviews.
- Internal Audit tests process in months 4 and 6, sampling incidents end-to-end.
- Run formal mock audit month 7 using same evidence populations; include interviews with IC, engineer, Support, Compliance.
- Correct failures through tracked actions, never by editing history.
- Freeze process wording after month 7; any change logged as exception.
25. Risk register and contingency planning (depends on: 1)
Name likely failure modes and pre-commit responses; review monthly with sponsor.
- Too few commander volunteers: make command rostered duty for managers and staff engineers until pool reaches 30.
- Compensation not approved in time: fallback to time-off-in-lieu plus phased stipend, never launch mandatory night on-call uncompensated.
- Tool migration slips: cut scope on incident record layer, never on paging consolidation.
- Noise pruning hides real failure: demote to ticket first, observe 30 days, keep recovery path, monthly missed-detection review.
- Burnout in experienced teams: weekly load monitoring with page caps.
- Major SEV1 mid-rollout: program lead becomes full-time responder, wave schedule slips one wave, sponsor told same day.
- Shared ledger concentration: if blast-radius workstream slips, escalate to board as accepted risk with dated plan.
26. Inspect and adapt; year-two sustainability (depends on: 22, 24)
Prevent decay after audit by revising on data and assigning permanent owners.
- At 90 days live, revise policy using measurements: severity calibration, Layer B/C membership from page data, uncovered shifts, commander burn, missed updates, action closure, survey results.
- Assign permanent owners for policy, paging platform, status page, service catalog, training, metrics.
- Re-baseline targets every six months; shift from lagging to leading metrics: error budget burn, near-miss rate, drill performance.
- Year-two candidates: follow-the-sun coverage, automated mitigation for top recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, completion of ledger blast-radius reduction.
- Keep annual policy review, certification renewal, exercise calendar, board reporting as permanent commitments.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
Answer with a single JSON object, and nothing else, that conforms to this JSON schema:
{"description": "The structure every proposing agent must answer with, in every round.", "properties": {"steps": {"items": {"description": "One step of a proposed plan. Dependencies are step_ids of the same proposal.", "properties": {"step_id": {"title": "Step Id", "type": "string"}, "title": {"title": "Title", "type": "string"}, "description": {"title": "Description", "type": "string"}, "dependencies": {"items": {"type": "string"}, "title": "Dependencies", "type": "array"}}, "required": ["step_id", "title", "description", "dependencies"], "title": "Step", "type": "object", "additionalProperties": false}, "title": "Steps", "type": "array"}, "estimated_complexity": {"title": "Estimated Complexity", "type": "string"}, "success_metrics": {"title": "Success Metrics", "type": "string"}}, "required": ["steps", "estimated_complexity", "success_metrics"], "title": "ProposalContent", "type": "object", "additionalProperties": false}
{"estimated_complexity": "high", "success_metrics": "- Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.\n- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.\n- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.\n- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.\n- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.\n- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.\n- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.\n- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.\n- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.\n- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.\n- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.\n- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.\n- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.\n- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.\n- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.\n- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.\n- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.\n- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.\n- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of \"not reportable\".\n- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.\n- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.\n- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.\n- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.", "steps": [{"step_id": "S1", "title": "Charter the program and start the audit clock", "description": "Turn the CEO email into a chartered company program with one accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.\n\n- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.\n- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.\n- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.\n- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.\n- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.\n- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.", "dependencies": []}, {"step_id": "S2", "title": "Seven-day operating floor so the next outage already has an owner", "description": "Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.\n\n- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.\n- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.\n- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.\n- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.\n- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.\n- Run a 15-minute daily operations review until the permanent process is live.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Forensic baseline of incidents, alerts and money lost", "description": "Rebuild the facts before designing anything. This is both the design input and the frozen \"before\" picture for the executive team and the auditor.\n\n- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.\n- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.\n- Reconstruct the two \"nobody in charge\" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.\n- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.\n- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.\n- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.\n- Confirm the SOC 2 Type II observation window with the auditor; map the future response controls to applicable Trust Services Criteria; define evidence set and retention requirements.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Listening tour, resistance map and the written on-call deal", "description": "Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.\n\n- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.\n- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.\n- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.\n- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.\n- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.", "dependencies": ["S1"]}, {"step_id": "S5", "title": "Service catalog, ownership, tiering and customer-journey map", "description": "You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.\n\n- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.\n- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).\n- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.\n- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.\n- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.\n- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.", "dependencies": ["S3"]}, {"step_id": "S6", "title": "Severity standard, declaration rights and incident lifecycle", "description": "Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.\n\n- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.\n- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.\n- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.\n- **SEV4:** no customer impact. Ticket only; never pages a human.\n- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.\n- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.\n- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.\n- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.\n- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.", "dependencies": ["S3", "S5"]}, {"step_id": "S7", "title": "Roles, decision authority and handover discipline", "description": "Solve \"nobody in charge for an hour\" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.\n\n- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.\n- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.\n- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.\n- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.\n- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.\n- Rules: command claimed within 5 minutes and stated in channel (\"I am IC\"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.\n- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.", "dependencies": ["S6"]}, {"step_id": "S8", "title": "Three-layer 24x7 coverage: command corps, domain rotations, triage desk", "description": "Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.\n\n- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.\n- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.\n- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.\n- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to \"I will not carry a pager for other teams' code.\"\n- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.\n- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.\n- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.", "dependencies": ["S5", "S7"]}, {"step_id": "S9", "title": "Paid on-call, New York labour compliance and fatigue safeguards", "description": "Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.\n\n- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.\n- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.\n- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.\n- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.\n- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.\n- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.\n- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.", "dependencies": ["S4", "S8"]}, {"step_id": "S10", "title": "Alert quality contract and page budget", "description": "3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.\n\n- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.\n- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.\n- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.\n- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.\n- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.\n- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.", "dependencies": ["S3", "S5"]}, {"step_id": "S11", "title": "Detection uplift on the money path, validated by incident replay", "description": "The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.\n\n- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.\n- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.\n- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.\n- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.\n- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish \"replay coverage\" as a leading metric and close the gaps it exposes.\n- Record detection source on every incident. \"Customer detected first\" becomes a named defect class with a mandatory tracked action.", "dependencies": ["S5", "S10"]}, {"step_id": "S12", "title": "One pager, one incident record, one status page", "description": "Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.\n\n- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.\n- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.\n- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.\n- Route every page through the service catalog: service label → owning domain schedule → escalation policy.\n- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.\n- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.\n- Set a hard date after which a page outside this platform creates no on-call obligation.", "dependencies": ["S6", "S8", "S10"]}, {"step_id": "S13", "title": "Escalation ladder and the five-minute command rule", "description": "Write one unskippable path from \"something looks wrong\" to \"someone is in charge\", and make the default action never be waiting.\n\n- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.\n- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.\n- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.\n- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.\n- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.\n- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.\n- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.", "dependencies": ["S7", "S8", "S12"]}, {"step_id": "S14", "title": "Live execution doctrine and payments safety rules", "description": "Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.\n\n- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.\n- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.\n- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.\n- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.\n- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.\n- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.\n- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.\n- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.\n- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.\n- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.", "dependencies": ["S7", "S13"]}, {"step_id": "S15", "title": "Internal communications protocol", "description": "Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.\n\n- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.\n- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.\n- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.\n- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.\n- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.\n- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.", "dependencies": ["S7", "S12"]}, {"step_id": "S16", "title": "Customer communications and status page policy", "description": "Replace \"whoever is around\" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.\n\n- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.\n- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.\n- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.\n- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.\n- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.\n- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.", "dependencies": ["S6", "S12", "S15"]}, {"step_id": "S17", "title": "Regulatory, partner and legal notification playbook", "description": "In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.\n\n- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.\n- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is \"not reportable\"**, with evidence, decision-maker and timestamp.\n- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.\n- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.\n- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.\n- Rehearse the playbook quarterly as part of the exercise programme.\n- Agree with Legal and Finance the availability measurement method per contract and per component; compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.\n- Decide credit posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain; attribute credits to root-cause families for investment decisions.", "dependencies": ["S6", "S16"]}, {"step_id": "S18", "title": "Blameless postmortem standard and Incident Review Board", "description": "Replace \"some incidents, various formats\" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.\n\n- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.\n- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.\n- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.\n- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**\n- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.\n- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.\n- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly \"top five recurring causes\" analysis, with restricted versions for security or privileged content.", "dependencies": ["S6", "S7", "S12"]}, {"step_id": "S19", "title": "Action ownership, reserved capacity and enforcement", "description": "Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.\n\n- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.\n- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.\n- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.\n- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.\n- Verify effectiveness before closing. A closed ticket without evidence does not close the action.\n- Report closure rate by team monthly and include it in engineering manager objectives.\n- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.", "dependencies": ["S12", "S18"]}, {"step_id": "S20", "title": "Training and certification academy", "description": "Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.\n\n- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.\n- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.\n- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.\n- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.\n- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.\n- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.\n- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.", "dependencies": ["S7", "S13", "S15", "S18"]}, {"step_id": "S21", "title": "Exercise programme: tabletops, game days and unannounced drills", "description": "The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.\n\n- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.\n- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.\n- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.\n- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.\n- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.\n- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.\n- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.", "dependencies": ["S12", "S14", "S20"]}, {"step_id": "S22", "title": "Publish Incident Management Policy v1 and the exception register", "description": "Collapse the design into a document people will actually open mid-outage, and make it official.\n\n- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.\n- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.\n- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.\n- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.\n- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.", "dependencies": ["S6", "S7", "S8", "S9", "S10", "S13", "S15", "S16", "S17", "S18", "S19"]}, {"step_id": "S23", "title": "Pilot on the payments critical path", "description": "Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.\n\n- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.\n- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.\n- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.\n- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.\n- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.\n- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.", "dependencies": ["S9", "S11", "S12", "S14", "S20", "S22"]}, {"step_id": "S24", "title": "Metrics, dashboards, review cadence and anti-gaming", "description": "Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.\n\n- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.\n- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.\n- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.\n- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.\n- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.\n- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.\n- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.\n- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.", "dependencies": ["S12", "S18", "S23"]}, {"step_id": "S25", "title": "Alert noise burn-down campaign", "description": "Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.\n\n- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.\n- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.\n- Publish a weekly noise leaderboard that names systems, never people.\n- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.\n- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.\n- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.", "dependencies": ["S10", "S12", "S23"]}, {"step_id": "S26", "title": "Wave rollout to all 28 teams with readiness gates", "description": "Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.\n\n- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.\n- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.\n- Each wave gets a named coach from the program office for three weeks.\n- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.\n- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.\n- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.", "dependencies": ["S23", "S24", "S25"]}, {"step_id": "S27", "title": "Change management, fairness and pager culture", "description": "Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.\n\n- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.\n- Publicly close the two historical \"nobody in charge\" incidents with a written account of what would be different now.\n- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.\n- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.\n- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.\n- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.", "dependencies": ["S4", "S9", "S23"]}, {"step_id": "S28", "title": "SOC 2 evidence by design, internal testing and mock audit", "description": "Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.\n\n- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.\n- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including \"not reportable\", postmortems, action closure with verification, training register, drill records, access reviews.\n- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.\n- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.\n- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.\n- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.", "dependencies": ["S22", "S24", "S26"]}, {"step_id": "S29", "title": "Risk register and pre-committed contingencies", "description": "Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.\n\n- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.\n- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.\n- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.\n- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.\n- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.\n- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.\n- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.", "dependencies": ["S1"]}, {"step_id": "S30", "title": "Ninety-day inspect-and-adapt, then year-two sustainability", "description": "Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.\n\n- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.\n- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.\n- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.\n- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.\n- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.\n- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.", "dependencies": ["S24", "S26", "S28"]}]}Votes
The calls of the vote
The votes
| Voter | Model | Voted for | Justification | Usage | Raw |
|---|---|---|---|---|---|
| opus5_voter_1 | opus5 · anthropic/claude-opus-5 |
Proposal A · opus5_refine_1 winner | Proposal 1 covers every element of the brief with the most operational depth, and it uniquely adds items the others omit or bury: week-one audit/legal and evidence scoping, change-intelligence and deploy-safety on the critical path (the main lever on a 3h10 MTTM), Support/account management as an instrumented detection tier, a multi-incident concurrency doctrine, and an incentive guardrail warning that credits may rise before they fall. Its staffing answer to the pager objection is the strongest: three coverage layers plus a paid overnight Duty Triage Desk, with published arithmetic showing why 28 night rotations are rejected and roughly 100 of 260 engineers carry a night duty at one week in six, gated on compensation being live in payroll. The success metrics are specific, baseline-anchored and mostly leading rather than lagging (replay coverage, page budget, action closure, mock-audit outcomes), and the 35-step dependency graph runs coherently from a seven-day operating floor through pilot, waves, control test and mock audit before the SOC 2 window. The only real cost is length; the competing proposals are near-identical in structure but consistently thinner on these differentiators. |
79.9k in · 1.3k out · 20 s | show[SYSTEM]
You are an expert and objective evaluator of project plan proposals.
Your task is to select the BEST proposal based on criteria of completeness, clarity, and alignment with the main objective.
Use your internal reasoning processes to thoroughly analyze each proposal, considering all aspects and implications.
Take as much time and space as you need to evaluate each proposal in depth before making your decision.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
Proposals to Evaluate:
--- PROPOSAL 1 ---
Proposal ID: 30f15b26-0ec0-4756-9f4a-6279eb889884
Content:
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90 a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95% internally and externally.
- Top-100 account outreach completed within 30 minutes of SEV1 in 95% of qualifying cases, using the approved briefing pack.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- Change-correlated incidents are identified within 10 minutes of declaration in 90% of cases, and the change-correlated share of incidents declines quarter over quarter.
- SLA credits fall to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, one concurrency drill and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and a tested read-only mode by month 6, with milestones reported to the board quarterly.
Steps (35):
1. Charter the programme: one owner, one mandate, funded, dated against the audit
Turn the CEO email into a chartered company programme with a single accountable owner and authority over all 28 teams. Incident response stops being a per-team preference and becomes a **company operating process**.
- Name the CTO as executive sponsor and a full-time Head of Reliability & Incident Management as accountable owner, supported by a programme office of three: programme lead, incident-platform engineer, reliability analyst.
- Form a decision group (Engineering, Platform/SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables on day one: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, named service ownership, and **paid on-call**.
- Publish the timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal control test month 5, mock audit month 7, SOC 2 fieldwork month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.3M/yr, training and exercises, plus 15% of engineering capacity reserved for reliability work.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence immediately.
- Publish a one-page interim severity card and a single declaration path: one chat command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim 24x7 Duty Incident Commander roster (primary plus backup) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document and one naming convention for every major incident, starting now.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is the design input, the executive narrative and the frozen "before" picture for the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the **detection backlog**.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Correlate incidents with deployments, config changes and feature flags to quantify how many started with a change we made.
- Quantify true cost beyond credits: failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost, churn risk on named accounts.
- Freeze the baselines in a signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Audit, legal and evidence scoping in week one (depends on: 1)
Design evidence as a by-product of operating, never as a reconstruction before fieldwork. Settle the scope and the legal handling of incident records now, not in month seven.
- Confirm with the external auditor the Type II observation window, the incident population definition, and the evidence they will sample. Everything from the day-7 floor onwards must count.
- Map incident response to the Trust Services Criteria with Compliance: CC7.2–CC7.5 (monitoring, identification, response, recovery), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), A1.2 (availability).
- Define the evidence set and where it is produced automatically: incident records, paging and acknowledgement logs, role assignments, status-page history, reportability decisions, postmortems, action closure with verification, training register, drill records, access reviews.
- Agree retention, confidentiality, legal hold and access rules. Decide with counsel which postmortem content is privileged and how privileged material is segregated **without making the ordinary postmortem secret**.
- Start the obligation matrix with Legal: NYDFS 23 NYCRR 500, state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network contractual windows, cyber-insurer notice.
- Begin monthly evidence sampling from month 1 so control operation is visible long before the audit.
5. Listening tour, resistance map and the written on-call deal (depends on: 1, 3)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written deal signed by the sponsor.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
- State the hard gate publicly: no new mandatory night rotation starts before compensation, training, runbooks and staffing rules are live.
6. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- Record for every service: one accountable team, engineering manager, chat channel, escalation policy, dashboards, runbook link, dependency list, regions and data stores.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record SLO, RTO, RPO, active-active or region-bound, failover method, and the dependencies that make nominal two-region redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
7. Severity standard, declaration rights and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning responders and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and responders paged, comms and scribe if customer-visible, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, credit-generating or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Attach a financial-integrity flag or a security flag to any severity. The flag forces dual control, Legal engagement and the reportability checkpoint without inventing a fifth level.
- Anchor on payments reality alongside error rates: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including a customer report.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
8. Roles, authority, concurrency and handover discipline (depends on: 7)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice to the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1 and SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, Finance and Vendor Management join on defined triggers rather than by invitation.
- Rules: command claimed within 5 minutes and stated in channel; distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- **Concurrency doctrine:** two simultaneous SEV1/SEV2 incidents activate the secondary commander and a surge roster; a designated Multi-Incident Coordinator arbitrates shared resources such as the ledger, the database platform and the deploy freeze.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
9. Three-layer 24x7 coverage: command corps, domain rotations, overnight triage desk (depends on: 5, 6, 8)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving roughly one primary week per person every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement and reconciliation, auth, API edge, Kubernetes platform, data and reporting, partner integrations). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Publish the staffing arithmetic: 10–12 domain rotations of six to eight people plus a 30-person command corps means roughly 100 of 260 engineers carry a night obligation, at about one week in six. Twenty-eight independent rotations would be unstaffable and is therefore rejected on the numbers.
- Teams too small for a fair rotation get headcount, service reassignment, or a time-limited executive exception. **Never a two-person 24x7 rotation.**
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: a human acknowledgement within 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
10. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 5, 9)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour, and a premium for the overnight triage desk.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per responder per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty since day 7 is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons without career penalty, and on-call load is published quarterly by team.
11. Alert quality contract and page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression can never hide an incident.
12. Detection uplift on the money path, validated by incident replay (depends on: 6, 11)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action and a missed-detection review.
13. Change intelligence and deployment safety on the critical path (depends on: 6, 12)
Most of these incidents start with something we changed. Make change the first hypothesis the tooling answers, and make changes safer to reverse.
- Stream every deployment, configuration change, feature-flag flip, schema migration and infrastructure change into the incident timeline with service labels and owners.
- Give the commander an automatic "what changed in the last 60 minutes on the affected journey" panel at declaration time.
- Require Tier 0/1 changes to be progressively delivered with a documented rollback that is tested, time-bounded and executable by the on-call responder without the author.
- Treat ledger schema migrations and settlement-affecting changes as a separate class: dual approval, rehearsed rollback, no deploys inside the settlement window.
- Enforce a change freeze during SEV1 and SEV2, lifted only by the commander and logged.
- Report change-correlated incidents monthly; a rising ratio is a signal to strengthen release safety, not to blame a team.
14. One pager, one incident record, one status page — migrated without a detection gap (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy paging path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access, MFA and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page delivered outside this platform creates no on-call obligation.
15. Escalation ladder and the five-minute command rule (depends on: 8, 9, 14)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts and support entitlements: AWS and database premium support, processors, sponsor banks, card networks.
16. Live execution doctrine and payments safety rules (depends on: 8, 14, 15)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed in advance rather than improvised.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and automatic reopening if impact recurs.
17. Service readiness bar, major-incident playbooks and ledger blast-radius reduction (depends on: 6, 9, 16)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human, and the shared ledger is the largest structural risk in the estate.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as a concentration risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers, stronger regional independence — with board-visible quarterly milestones.
- Enforcement: no readiness sign-off, no night paging alerts unless the manager accepts the gap in writing with an expiry date and a compensating control. Never respond to a gap by turning detection off.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
18. Internal communications protocol (depends on: 8, 14)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, Sales and Security.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change. First internal brief within 10 minutes of SEV1 and 15 of SEV2.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
19. Customer communications, status page and account-manager outreach (depends on: 7, 18)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; monitoring notice after mitigation; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, monitoring, resolution.
- Component-level status page mapped to customer journeys rather than internal service names, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1 and 60 of SEV2, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state observed impact, affected capabilities, any workaround and the next update time; **never speculate on cause, recovery time, data integrity or blame**.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
- Record any legally required restriction or delay of public detail, its approver, and the alternative stakeholder plan.
20. Support and account management as a detection and intake tier (depends on: 7, 12, 19)
Customers detected 40% of incidents first, which means the front line already holds the signal. Turn Support and account managers into an instrumented detection channel rather than a bystander.
- Give Support explicit declaration rights, a one-page trigger card, and a macro that opens an incident candidate directly in the incident platform.
- Automate the clustering rule: five similar tickets or calls in ten minutes auto-creates a triage incident assigned to the duty commander.
- Route credible partner, processor and sponsor-bank notifications into the same declaration path within five minutes.
- Generate the affected-customer list automatically from journey telemetry and the incident record, and push it to Support, CS and the account-manager briefing.
- Train Support and account managers on approved language and prohibit independent technical explanations to customers.
- Measure and publish "signal was in Support before it was in monitoring" as a detection defect, and feed each instance into the detection backlog.
21. Regulatory, partner and legal notification playbook (depends on: 4, 7, 19)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Complete the obligation matrix started in scoping and have counsel validate the triggers, deadlines, channels and submitting authority for each obligation.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with facts considered, decision-maker, timestamp and reassessment trigger.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, card networks, outside counsel, cyber-insurer, critical vendors.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the operational facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly inside the exercise programme, including one full regulator-notification simulation per year.
22. SLA credit workflow, true cost model and incentive guardrails (depends on: 7, 19)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component, including how partial degradation counts.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Guardrail against the perverse incentive: better detection will surface incidents that previously went unbilled, so credits may rise before they fall. Publish this expectation to the executive team in advance, and make it a written rule that no finance or commercial pressure may influence severity or declaration.
23. Blameless postmortem standard and Incident Review Board (depends on: 4, 7, 8)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger; any control failure.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning engineering director is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions where security, privacy or privileged content requires it.
24. Action ownership, reserved capacity and enforcement (depends on: 14, 23)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, accountable manager, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness through tests, telemetry, exercises or production evidence before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
25. Training and certification academy (depends on: 8, 15, 18, 23)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, the 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, contractual clocks, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days of joining a domain.
- Certification is valid 12 months and renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
26. Exercise programme: tabletops, game days, night drills and vendor rehearsals (depends on: 14, 17, 25)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander. First company-wide tabletop within 30 days.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times, plus a concurrency drill with two simultaneous incidents.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down. Rehearse joint escalation with AWS, a processor and a sponsor bank.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
27. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 15, 18, 19, 21, 23, 24)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, escalation and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix, Readiness Bar.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
28. Pilot on the payments critical path (depends on: 10, 12, 13, 14, 17, 25, 27)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement and reconciliation, PostgreSQL platform, Kubernetes platform, API edge, auth and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, action tracking, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The programme lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time in 95% of cases, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
29. Alert noise burn-down campaign (depends on: 11, 14, 28)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → under 1,500 pages a month by day 90, under 500 with noise below 15% by month 6.
30. Metrics, dashboards, review cadence and anti-gaming (depends on: 14, 23, 28)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time from first impact to detect, declare, commander, acknowledge, mitigate, resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors, change-correlated incident ratio.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
31. Wave rollout to all 28 teams with readiness gates (depends on: 28, 30)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the programme office for three weeks; the director signs the gate.
- **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard by team, tier and control gap.
32. Change management, fairness and pager culture (depends on: 5, 10, 28)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
33. SOC 2 evidence by design, internal control test and mock audit (depends on: 4, 27, 30, 31)
The real process is the evidence. Never build a parallel audit process, and never reconstruct records after the fact.
- Maintain the evidence set defined in scoping, produced automatically and indexed: versioned policies and exceptions, catalog records, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Sample incidents monthly from first signal through verified corrective action, deliberately including customer-reported events, downgraded incidents and missed timelines.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling end to end and reporting gaps to the sponsor.
- Run a formal mock audit in month 7 using the populations, evidence requests and interviews the auditor will use: a commander, a random engineer, Support, Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
34. Programme risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps and immediate staffing intervention.
- **A major SEV1 mid-rollout:** the programme lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
35. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 30, 31, 33)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, exercise calendar and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and quarterly board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 2 ---
Proposal ID: 1f1527bb-d992-451c-bbc4-9e61ffdb1f30
Content:
Estimated Complexity: high
Success Metrics: - By day 7, every suspected major incident uses one record, one coordination channel, and a named Incident Commander within 10 minutes.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes.
- By month 3, at least 95% of SEV1 and SEV2 incidents have a named commander within 5 minutes.
- By month 3, at least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and appropriate technical coverage.
- By week 18, all 180 services have an owner, tier, escalation path, and tested coverage model.
- No new mandatory night rotation begins before compensation, training, access, runbooks, and minimum staffing are active.
- Every direct 24x7 technical rotation has at least six qualified responders or an approved, expiring executive exception.
- No responder is routinely primary more often than one week in six or assigned to two simultaneous primary rotations.
- Median impact-to-detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- By day 90, incident replay identifies a current detector for at least 90% of the 31 historical incidents, including its expected detection minute.
- Median time to mitigation falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of applicable SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet the required internal and customer update cadence by month 3.
- Monthly human notification episodes fall from the current 3,400 alert events to no more than 1,500 by day 90 and 500 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without reducing Tier 0 or Tier 1 replay coverage.
- Average after-hours load remains at or below two notification episodes per responder per week; every sustained breach receives a dated remediation plan.
- All six monitoring sources route human pages through the controlled paging platform by week 12, with direct legacy routes retired by week 18.
- All required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 from month 3 onward.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due dates with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Reportability is assessed and recorded within 1 hour for 100% of SEV1 and qualifying SEV2 incidents, including not-reportable decisions.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Approved ledger RTO and RPO, controlled failover, read-only mode, and post-recovery reconciliation are exercised before month 6.
- Quarterly responder surveys reach at least 75% favorable responses for fairness, ownership boundaries, compensation, and sustainability by month 6.
- Monthly contracted availability meets or exceeds 99.95% by month 6 using the contractually authoritative measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- The month-7 mock audit finds no unowned high-risk control gap and at least 95% of sampled incidents contain complete operating evidence.
- The SOC 2 Type II incident-response controls complete external testing without an unresolved material exception.
Steps (28):
1. Charter the program and fund immediate action
Make incident management a **company operating process** within 48 hours. Give one accountable leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Include Engineering, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit in a decision group. The sponsor resolves blocked decisions within 48 hours.
- Approve non-negotiables: one severity model, one human-paging path, one incident record, paid on-call, named service ownership, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, resilience, and incident actions.
- Fund tooling, compensation, training, exercises, and program staff. Compare the cost with the existing $1.3M in annual SLA credits.
- Set the schedule: operating floor by day 7, design during weeks 1–4, pilot during weeks 5–10, rollout during weeks 11–18, control tests in months 4 and 6, and mock audit in month 7.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for the final policy or new tooling. Put a minimum viable process into operation immediately and retain evidence from the first day.
- Publish a one-page interim severity guide, declaration procedure, role card, and communications clock.
- Provide one monitored declaration route through chat, telephone, and the current paging environment.
- Create one channel, bridge, timeline, and incident identifier for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced managers and engineers.
- Require a named commander within 10 minutes. The duty engineering director assumes command if the command page is unclaimed.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Stop uncompensated mandatory after-hours expansion. Pay interim duty under a temporary stipend, retroactive to program launch.
- Hold a 15-minute daily operational review until the permanent process is active.
3. Create the factual, contractual, and control baseline (depends on: 1)
Build one defensible baseline for process design, investment decisions, and SOC 2 testing. Preserve the original data so progress cannot be created by changing definitions.
- Reconstruct all 31 incidents from impact start through detection, declaration, command, mitigation, recovery, communications, credits, and corrective actions.
- Reconstruct the two incidents with unclear command minute by minute.
- Identify the missing signal for every customer-first detection.
- Inventory all six alert sources, 3,400 monthly alert events, duplicates, noisy rules, missing owners, and missing runbooks.
- Record current rotations, unpaid duty, overnight activations, schedule size, and uncovered services.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, 85% noise, $1.3M in credits, and 11 of 64 actions closed.
- Inventory customer-specific availability definitions, notice periods, credit terms, sponsor-bank obligations, and other incident-related contracts.
- Confirm the required SOC 2 Type II observation period and evidence expectations with the auditor during week 1.
4. Publish the on-call fairness contract (depends on: 1)
Treat pager resistance as a legitimate design constraint. Adoption depends on a written agreement that separates command from technical ownership.
- Interview representatives from all 28 teams and from Support, Customer Success, Security, and Operations.
- Separate concerns about unpaid work, sleep loss, unfamiliar systems, noisy alerts, inadequate runbooks, and blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for systems they own or have formally accepted and trained to support, and assisted by a separate Incident Commander.
- State that Platform may temporarily triage unknown ownership but does not inherit another team's service.
- Provide confidential accommodations for health, disability, pregnancy, or caregiving constraints without career penalty.
- Measure baseline trust, fairness, fatigue, and alert confidence. Repeat at days 60 and 120, then quarterly.
5. Build the service catalog and customer-journey map (depends on: 3)
Make a machine-readable catalog the source of truth for routing, impact analysis, status-page components, and control evidence. Every production service must have one accountable owner.
- Record the owning team, manager, business capability, repository, channel, dashboard, runbook, escalation policy, dependencies, regions, and data stores for all 180 services.
- Map initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, reporting, and onboarding to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing systems that can create material contractual impact.
- Classify Tier 2 as deferrable internal or batch systems and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional recovery mode, contractual commitments, and critical vendors for Tier 0 and Tier 1.
- Keep separate but coordinated ownership for the ledger application and PostgreSQL platform.
- Give orphan services an owner or approved decommission date within 30 days. Treat unowned Tier 0 services as release-blocking executive risks.
6. Adopt the severity model and incident modifiers (depends on: 3, 5)
Classify incidents by credible customer, financial, security, regulatory, and contractual harm. Start at the higher plausible severity while scope or integrity remains unknown.
- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; broad loss of a core payment journey; both regions impaired; or processing must be stopped. Immediately page all response roles and the Executive Duty Officer. Open the bridge within 5 minutes, freeze unrelated changes, issue internal notice within 10 minutes, publish applicable customer status within 15 minutes, start legal assessment within 1 hour, and require a postmortem.
- **SEV2 — major:** material payment degradation; settlement deadline at risk; regional impairment with reduced resilience; a critical customer or material cohort unavailable; or an SLA breach is likely. Page command and technical roles immediately. Issue internal notice within 15 minutes, applicable customer status within 30 minutes, and require a postmortem.
- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads. Page only when immediate action can reduce harm.
- **SEV4 — operational event:** no current customer impact and no credible imminent harm. Create a ticket and handle during normal hours.
- Add FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, and VENDOR modifiers. These invoke specialist controls without distorting customer-impact severity.
- Use at least SEV2 posture for unknown impact lasting 15 minutes, credible ledger-integrity risk, or a cross-domain incident without clear ownership.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Permit anyone to declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Publish a decision tree and examples based on the 31 historical incidents.
7. Define roles, authority, and handoffs (depends on: 6)
Separate coordination, communications, recordkeeping, and technical repair. One named person holds command continuously throughout every SEV1 and SEV2.
- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not act as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, executive updates, and coordination with Legal.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- **Subject-Matter Responders:** diagnose and mitigate only systems for which they have ownership, access, training, or a formally accepted support agreement.
- **Executive Duty Officer:** removes organizational obstacles and makes exceptional business decisions without displacing the commander.
- Add Security, Legal, Compliance, Finance, and Vendor Management on modifier-specific triggers.
- Require distinct commander, Communications Lead, scribe, and technical lead for SEV1. Communications and scribe may combine for the first 10 minutes of a bounded SEV2.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, reconciliation, and change evidence for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, including the exact transfer time and unresolved risks.
8. Create sustainable 24x7 coverage (depends on: 4, 5, 7)
Use central command coverage and risk-based technical coverage instead of creating 28 fragile night rotations. Responders do not carry pagers for unfamiliar code.
- Create a 24x7 Incident Command corps of approximately 30 certified people with primary and backup schedules. Two schedules create about 104 weekly assignments a year, or roughly three to four weeks per person annually.
- Create a 24x7 Communications pool of 18–24 trained people from Support, Customer Operations, Engineering Operations, and management.
- Create a similarly sized scribe pool. The backup commander temporarily records the first minutes if a scribe has not joined.
- Maintain a 24x7 Executive Duty Officer schedule and specialist contact paths for Security and Legal.
- Group Tier 0 and Tier 1 systems into roughly 8–12 coherent response domains only where responders share training, access, runbooks, and explicit support acceptance.
- Staff every critical domain with primary and secondary responders and at least six qualified people. Target eight where overnight activation is frequent.
- Give Tier 2 and Tier 3 services business-hours ownership plus tested manager and director escalation.
- Reclassify any lower-tier service capable of causing severe overnight harm rather than hiding the risk behind manager callback.
- Route unknown-owner incidents to Duty Command and Platform temporarily. Record each occurrence as a catalog control defect.
- Prohibit simultaneous primary assignments and two-person 24x7 rotations.
9. Implement compensation and fatigue protections (depends on: 4, 8)
End unpaid on-call before expanding mandatory coverage. Use fixed duty compensation so responders are not rewarded for alert volume.
- Use planning bands of $900–$1,200 per Tier 0/1 primary week and $300–$500 per secondary week.
- Use planning bands of $1,000–$1,300 per Duty Commander week and $400–$700 for Communications or scribe primary duty.
- Pay holiday premiums. Compensate all legally compensable active and waiting time for non-exempt staff, including overtime where required.
- Have HR, Finance, Payroll, and employment counsel approve final bands, tax handling, FLSA classification, New York wage-hour treatment, and schedule constraints within 14 days.
- Provide a protected recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal delivery commitments by about 15% during a primary week.
- Prohibit consecutive primary weeks, on-call during leave, hidden schedule swaps, and primary duty more often than one week in six.
- Allow responders to declare temporary fatigue-related unfitness without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Make active payroll setup, training, access, and readiness hard gates before any new mandatory night rotation starts.
10. Codify the live incident lifecycle (depends on: 6, 7)
Give every incident the same operational flow from first signal through verified recovery. The first objective is limiting customer and financial harm, not proving root cause.
- Use the states Detected, Declared, Triaged, Mitigating, Mitigated, Monitoring, Resolved, and Reviewed.
- Record impact start separately from detection. Use the earliest defensible evidence and revise it transparently when later facts emerge.
- Open with a standard command message: severity, known impact, assigned roles, immediate objective, workstreams, and next update time.
- Freeze unrelated production changes during SEV1 and normally during SEV2. Record every exception.
- Prefer reversible mitigation: rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, partner rerouting, or controlled processing suspension.
- Separate mitigation and diagnosis workstreams when staffing allows.
- Keep decisions in the shared incident record rather than direct messages.
- Require formal command and technical handoffs for shift changes, fatigue, or incidents exceeding four hours.
- For payment incidents, verify backlog handling, duplicate protection, customer state, settlement exposure, and ledger reconciliation before resolution.
- Require a severity-specific stability period and explicit handback to the owning team, Support, and Customer Success.
11. Establish one paging and incident system of record (depends on: 5, 7, 8)
Monitoring tools may remain specialized, but every human page and major-incident record must enter one controlled platform. This provides consistent routing and an audit trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page within two weeks.
- Ingest events from all six monitoring tools before disabling their direct human-paging routes.
- Route pages using the service catalog and deduplicate events belonging to the same symptom.
- Provide one declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, decisions, severity changes, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status publishing, and conference facilities.
- Apply MFA, role-based access, periodic access reviews, and tamper-evident history.
- Keep privileged legal or security material in restricted linked records rather than exposing it in the general timeline.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, the status provider, or an AWS region.
- Retire each legacy paging route only after ownership review, end-to-end testing, and two weeks of verified operation.
12. Enforce the alert-quality contract (depends on: 3, 5)
Treat every human page as a production interface with an owner and a required action. Measure human notification episodes rather than raw monitoring events.
- Require every paging rule to identify the service, owner, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Define an actionable page as one that causes or materially informs a timely intervention or risk decision.
- Define noise as duplicate, non-urgent, unactionable, stale, test-generated, or incorrectly routed notification.
- Page on payment outcomes, error-budget burn, queue age against deadlines, and financial-integrity risk rather than raw CPU, memory, pod, or log thresholds.
- Run new rules in shadow mode for seven days unless a documented emergency exception applies.
- Review repeated no-action pages within two business days.
- Set a page budget of no more than two after-hours notification episodes per responder per week, measured over four weeks.
- Make a sustained budget breach trigger a tuning sprint and block additional non-emergency paging rules.
- Require compensating detection and central approval before suppressing a Tier 0 or Tier 1 rule.
- Never disable an existing critical detector solely because its metadata or runbook is incomplete. Track the gap with a dated remediation owner.
13. Detect payment failures before customers (depends on: 5, 12)
Move detection from infrastructure health to customer journeys and ledger truth. Validate coverage against actual historical failures.
- Define SLIs and internal SLOs for initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, APIs, webhooks, and reporting freshness.
- Set internal objectives with enough headroom to protect the contractual 99.95% availability commitment.
- Run external synthetic transactions through critical journeys at least once per minute from paths independent of the production platform.
- Test each region and expose dependencies that defeat nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed payment value relative to settlement deadlines.
- Add tenant and cohort anomaly detection for high-value customers and critical payment methods.
- Convert credible Support, account-manager, processor, bank, and network reports into incident candidates within five minutes.
- Replay all 31 historical incidents. Record which current detector would fire and at what minute.
- Treat customer-first detection as a mandatory missed-detection review with a tracked action.
14. Implement the detection and escalation ladder (depends on: 7, 8, 11)
Create one time-bound path from the first credible signal to named command and the correct technical owner. Device delivery does not count as acknowledgement.
- Converge automated alerts, engineer observations, support cases, account-manager reports, partner notices, and customer calls on the same declaration path.
- Page the Duty Incident Commander and owning critical-domain primary immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at 5 minutes, manager at 10 minutes, and director at 15 minutes.
- Escalate unclaimed command to the backup commander at 5 minutes. The Executive Duty Officer assumes temporary command at 10 minutes until a certified transfer occurs.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Keep command with the current commander when service ownership remains unclear. Assign a temporary technical lead and record the ownership gap.
- Maintain tested external escalation routes for AWS, database support, processors, sponsor banks, networks, and critical vendors.
- Test the full declaration, acknowledgement, fallback, conference, and status-publishing path weekly.
- Treat missed acknowledgement, failed routing, and unowned incidents as control failures requiring review.
15. Standardize internal communications (depends on: 7, 10, 11)
Give responders one working room and stakeholders one controlled source of truth. Executives must not interrupt the technical command path.
- Maintain one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue the initial internal notice within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within the same initial-notice window.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record every material decision and outbound message in the incident timeline.
- Require the executive team to sign the communication behavior rules.
16. Standardize customer, account, and regulatory communications (depends on: 3, 6, 15)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate observed impact before root cause is known.
- Publish customer status within 15 minutes of a customer-visible SEV1 and within 30 minutes of a customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map public status components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, available workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Have account managers contact affected strategic accounts within 30 minutes for SEV1 and 60 minutes for SEV2 using the approved briefing.
- Offer status subscriptions to all customers. Auto-enroll only where contracts, consent, and applicable communication rules permit.
- Encode customer-specific notice deadlines and channels in the customer record.
- Have Legal and Compliance maintain a counsel-validated matrix covering applicable NYDFS, breach, GLBA or FTC, PCI, money-transmitter, sponsor-bank, network, insurance, and customer obligations.
- Complete and record a reportability assessment within one hour for every SEV1 and every security, privacy, or integrity-related SEV2, including decisions of not reportable.
- Let Legal own regulatory text and submission while the Incident Commander owns operational facts. Record any legally required restriction of public detail and its alternative stakeholder plan.
17. Tie incidents to SLA credits and financial exposure (depends on: 3, 16)
Link the incident record to contractual and financial outcomes. Finance should not discover outages through later credit claims.
- Define the authoritative availability calculation for each contract and customer journey with Legal and Finance.
- Calculate affected customers and minutes from incident scope and journey telemetry.
- Produce a preliminary credit and contractual-exposure estimate within five business days of resolution.
- Record failed payment count, failed or delayed value, settlement exposure, reconciliation breaks, support effort, and engineering effort.
- Establish a documented approval path for proactive credits and claims-based credits.
- Attribute credits and financial harm to recurring failure families.
- Use the quarterly credit analysis to prioritize detection, resilience, and architectural investment.
18. Establish mandatory blameless postmortems (depends on: 6, 7)
Use one learning standard with fixed deadlines. Keep learning separate from disciplinary and misconduct processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact lasting more than two hours, SLA credits, contractual breach, repeated contributing factors, major control failures, and ledger-integrity near misses.
- Produce a factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the owning engineering director accountable for completion. The commander owns response analysis, and the scribe supplies the timeline.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit answers to why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the commander or primary technical responder.
- Describe decisions using the context and information available at the time. Do not name an individual as the root cause.
- Keep HR, misconduct, and personnel matters in separate processes.
- Publish broadly useful findings internally while restricting security, privacy, personnel, and privileged content appropriately.
19. Make corrective actions enforceable risk commitments (depends on: 18)
An action is not complete when its ticket is closed. It is complete when the intended risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within 7 days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Prefer actions that remove hazards or reduce blast radius over vague actions such as retraining or adding monitoring.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of engineering capacity for approved reliability work.
- Escalate overdue high-risk actions to the manager after 7 days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit unrepaired severe conditions to block related releases.
- Verify effectiveness using tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days as complete, re-planned, superseded with evidence, or formally risk-accepted.
20. Set the service-readiness bar and payments playbooks (depends on: 5, 10, 12, 13)
A critical service must be supportable at 3 a.m. before it enters direct overnight coverage. Existing critical detection remains active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and integrity constraints.
- Require responders to demonstrate access and safe execution before independent primary duty.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional impairment, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document split-brain prevention, replay protection, failover, controlled backlog recovery, and post-recovery reconciliation.
- Require executive-approved RTO and RPO for the shared ledger cluster.
- Exercise critical runbooks at least twice a year and after material changes.
- Block new Tier 0 or Tier 1 releases and paging rules when readiness requirements are missing.
- Handle existing gaps using named owners, compensating controls, executive-approved expiry dates, and remediation plans.
- Run a parallel architecture workstream to reduce ledger blast radius, isolate non-critical readers, and strengthen regional independence.
21. Train and certify every response role (depends on: 7, 14, 15, 18, 20)
Command and communication are learned skills. Use paid working time and certify people before independent duty.
- Give all employees a 30-minute module on recognizing impact, declaring incidents, and locating the status page.
- Give responders a half-day module on severity, acknowledgement, escalation, evidence, runbooks, and financial-integrity precautions.
- Train scribes for two hours on timeline quality and fact-versus-hypothesis labeling.
- Train Incident Commanders for two days on delegation, uncertainty, severity, mitigation strategy, fatigue, handoffs, and executive management.
- Require commander candidates to complete a simulated SEV1 and two shadowed incidents or exercises.
- Train Communications Leads for one day on status writing, customer segmentation, legal boundaries, and contractual clocks.
- Require domain responders to demonstrate dashboards, access, rollback, failover, escalation, and relevant playbooks.
- Require two shadow shifts before independent primary duty.
- Renew certification annually through simulation.
- Maintain the training, assessment, and certification register as operational and audit evidence.
- Nominate an incident-management champion in each of the 28 teams.
22. Exercise command, recovery, and tool failure (depends on: 11, 20, 21)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first cross-company command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous operational and security events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors when relevant.
- Conduct at least one unannounced after-hours paging test before the audit and two annually thereafter.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure acknowledgement, command, customer notice, mitigation decision, handoff, and recovery times.
- Create tracked corrective actions for every material exercise finding.
23. Publish the signed policy and control set (depends on: 6, 7, 8, 9, 10, 12, 14, 15, 16, 18, 19)
Convert the design into concise documents that people can use during an incident. The actual operating process must also be the documented and audited process.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, evidence, and exception standards.
- Include one-page cards for severity, roles, authority, escalation, and communication timings inside the incident tool.
- State explicitly that responders support only owned or formally accepted and trained service portfolios.
- Include the compensation structure, fatigue rules, declaration rights, and non-retaliation commitment.
- Obtain approval from the CTO, HR, Legal, Security, Compliance, and Internal Audit.
- Announce the policy at an all-hands and through team briefings.
- Create an exception register with owner, rationale, compensating control, approver, review date, and expiry.
- Version every policy change. Do not rewrite historical records when the process changes.
24. Instrument the scorecard and review forums (depends on: 3, 11, 18, 19)
Measure the process before the pilot so failures become visible immediately. Report medians and 90th percentiles rather than averages alone.
- Measure impact-to-detection, detection-to-declaration, declaration-to-command, acknowledgement, mitigation, recovery, and resolution.
- Split results by severity, service tier, customer journey, region, detection source, and business-hours status.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track notification episodes, actionability, duplicates, after-hours load, missed detection, routing errors, and page-budget breaches.
- Track postmortem timeliness, action age, due-date performance, verified effectiveness, and repeated contributing factors.
- Track journey availability, error-budget burn, failed or delayed value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight work, recovery days, exceptions, sentiment, and responder attrition.
- Hold a weekly Incident Review Board chaired by the Head of Reliability with relevant directors.
- Hold a monthly executive reliability review and a quarterly control and resilience review with Internal Audit.
- Reconcile incident records monthly against support cases, customer complaints, status history, credits, and major operational anomalies to detect under-reporting.
- Use team-level scorecards to direct help and investment. Never penalize an individual for good-faith declaration.
25. Pilot the complete process on the payment path (depends on: 9, 11, 13, 20, 21, 23, 24)
Run a four-to-six-week pilot across the highest-risk journey before expanding. The interim operating floor remains active for the rest of the company.
- Include payment orchestration, ledger application, PostgreSQL platform, API edge, authentication, settlement, reconciliation, Kubernetes platform, and Support intake.
- Include teams with existing on-call experience and teams new to the model.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking together.
- Run legacy and new paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently taking command.
- Correct critical process or tooling defects within 48 hours.
- Exit only after 95% timely command assignment, 95% communications compliance, no unpaid pages, complete required postmortems, tested fallbacks, and at least 50% lower pilot noise.
- Publish the pilot results, defects, and policy changes company-wide.
26. Roll out by risk with readiness gates (depends on: 25)
Expand in controlled waves and finish early enough to accumulate operating evidence before the audit. A calendar date does not override a failed readiness gate.
- Roll out remaining Tier 0 domains first, followed by Tier 1, Tier 2, and Tier 3.
- Use four waves of six to eight teams, each lasting two to three weeks.
- Gate each service on catalog ownership, appropriate coverage, active compensation, trained responders, access, tested escalation, alert quality, runbooks, and a passed tabletop.
- Require at least six responders only for direct 24x7 technical rotations. Use business-hours coverage for lower tiers.
- Give each wave a named coach and director sign-off.
- Reschedule failed gates or use a time-limited executive exception with compensating controls. Do not create silent waivers.
- Disable legacy human-paging routes after verified cutover for each wave.
- Run a quota-based noise reduction sprint in every wave, starting with the highest-volume rules.
- Pair every suppression with a compensating-detection check.
- Publish an internal adoption dashboard by team, service tier, coverage, and control gap.
- Complete critical coverage by approximately week 12 and all 28 teams by week 18.
27. Prove SOC 2 operating effectiveness (depends on: 22, 23, 24, 26)
Generate evidence through normal operation rather than reconstructing it before fieldwork. Test both control design and consistent execution.
- Map controls to the applicable Trust Services Criteria with Compliance and the auditor, including monitoring, incident identification, response, recovery, communications, and availability.
- Retain approved policies, exceptions, service ownership, schedules, compensation activation, access reviews, training, incidents, communications, reportability decisions, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, missed timelines, non-reportable decisions, and exercises in the testing population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the evidence populations and interviews expected from the external auditor.
- Correct deviations through tracked actions with owners and dates. Never edit history to create apparent compliance.
- Verify that evidence retention covers the full auditor-defined observation period.
- Brief commanders, engineers, Support, and Compliance on the actual process without scripting inaccurate answers.
28. Inspect, adapt, and institutionalize ownership (depends on: 26, 27)
Prevent the process from decaying after rollout or the audit. Change it using measured operating evidence rather than opinion.
- Review the policy after 90 days of live operation using severity calibration, page load, missed detection, communication compliance, action closure, fatigue, and survey results.
- Remove steps that create work without reducing risk. Add controls only where incidents, exercises, or evidence show a gap.
- Reassess Tier 0 and Tier 1 classification and domain boundaries every six months.
- Assign permanent owners for policy, catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review compensation bands, rotation burden, accommodations, and staffing annually.
- Report severe incidents, credits, overdue high-risk actions, and ledger concentration risk to the board or risk committee quarterly.
- Maintain the ledger blast-radius program as an executive risk until failover, degraded mode, reconciliation, and regional independence meet approved objectives.
- Evaluate follow-the-sun coverage using one year of actual activation and staffing data.
- Build year-two plans for automated mitigation, safer deployments, graceful degradation, and error-budget release controls.
--- PROPOSAL 3 ---
Proposal ID: f61aca40-a8a9-4e91-a5f3-ac1d03634441
Content:
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of 'not reportable'.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (29):
1. Charter the program, fund it, and start the audit clock
Convert the CEO email into a company operating process with one accountable owner, a budget, and a dated timeline that starts this week.
- Name the CTO as executive sponsor and appoint a Head of Reliability & Incident Management as the single accountable owner with full-time authority over all 28 teams.
- Stand up a three-person program office: program lead, platform engineer, reliability analyst.
- Form an eight-person steering group spanning Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, and HR. It proposes; the sponsor decides within 48 hours. Never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one human-paging path, one incident record, one postmortem format, paid on-call, mandatory action tracking, and named service ownership.
- Approve budget anchored against the $1.3M in SLA credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs, and 15% reserved engineering capacity.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Publish a one-page charter company-wide on day 2. Incident response is a company process, not a per-team preference.
- Timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day interim command floor (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place this week so the next outage has a named commander and the evidence clock starts immediately.
- Publish a one-page interim severity card and one declaration path: a Slack command, a phone number, and the existing pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs from the 12 teams already on-call. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical action items into complete, re-plan, or formally risk-accept within 30 days, prioritising ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline of incidents, alerts, and money lost (depends on: 1)
Rebuild the facts before locking any design. This is both the design input and the frozen 'before' picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines in one signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the written fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written, signed deal.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- State explicitly that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Map compliance, evidence, and audit requirements from day one (depends on: 1)
Design evidence as a by-product of operations, not a reconstruction before the auditor arrives. The interim process in week 1 is already evidence.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for internal and external communication; CC5 for control activities; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure with verification, training register, drill records, reportability decisions including 'not reportable'.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records. Require role-based access, MFA, and periodic access review.
- Version and approve all policy documents from day one: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard, Alert Quality Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has a named owner. A wrong owner recreates the pager objection. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Assign one accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. **Tier 1**: customer-facing but degradable. **Tier 2**: internal or batch. **Tier 3**: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for every Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform. They fail differently and need different hands.
- Orphan services get an owner within 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Missing ownership or runbooks blocks Tier 0/1 releases.
7. Adopt the severity scale, declaration rights, and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix. Anyone may declare. Nobody is penalised for over-declaring.
- **SEV1 (crisis):** money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; deploy freeze; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads; business-hours comms; postmortem if customer-detected, over two hours, a repeat, or credit-generating.
- **SEV4:** no customer impact. Ticket only. Never pages.
- Payments-specific anchors: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident become at least SEV2.
- Only the Incident Commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
8. Define incident roles, authority, dual-control, and handover discipline (depends on: 7)
Solve 'nobody in charge for an hour' by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Command claimed within 5 minutes and stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Staff 24x7 with a three-layer coverage model, not 28 night rotations (depends on: 6, 8)
Do not create 28 night rotations. That is precisely what engineers are rejecting. Centralise coordination in a trained command corps and keep technical ownership local.
- **Layer A — Incident Command corps:** approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. Paired with a Communications Lead pool of approximately 18 from Support, CS, and engineering management, and a scribe pool used as the training entry point.
- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat a Lisbon or APAC follow-the-sun cell as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
10. Approve paid on-call, New York labor compliance, and fatigue safeguards (depends on: 4, 9)
Unpaid on-call in New York is both a retention problem and a wage-hour exposure. Pay must be live in payroll before any new mandatory night rotation starts. Publish the actual numbers, then ask people to sign up.
- Indicative scheme locked by HR, Finance, and employment counsel within 14 days: approximately $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as approximately 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
- Publish on-call load by team quarterly.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
11. Set the alert-quality standard and a hard page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target: 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, noise below 15%, with no loss of Tier 0/1 detection coverage.
12. Detect payment and ledger failures before customers do (depends on: 6, 11)
The goal is blunt: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
- Run the replay test: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish replay coverage as a leading metric.
13. Consolidate to one paging platform, one incident record, one status page (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
14. Codify the five-minute escalation path and live execution doctrine (depends on: 8, 9, 13)
Write one unskippable path from 'something looks wrong' to 'someone is in charge'. The default action is never waiting. If nobody claims command within 5 minutes, the platform assigns it and announces it.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute.
- Separate diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Closure requires a stability observation window and explicit handback.
15. Set the service-readiness bar and write major-incident playbooks (depends on: 6, 9, 12)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human at night. The shared ledger cluster is the single largest structural risk.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction: partitioning, read replicas, isolation of non-critical readers, with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
16. Run one communications clock for internals, customers, and regulators (depends on: 7, 8, 13)
Replace 'whoever is around' with one timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure. State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is 'not reportable'.
- Obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
17. Tie incidents to SLA credits and true financial cost (depends on: 7, 16)
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
18. Make blameless postmortems mandatory with one format and fixed deadlines (depends on: 7, 8)
Replace 'some incidents, various formats' with one mandatory format, fixed deadlines, and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did?
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and are never the commander of the incident under review.
- Publish a searchable library and a quarterly top-five recurring causes analysis, with restricted versions for security or privileged content.
19. Enforce action ownership, reserved capacity, and tracking (depends on: 13, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity. Closing a ticket without evidence does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
- Target: 90% of high-priority actions closed by due date within two quarters.
20. Train and certify every role before independent duty (depends on: 8, 14, 16, 18)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it. The certification register is an audit artefact.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days plus shadowing): command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
21. Rehearse with tabletops, game days, and unannounced drills (depends on: 13, 15, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-mitigation.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
22. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 14, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 10, 12, 13, 15, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse.
- Publish a one-page result to the whole company.
24. Instrument metrics, dashboards, review cadence, and anti-gaming (depends on: 13, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
25. Roll out in risk-ordered waves with readiness gates (depends on: 23, 24)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, never waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave. Give every team its ranked list of noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Review missed detections monthly with the same seriousness as noise.
- Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
26. Run change management, fairness, and pager culture from day one (depends on: 4, 10)
Run this in parallel from day one. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
27. Produce SOC 2 evidence by operating, test internally, and mock-audit (depends on: 22, 24, 25)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence. Documented exceptions beat a claim of perfection.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5, CC2.2/CC2.3, CC5, and A1.2. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including 'not reportable', postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in months 4 and 6, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
28. Maintain the program risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and pre-commit the response. Review monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
29. Inspect at 90 days, lock year-two ownership, and prevent decay (depends on: 24, 25, 27)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 4 ---
Proposal ID: 91ea9817-3664-46dd-a758-888405cea636
Content:
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- By day 7, every suspected major incident uses one record, one channel, and a named commander; Support can declare without engineering confirmation.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of customer-visible SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (26):
1. Charter the program, fund it, and start the audit clock
Turn the CEO email into a chartered company program within 48 hours. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability as full-time owner, with a three-person office: program lead, platform engineer, and analyst.
- Form a small decision group of Engineering, SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit. It proposes. The sponsor decides within 48 hours.
- Lock the non-negotiables now: one severity scale, one human-paging platform, one incident record, one postmortem format, mandatory action tracking, paid on-call, and named service ownership.
- Confirm the SOC 2 Type II observation window with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, exercises, and reserved engineering capacity against the $1.3M in SLA credits.
- Reserve 15% of engineering capacity for detection, runbooks, and incident actions, protected by the sponsor.
- Publish a one-page charter on day 2. Clock: floor by day 7, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for policy, tooling, or the audit. Put a crude but real process in place this week so the next outage already has an owner.
- Publish a one-page interim severity card and one declaration path: chat command, phone number, and current pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline, and one naming convention for every suspected major incident.
- Triage the 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
- Replay one of the two nobody-in-charge incidents as a tabletop within 14 days using this floor process.
3. Rebuild the forensic baseline and incident-replay catalog (depends on: 1)
Rebuild the facts before locking design. This is the design input, the frozen before-picture for the CEO, and the test set for detection work.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact missing signal. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Name the top 50 noisy rules and every rule with no owner or runbook.
- Quantify true cost: credits, failed payment volume, delayed value, reconciliation breaks, and engineering hours lost.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11 of 64 actions closed, on-call in 12 of 28 teams.
- Build a **replay catalog**: for each historical incident, the detector that should now fire, at which minute, and the owner of the gap.
4. Publish the on-call fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders. Do not punish them by making them the company's permanent night watch.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the **fairness contract** in writing: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive exemption path for health, disability, or caregiving.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and journey tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has one owner. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Record owning team, manager, chat channel, escalation policy, dashboards, runbook, dependencies, regions, and data stores for all 180 services.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, refund, report, onboard — to services, data stores, regions, sponsor banks, and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform. They fail differently.
- Orphan services get an owner in 30 days or an approved decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Coverage follows the journey, not the org chart. Small teams that own Tier 0 pieces get headcount, service reassignment, or membership in a domain rotation. Never a two-person 24x7 rota.
6. Lock severity levels, integrity flags, and the incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Classify on actual or credible customer, financial, security, and contractual harm, never on who reported it or how hard the fix looks.
- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 contained: narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages a human.
- Attach an **integrity, security, settlement, or regulatory flag** to any severity. The flag forces dual-control, Legal, and the reportability checkpoint without inventing a fifth level. A one-customer ledger corruption is still a flagged crisis.
- Auto-escalate to at least SEV2: any ledger-cluster event, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident.
- Anyone may declare. Nobody is punished for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- Communications Lead: single voice for the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists. A human validates at SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory text. The commander owns the facts.
- Distinct people for command, comms, and technical lead at SEV1. Comms and scribe may combine only for bounded SEV2.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
- Every role assignment and handover is announced verbally and in writing with the exact time and open risks.
8. Staff 24x7 with command, domains, and a Duty Triage Desk (depends on: 4, 5, 7)
Do not create 28 night rotations. That is what engineers are rejecting. Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified people from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to eight months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate ownership into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- **Duty Triage Desk**: a small paid overnight first-line rotation that owns the first ten minutes of ambiguous, unowned, or low-confidence pages. It verifies, enriches, and applies only the runbook's safe steps, then wakes the owning domain. It never sits on a suspected ledger, payment-halt, or security event. Those page commander and likely domains immediately, in parallel.
- Overnight comms: SEV1 pages a Communications Lead 24x7. Customer-visible SEV2 lets the commander publish the first status from a template; Comms is paged if the incident is still open at 30 minutes or a top-100 account is affected.
- Staffing math: 260 engineers can sustain 10–12 domain rotations, one command corps, and one triage desk. They cannot sustain 28 night rotas. Role exclusivity: nobody is primary on two rotations in the same week. Commanders may also be domain responders in different weeks.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
- Evaluate follow-the-sun as a 12-month option, not a year-one dependency. Seed Layer B from the 12 teams already on-call.
9. Pay on-call, meet New York labor rules, and cap fatigue (depends on: 4, 8)
Unpaid on-call in New York is a retention problem and a wage-hour exposure. Pay must be in payroll before any new mandatory night rotation starts.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, $400–$800 for Comms or triage-desk duty, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- **Hard gate: no mandatory night rotation starts before compensation is live in payroll.** Interim duty is paid retroactively. Budget roughly $0.9M–$1.2M a year, then refine with actual rotation count.
10. Enforce an alert-quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers, then replay the past (depends on: 5, 10)
Stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. A detector is not done until it would have caught the last 12 months.
- Define SLIs and SLOs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner, processor, and sponsor-bank notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a named defect class with a mandatory tracked action.
- Run the **replay test** against the S3 catalog. For each of the 31 incidents, name the detector that would now fire and at which minute. Close gaps the replay exposes before calling detection improved.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One chat command creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy. Unowned pages go to the Duty Triage Desk and Duty Command, and log a catalog defect.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Codify escalation, the five-minute command rule, and vendor incidents (depends on: 7, 8, 12)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting. Payments also fail at processors and banks, which you cannot patch.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Vendor-incident class: processor, sponsor-bank, card-network, or cloud-control-plane failure. Command is still required. Success is time-to-customer-notice, time-to-failover-decision, and queue management, not root cause at the vendor.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Write live execution doctrine, readiness bars, and ledger playbooks (depends on: 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving root cause. A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis. Split diagnosis and mitigation workstreams once staffing allows.
- Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Drain backlogs under control. Complete reconciliation before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date. Existing detection stays on while gaps are repaired.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers, with board-visible milestones and executive-signed RPO/RTO.
15. Run one communications clock for internals, customers, and account managers (depends on: 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for customer-visible SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys, not internal service names. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer record.
- Language rules: state impact and the next update time. **Never speculate** on cause, recovery time, data integrity, or blame.
16. Operationalize regulatory notice, partner clocks, and SLA credits (depends on: 6, 15)
In payments, some incidents start a legal clock at detection. Tie incidents to money so Finance does not learn about outages from invoices.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when not reportable**, with facts, decision-maker, and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families. Track failed payment volume, delayed value, and reconciliation breaks as the true incident cost.
17. Make postmortems mandatory and actions enforceable (depends on: 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth. Eleven of 64 closed is a process nobody enforces.
- Mandatory for every SEV1 and SEV2, any incident a customer detected first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and gap, timeline, response analysis, contributing conditions, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items. Facilitators are trained and are never the commander of the incident under review.
- Every action gets one named owner, priority, due date, expected risk reduction, verification method, and a ticket created automatically. If it is not on the board, it does not exist.
- Classes: containment in 7 days, corrective in 30, strategic in 90. SEV1 recurrence-prevention enters the next sprint ahead of roadmap work. Reserve 15% capacity.
- Overdue ladder: manager at +7 days, director at +14, CTO at +30. Overdue high-risk items need written residual-risk acceptance and can block related releases. Verify effectiveness before closing.
18. Train and certify every response role before independent duty (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Duty Triage Desk, half day: enrichment, safe-step limits, when to wake a domain immediately, when not to delay.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
19. Publish Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17)
Collapse the design into a document people will actually open mid-outage, and make it official. Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
20. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 18, 19)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, Duty Triage Desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
21. Instrument metrics, reviews, and anti-gaming (depends on: 12, 17, 20)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- Response: time from first impact to detect, declare, commander, acknowledge, mitigate, and resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
22. Rehearse with tabletops, game days, and night drills (depends on: 12, 14, 18, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced overnight paging drill to measure real acknowledgement times.
- One combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO, before the audit.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
23. Burn down alert noise and roll out by risk with gates (depends on: 20, 21)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Noise reduction is a quota inside each wave, not a background hope.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy paging paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Give every team its noisiest rules from the baseline. Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed. Pair every suppression with a compensating-detection check.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
- If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
24. Operate the fairness and culture program in parallel (depends on: 4, 9)
Run this from the moment the deal is published. Engineers judge the process on fairness. Executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
25. Prove SOC 2 operating effectiveness before fieldwork (depends on: 19, 21, 22, 23)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (depends on: 21, 23, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, metrics, and the ledger blast-radius workstream, independent of the audit cycle.
- Pre-committed contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 5 ---
Proposal ID: b427ef6b-42c2-41ac-8f9d-992d99ba5e20
Content:
Estimated Complexity: high
Success Metrics: - Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (30):
1. Charter the program and start the audit clock
Turn the CEO email into a chartered company program with one accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
- Confirm the SOC 2 Type II observation window with the auditor; map the future response controls to applicable Trust Services Criteria; define evidence set and retention requirements.
4. Listening tour, resistance map and the written on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (depends on: 5, 7)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 4, 8)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (depends on: 5, 10)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (depends on: 7, 13)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 12, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
- Agree with Legal and Finance the availability measurement method per contract and per component; compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide credit posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain; attribute credits to root-cause families for investment decisions.
18. Blameless postmortem standard and Incident Review Board (depends on: 6, 7, 12)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
19. Action ownership, reserved capacity and enforcement (depends on: 12, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
20. Training and certification academy (depends on: 7, 13, 15, 18)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
21. Exercise programme: tabletops, game days and unannounced drills (depends on: 12, 14, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
22. Publish Incident Management Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
24. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
25. Alert noise burn-down campaign (depends on: 10, 12, 23)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
26. Wave rollout to all 28 teams with readiness gates (depends on: 23, 24, 25)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
27. Change management, fairness and pager culture (depends on: 4, 9, 23)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
28. SOC 2 evidence by design, internal testing and mock audit (depends on: 22, 24, 26)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
29. Risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
30. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 24, 26, 28)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Instructions:
1. Carefully evaluate each proposal considering:
- Clarity and detail of the proposed steps.
- Overall alignment with the Main Objective.
- Quality and relevance of the proposed success metrics.
- Coherence of dependencies between steps (if any).
2. Think deeply about each proposal, analyzing its strengths and weaknesses. Consider how well each proposal addresses the main objective and which one would be most effective.
After your thorough analysis, give the Proposal ID of the selected proposal and a one- or two-sentence justification of your choice.
{"selected_proposal_id": "30f15b26-0ec0-4756-9f4a-6279eb889884", "justification": "Proposal 1 covers every element of the brief with the most operational depth, and it uniquely adds items the others omit or bury: week-one audit/legal and evidence scoping, change-intelligence and deploy-safety on the critical path (the main lever on a 3h10 MTTM), Support/account management as an instrumented detection tier, a multi-incident concurrency doctrine, and an incentive guardrail warning that credits may rise before they fall.\n\nIts staffing answer to the pager objection is the strongest: three coverage layers plus a paid overnight Duty Triage Desk, with published arithmetic showing why 28 night rotations are rejected and roughly 100 of 260 engineers carry a night duty at one week in six, gated on compensation being live in payroll.\n\nThe success metrics are specific, baseline-anchored and mostly leading rather than lagging (replay coverage, page budget, action closure, mock-audit outcomes), and the 35-step dependency graph runs coherently from a seven-day operating floor through pilot, waves, control test and mock audit before the SOC 2 window. The only real cost is length; the competing proposals are near-identical in structure but consistently thinner on these differentiators."} |
| gpt5.6-sol_voter_2 | gpt5.6-sol · openai/gpt-5.6-sol |
Proposal B · gpt5.6-sol_refine_2 | Proposal 2 provides the best balance of completeness, operational clarity, and feasibility. It covers every required control with coherent dependencies, sustainable paid domain-based coverage, careful alert-suppression safeguards, precise communication rules, and an evidence-producing rollout that begins immediately rather than waiting for the audit. |
49.4k in · 2.2k out · 35 s | show[SYSTEM]
You are an expert and objective evaluator of project plan proposals.
Your task is to select the BEST proposal based on criteria of completeness, clarity, and alignment with the main objective.
Use your internal reasoning processes to thoroughly analyze each proposal, considering all aspects and implications.
Take as much time and space as you need to evaluate each proposal in depth before making your decision.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
Proposals to Evaluate:
--- PROPOSAL 1 ---
Proposal ID: 30f15b26-0ec0-4756-9f4a-6279eb889884
Content:
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90 a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95% internally and externally.
- Top-100 account outreach completed within 30 minutes of SEV1 in 95% of qualifying cases, using the approved briefing pack.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- Change-correlated incidents are identified within 10 minutes of declaration in 90% of cases, and the change-correlated share of incidents declines quarter over quarter.
- SLA credits fall to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, one concurrency drill and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and a tested read-only mode by month 6, with milestones reported to the board quarterly.
Steps (35):
1. Charter the programme: one owner, one mandate, funded, dated against the audit
Turn the CEO email into a chartered company programme with a single accountable owner and authority over all 28 teams. Incident response stops being a per-team preference and becomes a **company operating process**.
- Name the CTO as executive sponsor and a full-time Head of Reliability & Incident Management as accountable owner, supported by a programme office of three: programme lead, incident-platform engineer, reliability analyst.
- Form a decision group (Engineering, Platform/SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables on day one: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, named service ownership, and **paid on-call**.
- Publish the timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal control test month 5, mock audit month 7, SOC 2 fieldwork month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.3M/yr, training and exercises, plus 15% of engineering capacity reserved for reliability work.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence immediately.
- Publish a one-page interim severity card and a single declaration path: one chat command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim 24x7 Duty Incident Commander roster (primary plus backup) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document and one naming convention for every major incident, starting now.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is the design input, the executive narrative and the frozen "before" picture for the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the **detection backlog**.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Correlate incidents with deployments, config changes and feature flags to quantify how many started with a change we made.
- Quantify true cost beyond credits: failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost, churn risk on named accounts.
- Freeze the baselines in a signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Audit, legal and evidence scoping in week one (depends on: 1)
Design evidence as a by-product of operating, never as a reconstruction before fieldwork. Settle the scope and the legal handling of incident records now, not in month seven.
- Confirm with the external auditor the Type II observation window, the incident population definition, and the evidence they will sample. Everything from the day-7 floor onwards must count.
- Map incident response to the Trust Services Criteria with Compliance: CC7.2–CC7.5 (monitoring, identification, response, recovery), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), A1.2 (availability).
- Define the evidence set and where it is produced automatically: incident records, paging and acknowledgement logs, role assignments, status-page history, reportability decisions, postmortems, action closure with verification, training register, drill records, access reviews.
- Agree retention, confidentiality, legal hold and access rules. Decide with counsel which postmortem content is privileged and how privileged material is segregated **without making the ordinary postmortem secret**.
- Start the obligation matrix with Legal: NYDFS 23 NYCRR 500, state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network contractual windows, cyber-insurer notice.
- Begin monthly evidence sampling from month 1 so control operation is visible long before the audit.
5. Listening tour, resistance map and the written on-call deal (depends on: 1, 3)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written deal signed by the sponsor.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
- State the hard gate publicly: no new mandatory night rotation starts before compensation, training, runbooks and staffing rules are live.
6. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- Record for every service: one accountable team, engineering manager, chat channel, escalation policy, dashboards, runbook link, dependency list, regions and data stores.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record SLO, RTO, RPO, active-active or region-bound, failover method, and the dependencies that make nominal two-region redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
7. Severity standard, declaration rights and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning responders and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and responders paged, comms and scribe if customer-visible, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, credit-generating or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Attach a financial-integrity flag or a security flag to any severity. The flag forces dual control, Legal engagement and the reportability checkpoint without inventing a fifth level.
- Anchor on payments reality alongside error rates: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including a customer report.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
8. Roles, authority, concurrency and handover discipline (depends on: 7)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice to the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1 and SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, Finance and Vendor Management join on defined triggers rather than by invitation.
- Rules: command claimed within 5 minutes and stated in channel; distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- **Concurrency doctrine:** two simultaneous SEV1/SEV2 incidents activate the secondary commander and a surge roster; a designated Multi-Incident Coordinator arbitrates shared resources such as the ledger, the database platform and the deploy freeze.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
9. Three-layer 24x7 coverage: command corps, domain rotations, overnight triage desk (depends on: 5, 6, 8)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving roughly one primary week per person every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement and reconciliation, auth, API edge, Kubernetes platform, data and reporting, partner integrations). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Publish the staffing arithmetic: 10–12 domain rotations of six to eight people plus a 30-person command corps means roughly 100 of 260 engineers carry a night obligation, at about one week in six. Twenty-eight independent rotations would be unstaffable and is therefore rejected on the numbers.
- Teams too small for a fair rotation get headcount, service reassignment, or a time-limited executive exception. **Never a two-person 24x7 rotation.**
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: a human acknowledgement within 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
10. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 5, 9)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour, and a premium for the overnight triage desk.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per responder per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty since day 7 is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons without career penalty, and on-call load is published quarterly by team.
11. Alert quality contract and page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression can never hide an incident.
12. Detection uplift on the money path, validated by incident replay (depends on: 6, 11)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action and a missed-detection review.
13. Change intelligence and deployment safety on the critical path (depends on: 6, 12)
Most of these incidents start with something we changed. Make change the first hypothesis the tooling answers, and make changes safer to reverse.
- Stream every deployment, configuration change, feature-flag flip, schema migration and infrastructure change into the incident timeline with service labels and owners.
- Give the commander an automatic "what changed in the last 60 minutes on the affected journey" panel at declaration time.
- Require Tier 0/1 changes to be progressively delivered with a documented rollback that is tested, time-bounded and executable by the on-call responder without the author.
- Treat ledger schema migrations and settlement-affecting changes as a separate class: dual approval, rehearsed rollback, no deploys inside the settlement window.
- Enforce a change freeze during SEV1 and SEV2, lifted only by the commander and logged.
- Report change-correlated incidents monthly; a rising ratio is a signal to strengthen release safety, not to blame a team.
14. One pager, one incident record, one status page — migrated without a detection gap (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy paging path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access, MFA and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page delivered outside this platform creates no on-call obligation.
15. Escalation ladder and the five-minute command rule (depends on: 8, 9, 14)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts and support entitlements: AWS and database premium support, processors, sponsor banks, card networks.
16. Live execution doctrine and payments safety rules (depends on: 8, 14, 15)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed in advance rather than improvised.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and automatic reopening if impact recurs.
17. Service readiness bar, major-incident playbooks and ledger blast-radius reduction (depends on: 6, 9, 16)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human, and the shared ledger is the largest structural risk in the estate.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as a concentration risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers, stronger regional independence — with board-visible quarterly milestones.
- Enforcement: no readiness sign-off, no night paging alerts unless the manager accepts the gap in writing with an expiry date and a compensating control. Never respond to a gap by turning detection off.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
18. Internal communications protocol (depends on: 8, 14)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, Sales and Security.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change. First internal brief within 10 minutes of SEV1 and 15 of SEV2.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
19. Customer communications, status page and account-manager outreach (depends on: 7, 18)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; monitoring notice after mitigation; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, monitoring, resolution.
- Component-level status page mapped to customer journeys rather than internal service names, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1 and 60 of SEV2, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state observed impact, affected capabilities, any workaround and the next update time; **never speculate on cause, recovery time, data integrity or blame**.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
- Record any legally required restriction or delay of public detail, its approver, and the alternative stakeholder plan.
20. Support and account management as a detection and intake tier (depends on: 7, 12, 19)
Customers detected 40% of incidents first, which means the front line already holds the signal. Turn Support and account managers into an instrumented detection channel rather than a bystander.
- Give Support explicit declaration rights, a one-page trigger card, and a macro that opens an incident candidate directly in the incident platform.
- Automate the clustering rule: five similar tickets or calls in ten minutes auto-creates a triage incident assigned to the duty commander.
- Route credible partner, processor and sponsor-bank notifications into the same declaration path within five minutes.
- Generate the affected-customer list automatically from journey telemetry and the incident record, and push it to Support, CS and the account-manager briefing.
- Train Support and account managers on approved language and prohibit independent technical explanations to customers.
- Measure and publish "signal was in Support before it was in monitoring" as a detection defect, and feed each instance into the detection backlog.
21. Regulatory, partner and legal notification playbook (depends on: 4, 7, 19)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Complete the obligation matrix started in scoping and have counsel validate the triggers, deadlines, channels and submitting authority for each obligation.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with facts considered, decision-maker, timestamp and reassessment trigger.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, card networks, outside counsel, cyber-insurer, critical vendors.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the operational facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly inside the exercise programme, including one full regulator-notification simulation per year.
22. SLA credit workflow, true cost model and incentive guardrails (depends on: 7, 19)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component, including how partial degradation counts.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Guardrail against the perverse incentive: better detection will surface incidents that previously went unbilled, so credits may rise before they fall. Publish this expectation to the executive team in advance, and make it a written rule that no finance or commercial pressure may influence severity or declaration.
23. Blameless postmortem standard and Incident Review Board (depends on: 4, 7, 8)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger; any control failure.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning engineering director is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions where security, privacy or privileged content requires it.
24. Action ownership, reserved capacity and enforcement (depends on: 14, 23)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, accountable manager, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness through tests, telemetry, exercises or production evidence before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
25. Training and certification academy (depends on: 8, 15, 18, 23)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, the 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, contractual clocks, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days of joining a domain.
- Certification is valid 12 months and renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
26. Exercise programme: tabletops, game days, night drills and vendor rehearsals (depends on: 14, 17, 25)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander. First company-wide tabletop within 30 days.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times, plus a concurrency drill with two simultaneous incidents.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down. Rehearse joint escalation with AWS, a processor and a sponsor bank.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
27. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 15, 18, 19, 21, 23, 24)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, escalation and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix, Readiness Bar.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
28. Pilot on the payments critical path (depends on: 10, 12, 13, 14, 17, 25, 27)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement and reconciliation, PostgreSQL platform, Kubernetes platform, API edge, auth and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, action tracking, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The programme lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time in 95% of cases, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
29. Alert noise burn-down campaign (depends on: 11, 14, 28)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → under 1,500 pages a month by day 90, under 500 with noise below 15% by month 6.
30. Metrics, dashboards, review cadence and anti-gaming (depends on: 14, 23, 28)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time from first impact to detect, declare, commander, acknowledge, mitigate, resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors, change-correlated incident ratio.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
31. Wave rollout to all 28 teams with readiness gates (depends on: 28, 30)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the programme office for three weeks; the director signs the gate.
- **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard by team, tier and control gap.
32. Change management, fairness and pager culture (depends on: 5, 10, 28)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
33. SOC 2 evidence by design, internal control test and mock audit (depends on: 4, 27, 30, 31)
The real process is the evidence. Never build a parallel audit process, and never reconstruct records after the fact.
- Maintain the evidence set defined in scoping, produced automatically and indexed: versioned policies and exceptions, catalog records, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Sample incidents monthly from first signal through verified corrective action, deliberately including customer-reported events, downgraded incidents and missed timelines.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling end to end and reporting gaps to the sponsor.
- Run a formal mock audit in month 7 using the populations, evidence requests and interviews the auditor will use: a commander, a random engineer, Support, Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
34. Programme risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps and immediate staffing intervention.
- **A major SEV1 mid-rollout:** the programme lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
35. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 30, 31, 33)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, exercise calendar and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and quarterly board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 2 ---
Proposal ID: 1f1527bb-d992-451c-bbc4-9e61ffdb1f30
Content:
Estimated Complexity: high
Success Metrics: - By day 7, every suspected major incident uses one record, one coordination channel, and a named Incident Commander within 10 minutes.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes.
- By month 3, at least 95% of SEV1 and SEV2 incidents have a named commander within 5 minutes.
- By month 3, at least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and appropriate technical coverage.
- By week 18, all 180 services have an owner, tier, escalation path, and tested coverage model.
- No new mandatory night rotation begins before compensation, training, access, runbooks, and minimum staffing are active.
- Every direct 24x7 technical rotation has at least six qualified responders or an approved, expiring executive exception.
- No responder is routinely primary more often than one week in six or assigned to two simultaneous primary rotations.
- Median impact-to-detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- By day 90, incident replay identifies a current detector for at least 90% of the 31 historical incidents, including its expected detection minute.
- Median time to mitigation falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of applicable SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet the required internal and customer update cadence by month 3.
- Monthly human notification episodes fall from the current 3,400 alert events to no more than 1,500 by day 90 and 500 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without reducing Tier 0 or Tier 1 replay coverage.
- Average after-hours load remains at or below two notification episodes per responder per week; every sustained breach receives a dated remediation plan.
- All six monitoring sources route human pages through the controlled paging platform by week 12, with direct legacy routes retired by week 18.
- All required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 from month 3 onward.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due dates with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Reportability is assessed and recorded within 1 hour for 100% of SEV1 and qualifying SEV2 incidents, including not-reportable decisions.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Approved ledger RTO and RPO, controlled failover, read-only mode, and post-recovery reconciliation are exercised before month 6.
- Quarterly responder surveys reach at least 75% favorable responses for fairness, ownership boundaries, compensation, and sustainability by month 6.
- Monthly contracted availability meets or exceeds 99.95% by month 6 using the contractually authoritative measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- The month-7 mock audit finds no unowned high-risk control gap and at least 95% of sampled incidents contain complete operating evidence.
- The SOC 2 Type II incident-response controls complete external testing without an unresolved material exception.
Steps (28):
1. Charter the program and fund immediate action
Make incident management a **company operating process** within 48 hours. Give one accountable leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Include Engineering, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit in a decision group. The sponsor resolves blocked decisions within 48 hours.
- Approve non-negotiables: one severity model, one human-paging path, one incident record, paid on-call, named service ownership, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, resilience, and incident actions.
- Fund tooling, compensation, training, exercises, and program staff. Compare the cost with the existing $1.3M in annual SLA credits.
- Set the schedule: operating floor by day 7, design during weeks 1–4, pilot during weeks 5–10, rollout during weeks 11–18, control tests in months 4 and 6, and mock audit in month 7.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for the final policy or new tooling. Put a minimum viable process into operation immediately and retain evidence from the first day.
- Publish a one-page interim severity guide, declaration procedure, role card, and communications clock.
- Provide one monitored declaration route through chat, telephone, and the current paging environment.
- Create one channel, bridge, timeline, and incident identifier for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced managers and engineers.
- Require a named commander within 10 minutes. The duty engineering director assumes command if the command page is unclaimed.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Stop uncompensated mandatory after-hours expansion. Pay interim duty under a temporary stipend, retroactive to program launch.
- Hold a 15-minute daily operational review until the permanent process is active.
3. Create the factual, contractual, and control baseline (depends on: 1)
Build one defensible baseline for process design, investment decisions, and SOC 2 testing. Preserve the original data so progress cannot be created by changing definitions.
- Reconstruct all 31 incidents from impact start through detection, declaration, command, mitigation, recovery, communications, credits, and corrective actions.
- Reconstruct the two incidents with unclear command minute by minute.
- Identify the missing signal for every customer-first detection.
- Inventory all six alert sources, 3,400 monthly alert events, duplicates, noisy rules, missing owners, and missing runbooks.
- Record current rotations, unpaid duty, overnight activations, schedule size, and uncovered services.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, 85% noise, $1.3M in credits, and 11 of 64 actions closed.
- Inventory customer-specific availability definitions, notice periods, credit terms, sponsor-bank obligations, and other incident-related contracts.
- Confirm the required SOC 2 Type II observation period and evidence expectations with the auditor during week 1.
4. Publish the on-call fairness contract (depends on: 1)
Treat pager resistance as a legitimate design constraint. Adoption depends on a written agreement that separates command from technical ownership.
- Interview representatives from all 28 teams and from Support, Customer Success, Security, and Operations.
- Separate concerns about unpaid work, sleep loss, unfamiliar systems, noisy alerts, inadequate runbooks, and blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for systems they own or have formally accepted and trained to support, and assisted by a separate Incident Commander.
- State that Platform may temporarily triage unknown ownership but does not inherit another team's service.
- Provide confidential accommodations for health, disability, pregnancy, or caregiving constraints without career penalty.
- Measure baseline trust, fairness, fatigue, and alert confidence. Repeat at days 60 and 120, then quarterly.
5. Build the service catalog and customer-journey map (depends on: 3)
Make a machine-readable catalog the source of truth for routing, impact analysis, status-page components, and control evidence. Every production service must have one accountable owner.
- Record the owning team, manager, business capability, repository, channel, dashboard, runbook, escalation policy, dependencies, regions, and data stores for all 180 services.
- Map initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, reporting, and onboarding to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing systems that can create material contractual impact.
- Classify Tier 2 as deferrable internal or batch systems and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional recovery mode, contractual commitments, and critical vendors for Tier 0 and Tier 1.
- Keep separate but coordinated ownership for the ledger application and PostgreSQL platform.
- Give orphan services an owner or approved decommission date within 30 days. Treat unowned Tier 0 services as release-blocking executive risks.
6. Adopt the severity model and incident modifiers (depends on: 3, 5)
Classify incidents by credible customer, financial, security, regulatory, and contractual harm. Start at the higher plausible severity while scope or integrity remains unknown.
- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; broad loss of a core payment journey; both regions impaired; or processing must be stopped. Immediately page all response roles and the Executive Duty Officer. Open the bridge within 5 minutes, freeze unrelated changes, issue internal notice within 10 minutes, publish applicable customer status within 15 minutes, start legal assessment within 1 hour, and require a postmortem.
- **SEV2 — major:** material payment degradation; settlement deadline at risk; regional impairment with reduced resilience; a critical customer or material cohort unavailable; or an SLA breach is likely. Page command and technical roles immediately. Issue internal notice within 15 minutes, applicable customer status within 30 minutes, and require a postmortem.
- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads. Page only when immediate action can reduce harm.
- **SEV4 — operational event:** no current customer impact and no credible imminent harm. Create a ticket and handle during normal hours.
- Add FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, and VENDOR modifiers. These invoke specialist controls without distorting customer-impact severity.
- Use at least SEV2 posture for unknown impact lasting 15 minutes, credible ledger-integrity risk, or a cross-domain incident without clear ownership.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Permit anyone to declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Publish a decision tree and examples based on the 31 historical incidents.
7. Define roles, authority, and handoffs (depends on: 6)
Separate coordination, communications, recordkeeping, and technical repair. One named person holds command continuously throughout every SEV1 and SEV2.
- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not act as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, executive updates, and coordination with Legal.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- **Subject-Matter Responders:** diagnose and mitigate only systems for which they have ownership, access, training, or a formally accepted support agreement.
- **Executive Duty Officer:** removes organizational obstacles and makes exceptional business decisions without displacing the commander.
- Add Security, Legal, Compliance, Finance, and Vendor Management on modifier-specific triggers.
- Require distinct commander, Communications Lead, scribe, and technical lead for SEV1. Communications and scribe may combine for the first 10 minutes of a bounded SEV2.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, reconciliation, and change evidence for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, including the exact transfer time and unresolved risks.
8. Create sustainable 24x7 coverage (depends on: 4, 5, 7)
Use central command coverage and risk-based technical coverage instead of creating 28 fragile night rotations. Responders do not carry pagers for unfamiliar code.
- Create a 24x7 Incident Command corps of approximately 30 certified people with primary and backup schedules. Two schedules create about 104 weekly assignments a year, or roughly three to four weeks per person annually.
- Create a 24x7 Communications pool of 18–24 trained people from Support, Customer Operations, Engineering Operations, and management.
- Create a similarly sized scribe pool. The backup commander temporarily records the first minutes if a scribe has not joined.
- Maintain a 24x7 Executive Duty Officer schedule and specialist contact paths for Security and Legal.
- Group Tier 0 and Tier 1 systems into roughly 8–12 coherent response domains only where responders share training, access, runbooks, and explicit support acceptance.
- Staff every critical domain with primary and secondary responders and at least six qualified people. Target eight where overnight activation is frequent.
- Give Tier 2 and Tier 3 services business-hours ownership plus tested manager and director escalation.
- Reclassify any lower-tier service capable of causing severe overnight harm rather than hiding the risk behind manager callback.
- Route unknown-owner incidents to Duty Command and Platform temporarily. Record each occurrence as a catalog control defect.
- Prohibit simultaneous primary assignments and two-person 24x7 rotations.
9. Implement compensation and fatigue protections (depends on: 4, 8)
End unpaid on-call before expanding mandatory coverage. Use fixed duty compensation so responders are not rewarded for alert volume.
- Use planning bands of $900–$1,200 per Tier 0/1 primary week and $300–$500 per secondary week.
- Use planning bands of $1,000–$1,300 per Duty Commander week and $400–$700 for Communications or scribe primary duty.
- Pay holiday premiums. Compensate all legally compensable active and waiting time for non-exempt staff, including overtime where required.
- Have HR, Finance, Payroll, and employment counsel approve final bands, tax handling, FLSA classification, New York wage-hour treatment, and schedule constraints within 14 days.
- Provide a protected recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal delivery commitments by about 15% during a primary week.
- Prohibit consecutive primary weeks, on-call during leave, hidden schedule swaps, and primary duty more often than one week in six.
- Allow responders to declare temporary fatigue-related unfitness without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Make active payroll setup, training, access, and readiness hard gates before any new mandatory night rotation starts.
10. Codify the live incident lifecycle (depends on: 6, 7)
Give every incident the same operational flow from first signal through verified recovery. The first objective is limiting customer and financial harm, not proving root cause.
- Use the states Detected, Declared, Triaged, Mitigating, Mitigated, Monitoring, Resolved, and Reviewed.
- Record impact start separately from detection. Use the earliest defensible evidence and revise it transparently when later facts emerge.
- Open with a standard command message: severity, known impact, assigned roles, immediate objective, workstreams, and next update time.
- Freeze unrelated production changes during SEV1 and normally during SEV2. Record every exception.
- Prefer reversible mitigation: rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, partner rerouting, or controlled processing suspension.
- Separate mitigation and diagnosis workstreams when staffing allows.
- Keep decisions in the shared incident record rather than direct messages.
- Require formal command and technical handoffs for shift changes, fatigue, or incidents exceeding four hours.
- For payment incidents, verify backlog handling, duplicate protection, customer state, settlement exposure, and ledger reconciliation before resolution.
- Require a severity-specific stability period and explicit handback to the owning team, Support, and Customer Success.
11. Establish one paging and incident system of record (depends on: 5, 7, 8)
Monitoring tools may remain specialized, but every human page and major-incident record must enter one controlled platform. This provides consistent routing and an audit trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page within two weeks.
- Ingest events from all six monitoring tools before disabling their direct human-paging routes.
- Route pages using the service catalog and deduplicate events belonging to the same symptom.
- Provide one declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, decisions, severity changes, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status publishing, and conference facilities.
- Apply MFA, role-based access, periodic access reviews, and tamper-evident history.
- Keep privileged legal or security material in restricted linked records rather than exposing it in the general timeline.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, the status provider, or an AWS region.
- Retire each legacy paging route only after ownership review, end-to-end testing, and two weeks of verified operation.
12. Enforce the alert-quality contract (depends on: 3, 5)
Treat every human page as a production interface with an owner and a required action. Measure human notification episodes rather than raw monitoring events.
- Require every paging rule to identify the service, owner, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Define an actionable page as one that causes or materially informs a timely intervention or risk decision.
- Define noise as duplicate, non-urgent, unactionable, stale, test-generated, or incorrectly routed notification.
- Page on payment outcomes, error-budget burn, queue age against deadlines, and financial-integrity risk rather than raw CPU, memory, pod, or log thresholds.
- Run new rules in shadow mode for seven days unless a documented emergency exception applies.
- Review repeated no-action pages within two business days.
- Set a page budget of no more than two after-hours notification episodes per responder per week, measured over four weeks.
- Make a sustained budget breach trigger a tuning sprint and block additional non-emergency paging rules.
- Require compensating detection and central approval before suppressing a Tier 0 or Tier 1 rule.
- Never disable an existing critical detector solely because its metadata or runbook is incomplete. Track the gap with a dated remediation owner.
13. Detect payment failures before customers (depends on: 5, 12)
Move detection from infrastructure health to customer journeys and ledger truth. Validate coverage against actual historical failures.
- Define SLIs and internal SLOs for initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, APIs, webhooks, and reporting freshness.
- Set internal objectives with enough headroom to protect the contractual 99.95% availability commitment.
- Run external synthetic transactions through critical journeys at least once per minute from paths independent of the production platform.
- Test each region and expose dependencies that defeat nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed payment value relative to settlement deadlines.
- Add tenant and cohort anomaly detection for high-value customers and critical payment methods.
- Convert credible Support, account-manager, processor, bank, and network reports into incident candidates within five minutes.
- Replay all 31 historical incidents. Record which current detector would fire and at what minute.
- Treat customer-first detection as a mandatory missed-detection review with a tracked action.
14. Implement the detection and escalation ladder (depends on: 7, 8, 11)
Create one time-bound path from the first credible signal to named command and the correct technical owner. Device delivery does not count as acknowledgement.
- Converge automated alerts, engineer observations, support cases, account-manager reports, partner notices, and customer calls on the same declaration path.
- Page the Duty Incident Commander and owning critical-domain primary immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at 5 minutes, manager at 10 minutes, and director at 15 minutes.
- Escalate unclaimed command to the backup commander at 5 minutes. The Executive Duty Officer assumes temporary command at 10 minutes until a certified transfer occurs.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Keep command with the current commander when service ownership remains unclear. Assign a temporary technical lead and record the ownership gap.
- Maintain tested external escalation routes for AWS, database support, processors, sponsor banks, networks, and critical vendors.
- Test the full declaration, acknowledgement, fallback, conference, and status-publishing path weekly.
- Treat missed acknowledgement, failed routing, and unowned incidents as control failures requiring review.
15. Standardize internal communications (depends on: 7, 10, 11)
Give responders one working room and stakeholders one controlled source of truth. Executives must not interrupt the technical command path.
- Maintain one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue the initial internal notice within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within the same initial-notice window.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record every material decision and outbound message in the incident timeline.
- Require the executive team to sign the communication behavior rules.
16. Standardize customer, account, and regulatory communications (depends on: 3, 6, 15)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate observed impact before root cause is known.
- Publish customer status within 15 minutes of a customer-visible SEV1 and within 30 minutes of a customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map public status components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, available workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Have account managers contact affected strategic accounts within 30 minutes for SEV1 and 60 minutes for SEV2 using the approved briefing.
- Offer status subscriptions to all customers. Auto-enroll only where contracts, consent, and applicable communication rules permit.
- Encode customer-specific notice deadlines and channels in the customer record.
- Have Legal and Compliance maintain a counsel-validated matrix covering applicable NYDFS, breach, GLBA or FTC, PCI, money-transmitter, sponsor-bank, network, insurance, and customer obligations.
- Complete and record a reportability assessment within one hour for every SEV1 and every security, privacy, or integrity-related SEV2, including decisions of not reportable.
- Let Legal own regulatory text and submission while the Incident Commander owns operational facts. Record any legally required restriction of public detail and its alternative stakeholder plan.
17. Tie incidents to SLA credits and financial exposure (depends on: 3, 16)
Link the incident record to contractual and financial outcomes. Finance should not discover outages through later credit claims.
- Define the authoritative availability calculation for each contract and customer journey with Legal and Finance.
- Calculate affected customers and minutes from incident scope and journey telemetry.
- Produce a preliminary credit and contractual-exposure estimate within five business days of resolution.
- Record failed payment count, failed or delayed value, settlement exposure, reconciliation breaks, support effort, and engineering effort.
- Establish a documented approval path for proactive credits and claims-based credits.
- Attribute credits and financial harm to recurring failure families.
- Use the quarterly credit analysis to prioritize detection, resilience, and architectural investment.
18. Establish mandatory blameless postmortems (depends on: 6, 7)
Use one learning standard with fixed deadlines. Keep learning separate from disciplinary and misconduct processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact lasting more than two hours, SLA credits, contractual breach, repeated contributing factors, major control failures, and ledger-integrity near misses.
- Produce a factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the owning engineering director accountable for completion. The commander owns response analysis, and the scribe supplies the timeline.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit answers to why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the commander or primary technical responder.
- Describe decisions using the context and information available at the time. Do not name an individual as the root cause.
- Keep HR, misconduct, and personnel matters in separate processes.
- Publish broadly useful findings internally while restricting security, privacy, personnel, and privileged content appropriately.
19. Make corrective actions enforceable risk commitments (depends on: 18)
An action is not complete when its ticket is closed. It is complete when the intended risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within 7 days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Prefer actions that remove hazards or reduce blast radius over vague actions such as retraining or adding monitoring.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of engineering capacity for approved reliability work.
- Escalate overdue high-risk actions to the manager after 7 days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit unrepaired severe conditions to block related releases.
- Verify effectiveness using tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days as complete, re-planned, superseded with evidence, or formally risk-accepted.
20. Set the service-readiness bar and payments playbooks (depends on: 5, 10, 12, 13)
A critical service must be supportable at 3 a.m. before it enters direct overnight coverage. Existing critical detection remains active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and integrity constraints.
- Require responders to demonstrate access and safe execution before independent primary duty.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional impairment, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document split-brain prevention, replay protection, failover, controlled backlog recovery, and post-recovery reconciliation.
- Require executive-approved RTO and RPO for the shared ledger cluster.
- Exercise critical runbooks at least twice a year and after material changes.
- Block new Tier 0 or Tier 1 releases and paging rules when readiness requirements are missing.
- Handle existing gaps using named owners, compensating controls, executive-approved expiry dates, and remediation plans.
- Run a parallel architecture workstream to reduce ledger blast radius, isolate non-critical readers, and strengthen regional independence.
21. Train and certify every response role (depends on: 7, 14, 15, 18, 20)
Command and communication are learned skills. Use paid working time and certify people before independent duty.
- Give all employees a 30-minute module on recognizing impact, declaring incidents, and locating the status page.
- Give responders a half-day module on severity, acknowledgement, escalation, evidence, runbooks, and financial-integrity precautions.
- Train scribes for two hours on timeline quality and fact-versus-hypothesis labeling.
- Train Incident Commanders for two days on delegation, uncertainty, severity, mitigation strategy, fatigue, handoffs, and executive management.
- Require commander candidates to complete a simulated SEV1 and two shadowed incidents or exercises.
- Train Communications Leads for one day on status writing, customer segmentation, legal boundaries, and contractual clocks.
- Require domain responders to demonstrate dashboards, access, rollback, failover, escalation, and relevant playbooks.
- Require two shadow shifts before independent primary duty.
- Renew certification annually through simulation.
- Maintain the training, assessment, and certification register as operational and audit evidence.
- Nominate an incident-management champion in each of the 28 teams.
22. Exercise command, recovery, and tool failure (depends on: 11, 20, 21)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first cross-company command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous operational and security events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors when relevant.
- Conduct at least one unannounced after-hours paging test before the audit and two annually thereafter.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure acknowledgement, command, customer notice, mitigation decision, handoff, and recovery times.
- Create tracked corrective actions for every material exercise finding.
23. Publish the signed policy and control set (depends on: 6, 7, 8, 9, 10, 12, 14, 15, 16, 18, 19)
Convert the design into concise documents that people can use during an incident. The actual operating process must also be the documented and audited process.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, evidence, and exception standards.
- Include one-page cards for severity, roles, authority, escalation, and communication timings inside the incident tool.
- State explicitly that responders support only owned or formally accepted and trained service portfolios.
- Include the compensation structure, fatigue rules, declaration rights, and non-retaliation commitment.
- Obtain approval from the CTO, HR, Legal, Security, Compliance, and Internal Audit.
- Announce the policy at an all-hands and through team briefings.
- Create an exception register with owner, rationale, compensating control, approver, review date, and expiry.
- Version every policy change. Do not rewrite historical records when the process changes.
24. Instrument the scorecard and review forums (depends on: 3, 11, 18, 19)
Measure the process before the pilot so failures become visible immediately. Report medians and 90th percentiles rather than averages alone.
- Measure impact-to-detection, detection-to-declaration, declaration-to-command, acknowledgement, mitigation, recovery, and resolution.
- Split results by severity, service tier, customer journey, region, detection source, and business-hours status.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track notification episodes, actionability, duplicates, after-hours load, missed detection, routing errors, and page-budget breaches.
- Track postmortem timeliness, action age, due-date performance, verified effectiveness, and repeated contributing factors.
- Track journey availability, error-budget burn, failed or delayed value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight work, recovery days, exceptions, sentiment, and responder attrition.
- Hold a weekly Incident Review Board chaired by the Head of Reliability with relevant directors.
- Hold a monthly executive reliability review and a quarterly control and resilience review with Internal Audit.
- Reconcile incident records monthly against support cases, customer complaints, status history, credits, and major operational anomalies to detect under-reporting.
- Use team-level scorecards to direct help and investment. Never penalize an individual for good-faith declaration.
25. Pilot the complete process on the payment path (depends on: 9, 11, 13, 20, 21, 23, 24)
Run a four-to-six-week pilot across the highest-risk journey before expanding. The interim operating floor remains active for the rest of the company.
- Include payment orchestration, ledger application, PostgreSQL platform, API edge, authentication, settlement, reconciliation, Kubernetes platform, and Support intake.
- Include teams with existing on-call experience and teams new to the model.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking together.
- Run legacy and new paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently taking command.
- Correct critical process or tooling defects within 48 hours.
- Exit only after 95% timely command assignment, 95% communications compliance, no unpaid pages, complete required postmortems, tested fallbacks, and at least 50% lower pilot noise.
- Publish the pilot results, defects, and policy changes company-wide.
26. Roll out by risk with readiness gates (depends on: 25)
Expand in controlled waves and finish early enough to accumulate operating evidence before the audit. A calendar date does not override a failed readiness gate.
- Roll out remaining Tier 0 domains first, followed by Tier 1, Tier 2, and Tier 3.
- Use four waves of six to eight teams, each lasting two to three weeks.
- Gate each service on catalog ownership, appropriate coverage, active compensation, trained responders, access, tested escalation, alert quality, runbooks, and a passed tabletop.
- Require at least six responders only for direct 24x7 technical rotations. Use business-hours coverage for lower tiers.
- Give each wave a named coach and director sign-off.
- Reschedule failed gates or use a time-limited executive exception with compensating controls. Do not create silent waivers.
- Disable legacy human-paging routes after verified cutover for each wave.
- Run a quota-based noise reduction sprint in every wave, starting with the highest-volume rules.
- Pair every suppression with a compensating-detection check.
- Publish an internal adoption dashboard by team, service tier, coverage, and control gap.
- Complete critical coverage by approximately week 12 and all 28 teams by week 18.
27. Prove SOC 2 operating effectiveness (depends on: 22, 23, 24, 26)
Generate evidence through normal operation rather than reconstructing it before fieldwork. Test both control design and consistent execution.
- Map controls to the applicable Trust Services Criteria with Compliance and the auditor, including monitoring, incident identification, response, recovery, communications, and availability.
- Retain approved policies, exceptions, service ownership, schedules, compensation activation, access reviews, training, incidents, communications, reportability decisions, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, missed timelines, non-reportable decisions, and exercises in the testing population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the evidence populations and interviews expected from the external auditor.
- Correct deviations through tracked actions with owners and dates. Never edit history to create apparent compliance.
- Verify that evidence retention covers the full auditor-defined observation period.
- Brief commanders, engineers, Support, and Compliance on the actual process without scripting inaccurate answers.
28. Inspect, adapt, and institutionalize ownership (depends on: 26, 27)
Prevent the process from decaying after rollout or the audit. Change it using measured operating evidence rather than opinion.
- Review the policy after 90 days of live operation using severity calibration, page load, missed detection, communication compliance, action closure, fatigue, and survey results.
- Remove steps that create work without reducing risk. Add controls only where incidents, exercises, or evidence show a gap.
- Reassess Tier 0 and Tier 1 classification and domain boundaries every six months.
- Assign permanent owners for policy, catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review compensation bands, rotation burden, accommodations, and staffing annually.
- Report severe incidents, credits, overdue high-risk actions, and ledger concentration risk to the board or risk committee quarterly.
- Maintain the ledger blast-radius program as an executive risk until failover, degraded mode, reconciliation, and regional independence meet approved objectives.
- Evaluate follow-the-sun coverage using one year of actual activation and staffing data.
- Build year-two plans for automated mitigation, safer deployments, graceful degradation, and error-budget release controls.
--- PROPOSAL 3 ---
Proposal ID: f61aca40-a8a9-4e91-a5f3-ac1d03634441
Content:
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of 'not reportable'.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (29):
1. Charter the program, fund it, and start the audit clock
Convert the CEO email into a company operating process with one accountable owner, a budget, and a dated timeline that starts this week.
- Name the CTO as executive sponsor and appoint a Head of Reliability & Incident Management as the single accountable owner with full-time authority over all 28 teams.
- Stand up a three-person program office: program lead, platform engineer, reliability analyst.
- Form an eight-person steering group spanning Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, and HR. It proposes; the sponsor decides within 48 hours. Never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one human-paging path, one incident record, one postmortem format, paid on-call, mandatory action tracking, and named service ownership.
- Approve budget anchored against the $1.3M in SLA credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs, and 15% reserved engineering capacity.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Publish a one-page charter company-wide on day 2. Incident response is a company process, not a per-team preference.
- Timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day interim command floor (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place this week so the next outage has a named commander and the evidence clock starts immediately.
- Publish a one-page interim severity card and one declaration path: a Slack command, a phone number, and the existing pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs from the 12 teams already on-call. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical action items into complete, re-plan, or formally risk-accept within 30 days, prioritising ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline of incidents, alerts, and money lost (depends on: 1)
Rebuild the facts before locking any design. This is both the design input and the frozen 'before' picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines in one signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the written fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written, signed deal.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- State explicitly that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Map compliance, evidence, and audit requirements from day one (depends on: 1)
Design evidence as a by-product of operations, not a reconstruction before the auditor arrives. The interim process in week 1 is already evidence.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for internal and external communication; CC5 for control activities; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure with verification, training register, drill records, reportability decisions including 'not reportable'.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records. Require role-based access, MFA, and periodic access review.
- Version and approve all policy documents from day one: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard, Alert Quality Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has a named owner. A wrong owner recreates the pager objection. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Assign one accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. **Tier 1**: customer-facing but degradable. **Tier 2**: internal or batch. **Tier 3**: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for every Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform. They fail differently and need different hands.
- Orphan services get an owner within 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Missing ownership or runbooks blocks Tier 0/1 releases.
7. Adopt the severity scale, declaration rights, and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix. Anyone may declare. Nobody is penalised for over-declaring.
- **SEV1 (crisis):** money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; deploy freeze; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads; business-hours comms; postmortem if customer-detected, over two hours, a repeat, or credit-generating.
- **SEV4:** no customer impact. Ticket only. Never pages.
- Payments-specific anchors: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident become at least SEV2.
- Only the Incident Commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
8. Define incident roles, authority, dual-control, and handover discipline (depends on: 7)
Solve 'nobody in charge for an hour' by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Command claimed within 5 minutes and stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Staff 24x7 with a three-layer coverage model, not 28 night rotations (depends on: 6, 8)
Do not create 28 night rotations. That is precisely what engineers are rejecting. Centralise coordination in a trained command corps and keep technical ownership local.
- **Layer A — Incident Command corps:** approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. Paired with a Communications Lead pool of approximately 18 from Support, CS, and engineering management, and a scribe pool used as the training entry point.
- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat a Lisbon or APAC follow-the-sun cell as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
10. Approve paid on-call, New York labor compliance, and fatigue safeguards (depends on: 4, 9)
Unpaid on-call in New York is both a retention problem and a wage-hour exposure. Pay must be live in payroll before any new mandatory night rotation starts. Publish the actual numbers, then ask people to sign up.
- Indicative scheme locked by HR, Finance, and employment counsel within 14 days: approximately $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as approximately 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
- Publish on-call load by team quarterly.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
11. Set the alert-quality standard and a hard page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target: 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, noise below 15%, with no loss of Tier 0/1 detection coverage.
12. Detect payment and ledger failures before customers do (depends on: 6, 11)
The goal is blunt: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
- Run the replay test: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish replay coverage as a leading metric.
13. Consolidate to one paging platform, one incident record, one status page (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
14. Codify the five-minute escalation path and live execution doctrine (depends on: 8, 9, 13)
Write one unskippable path from 'something looks wrong' to 'someone is in charge'. The default action is never waiting. If nobody claims command within 5 minutes, the platform assigns it and announces it.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute.
- Separate diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Closure requires a stability observation window and explicit handback.
15. Set the service-readiness bar and write major-incident playbooks (depends on: 6, 9, 12)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human at night. The shared ledger cluster is the single largest structural risk.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction: partitioning, read replicas, isolation of non-critical readers, with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
16. Run one communications clock for internals, customers, and regulators (depends on: 7, 8, 13)
Replace 'whoever is around' with one timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure. State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is 'not reportable'.
- Obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
17. Tie incidents to SLA credits and true financial cost (depends on: 7, 16)
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
18. Make blameless postmortems mandatory with one format and fixed deadlines (depends on: 7, 8)
Replace 'some incidents, various formats' with one mandatory format, fixed deadlines, and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did?
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and are never the commander of the incident under review.
- Publish a searchable library and a quarterly top-five recurring causes analysis, with restricted versions for security or privileged content.
19. Enforce action ownership, reserved capacity, and tracking (depends on: 13, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity. Closing a ticket without evidence does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
- Target: 90% of high-priority actions closed by due date within two quarters.
20. Train and certify every role before independent duty (depends on: 8, 14, 16, 18)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it. The certification register is an audit artefact.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days plus shadowing): command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
21. Rehearse with tabletops, game days, and unannounced drills (depends on: 13, 15, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-mitigation.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
22. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 14, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 10, 12, 13, 15, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse.
- Publish a one-page result to the whole company.
24. Instrument metrics, dashboards, review cadence, and anti-gaming (depends on: 13, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
25. Roll out in risk-ordered waves with readiness gates (depends on: 23, 24)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, never waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave. Give every team its ranked list of noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Review missed detections monthly with the same seriousness as noise.
- Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
26. Run change management, fairness, and pager culture from day one (depends on: 4, 10)
Run this in parallel from day one. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
27. Produce SOC 2 evidence by operating, test internally, and mock-audit (depends on: 22, 24, 25)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence. Documented exceptions beat a claim of perfection.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5, CC2.2/CC2.3, CC5, and A1.2. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including 'not reportable', postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in months 4 and 6, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
28. Maintain the program risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and pre-commit the response. Review monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
29. Inspect at 90 days, lock year-two ownership, and prevent decay (depends on: 24, 25, 27)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 4 ---
Proposal ID: 91ea9817-3664-46dd-a758-888405cea636
Content:
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- By day 7, every suspected major incident uses one record, one channel, and a named commander; Support can declare without engineering confirmation.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of customer-visible SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (26):
1. Charter the program, fund it, and start the audit clock
Turn the CEO email into a chartered company program within 48 hours. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability as full-time owner, with a three-person office: program lead, platform engineer, and analyst.
- Form a small decision group of Engineering, SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit. It proposes. The sponsor decides within 48 hours.
- Lock the non-negotiables now: one severity scale, one human-paging platform, one incident record, one postmortem format, mandatory action tracking, paid on-call, and named service ownership.
- Confirm the SOC 2 Type II observation window with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, exercises, and reserved engineering capacity against the $1.3M in SLA credits.
- Reserve 15% of engineering capacity for detection, runbooks, and incident actions, protected by the sponsor.
- Publish a one-page charter on day 2. Clock: floor by day 7, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for policy, tooling, or the audit. Put a crude but real process in place this week so the next outage already has an owner.
- Publish a one-page interim severity card and one declaration path: chat command, phone number, and current pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline, and one naming convention for every suspected major incident.
- Triage the 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
- Replay one of the two nobody-in-charge incidents as a tabletop within 14 days using this floor process.
3. Rebuild the forensic baseline and incident-replay catalog (depends on: 1)
Rebuild the facts before locking design. This is the design input, the frozen before-picture for the CEO, and the test set for detection work.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact missing signal. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Name the top 50 noisy rules and every rule with no owner or runbook.
- Quantify true cost: credits, failed payment volume, delayed value, reconciliation breaks, and engineering hours lost.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11 of 64 actions closed, on-call in 12 of 28 teams.
- Build a **replay catalog**: for each historical incident, the detector that should now fire, at which minute, and the owner of the gap.
4. Publish the on-call fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders. Do not punish them by making them the company's permanent night watch.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the **fairness contract** in writing: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive exemption path for health, disability, or caregiving.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and journey tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has one owner. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Record owning team, manager, chat channel, escalation policy, dashboards, runbook, dependencies, regions, and data stores for all 180 services.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, refund, report, onboard — to services, data stores, regions, sponsor banks, and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform. They fail differently.
- Orphan services get an owner in 30 days or an approved decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Coverage follows the journey, not the org chart. Small teams that own Tier 0 pieces get headcount, service reassignment, or membership in a domain rotation. Never a two-person 24x7 rota.
6. Lock severity levels, integrity flags, and the incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Classify on actual or credible customer, financial, security, and contractual harm, never on who reported it or how hard the fix looks.
- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 contained: narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages a human.
- Attach an **integrity, security, settlement, or regulatory flag** to any severity. The flag forces dual-control, Legal, and the reportability checkpoint without inventing a fifth level. A one-customer ledger corruption is still a flagged crisis.
- Auto-escalate to at least SEV2: any ledger-cluster event, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident.
- Anyone may declare. Nobody is punished for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- Communications Lead: single voice for the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists. A human validates at SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory text. The commander owns the facts.
- Distinct people for command, comms, and technical lead at SEV1. Comms and scribe may combine only for bounded SEV2.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
- Every role assignment and handover is announced verbally and in writing with the exact time and open risks.
8. Staff 24x7 with command, domains, and a Duty Triage Desk (depends on: 4, 5, 7)
Do not create 28 night rotations. That is what engineers are rejecting. Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified people from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to eight months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate ownership into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- **Duty Triage Desk**: a small paid overnight first-line rotation that owns the first ten minutes of ambiguous, unowned, or low-confidence pages. It verifies, enriches, and applies only the runbook's safe steps, then wakes the owning domain. It never sits on a suspected ledger, payment-halt, or security event. Those page commander and likely domains immediately, in parallel.
- Overnight comms: SEV1 pages a Communications Lead 24x7. Customer-visible SEV2 lets the commander publish the first status from a template; Comms is paged if the incident is still open at 30 minutes or a top-100 account is affected.
- Staffing math: 260 engineers can sustain 10–12 domain rotations, one command corps, and one triage desk. They cannot sustain 28 night rotas. Role exclusivity: nobody is primary on two rotations in the same week. Commanders may also be domain responders in different weeks.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
- Evaluate follow-the-sun as a 12-month option, not a year-one dependency. Seed Layer B from the 12 teams already on-call.
9. Pay on-call, meet New York labor rules, and cap fatigue (depends on: 4, 8)
Unpaid on-call in New York is a retention problem and a wage-hour exposure. Pay must be in payroll before any new mandatory night rotation starts.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, $400–$800 for Comms or triage-desk duty, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- **Hard gate: no mandatory night rotation starts before compensation is live in payroll.** Interim duty is paid retroactively. Budget roughly $0.9M–$1.2M a year, then refine with actual rotation count.
10. Enforce an alert-quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers, then replay the past (depends on: 5, 10)
Stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. A detector is not done until it would have caught the last 12 months.
- Define SLIs and SLOs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner, processor, and sponsor-bank notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a named defect class with a mandatory tracked action.
- Run the **replay test** against the S3 catalog. For each of the 31 incidents, name the detector that would now fire and at which minute. Close gaps the replay exposes before calling detection improved.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One chat command creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy. Unowned pages go to the Duty Triage Desk and Duty Command, and log a catalog defect.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Codify escalation, the five-minute command rule, and vendor incidents (depends on: 7, 8, 12)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting. Payments also fail at processors and banks, which you cannot patch.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Vendor-incident class: processor, sponsor-bank, card-network, or cloud-control-plane failure. Command is still required. Success is time-to-customer-notice, time-to-failover-decision, and queue management, not root cause at the vendor.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Write live execution doctrine, readiness bars, and ledger playbooks (depends on: 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving root cause. A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis. Split diagnosis and mitigation workstreams once staffing allows.
- Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Drain backlogs under control. Complete reconciliation before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date. Existing detection stays on while gaps are repaired.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers, with board-visible milestones and executive-signed RPO/RTO.
15. Run one communications clock for internals, customers, and account managers (depends on: 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for customer-visible SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys, not internal service names. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer record.
- Language rules: state impact and the next update time. **Never speculate** on cause, recovery time, data integrity, or blame.
16. Operationalize regulatory notice, partner clocks, and SLA credits (depends on: 6, 15)
In payments, some incidents start a legal clock at detection. Tie incidents to money so Finance does not learn about outages from invoices.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when not reportable**, with facts, decision-maker, and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families. Track failed payment volume, delayed value, and reconciliation breaks as the true incident cost.
17. Make postmortems mandatory and actions enforceable (depends on: 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth. Eleven of 64 closed is a process nobody enforces.
- Mandatory for every SEV1 and SEV2, any incident a customer detected first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and gap, timeline, response analysis, contributing conditions, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items. Facilitators are trained and are never the commander of the incident under review.
- Every action gets one named owner, priority, due date, expected risk reduction, verification method, and a ticket created automatically. If it is not on the board, it does not exist.
- Classes: containment in 7 days, corrective in 30, strategic in 90. SEV1 recurrence-prevention enters the next sprint ahead of roadmap work. Reserve 15% capacity.
- Overdue ladder: manager at +7 days, director at +14, CTO at +30. Overdue high-risk items need written residual-risk acceptance and can block related releases. Verify effectiveness before closing.
18. Train and certify every response role before independent duty (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Duty Triage Desk, half day: enrichment, safe-step limits, when to wake a domain immediately, when not to delay.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
19. Publish Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17)
Collapse the design into a document people will actually open mid-outage, and make it official. Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
20. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 18, 19)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, Duty Triage Desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
21. Instrument metrics, reviews, and anti-gaming (depends on: 12, 17, 20)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- Response: time from first impact to detect, declare, commander, acknowledge, mitigate, and resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
22. Rehearse with tabletops, game days, and night drills (depends on: 12, 14, 18, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced overnight paging drill to measure real acknowledgement times.
- One combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO, before the audit.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
23. Burn down alert noise and roll out by risk with gates (depends on: 20, 21)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Noise reduction is a quota inside each wave, not a background hope.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy paging paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Give every team its noisiest rules from the baseline. Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed. Pair every suppression with a compensating-detection check.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
- If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
24. Operate the fairness and culture program in parallel (depends on: 4, 9)
Run this from the moment the deal is published. Engineers judge the process on fairness. Executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
25. Prove SOC 2 operating effectiveness before fieldwork (depends on: 19, 21, 22, 23)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (depends on: 21, 23, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, metrics, and the ledger blast-radius workstream, independent of the audit cycle.
- Pre-committed contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 5 ---
Proposal ID: b427ef6b-42c2-41ac-8f9d-992d99ba5e20
Content:
Estimated Complexity: high
Success Metrics: - Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (30):
1. Charter the program and start the audit clock
Turn the CEO email into a chartered company program with one accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
- Confirm the SOC 2 Type II observation window with the auditor; map the future response controls to applicable Trust Services Criteria; define evidence set and retention requirements.
4. Listening tour, resistance map and the written on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (depends on: 5, 7)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 4, 8)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (depends on: 5, 10)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (depends on: 7, 13)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 12, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
- Agree with Legal and Finance the availability measurement method per contract and per component; compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide credit posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain; attribute credits to root-cause families for investment decisions.
18. Blameless postmortem standard and Incident Review Board (depends on: 6, 7, 12)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
19. Action ownership, reserved capacity and enforcement (depends on: 12, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
20. Training and certification academy (depends on: 7, 13, 15, 18)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
21. Exercise programme: tabletops, game days and unannounced drills (depends on: 12, 14, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
22. Publish Incident Management Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
24. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
25. Alert noise burn-down campaign (depends on: 10, 12, 23)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
26. Wave rollout to all 28 teams with readiness gates (depends on: 23, 24, 25)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
27. Change management, fairness and pager culture (depends on: 4, 9, 23)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
28. SOC 2 evidence by design, internal testing and mock audit (depends on: 22, 24, 26)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
29. Risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
30. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 24, 26, 28)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Instructions:
1. Carefully evaluate each proposal considering:
- Clarity and detail of the proposed steps.
- Overall alignment with the Main Objective.
- Quality and relevance of the proposed success metrics.
- Coherence of dependencies between steps (if any).
2. Think deeply about each proposal, analyzing its strengths and weaknesses. Consider how well each proposal addresses the main objective and which one would be most effective.
After your thorough analysis, give the Proposal ID of the selected proposal and a one- or two-sentence justification of your choice.
{"selected_proposal_id": "1f1527bb-d992-451c-bbc4-9e61ffdb1f30", "justification": "Proposal 2 provides the best balance of completeness, operational clarity, and feasibility. It covers every required control with coherent dependencies, sustainable paid domain-based coverage, careful alert-suppression safeguards, precise communication rules, and an evidence-producing rollout that begins immediately rather than waiting for the audit."} |
| qwen3.8-max_voter_3 | qwen3.8-max · alibaba/qwen3.8-max |
Proposal A · opus5_refine_1 winner | Proposal 1 is the most comprehensive plan, with 35 steps that uniquely include a dedicated change intelligence step, a concurrency doctrine for simultaneous incidents, and an explicit Support-as-detection-tier step, while also carrying the most detailed success metrics (including change-correlated incident tracking and top-100 account outreach). Its dependencies are internally consistent and it addresses every element of the main objective with the greatest operational specificity. |
52.7k in · 3.0k out · 1 min 8 s | show[SYSTEM]
You are an expert and objective evaluator of project plan proposals.
Your task is to select the BEST proposal based on criteria of completeness, clarity, and alignment with the main objective.
Use your internal reasoning processes to thoroughly analyze each proposal, considering all aspects and implications.
Take as much time and space as you need to evaluate each proposal in depth before making your decision.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
Proposals to Evaluate:
--- PROPOSAL 1 ---
Proposal ID: 30f15b26-0ec0-4756-9f4a-6279eb889884
Content:
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90 a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95% internally and externally.
- Top-100 account outreach completed within 30 minutes of SEV1 in 95% of qualifying cases, using the approved briefing pack.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- Change-correlated incidents are identified within 10 minutes of declaration in 90% of cases, and the change-correlated share of incidents declines quarter over quarter.
- SLA credits fall to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, one concurrency drill and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and a tested read-only mode by month 6, with milestones reported to the board quarterly.
Steps (35):
1. Charter the programme: one owner, one mandate, funded, dated against the audit
Turn the CEO email into a chartered company programme with a single accountable owner and authority over all 28 teams. Incident response stops being a per-team preference and becomes a **company operating process**.
- Name the CTO as executive sponsor and a full-time Head of Reliability & Incident Management as accountable owner, supported by a programme office of three: programme lead, incident-platform engineer, reliability analyst.
- Form a decision group (Engineering, Platform/SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables on day one: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, named service ownership, and **paid on-call**.
- Publish the timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal control test month 5, mock audit month 7, SOC 2 fieldwork month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.3M/yr, training and exercises, plus 15% of engineering capacity reserved for reliability work.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence immediately.
- Publish a one-page interim severity card and a single declaration path: one chat command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim 24x7 Duty Incident Commander roster (primary plus backup) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document and one naming convention for every major incident, starting now.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is the design input, the executive narrative and the frozen "before" picture for the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the **detection backlog**.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Correlate incidents with deployments, config changes and feature flags to quantify how many started with a change we made.
- Quantify true cost beyond credits: failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost, churn risk on named accounts.
- Freeze the baselines in a signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Audit, legal and evidence scoping in week one (depends on: 1)
Design evidence as a by-product of operating, never as a reconstruction before fieldwork. Settle the scope and the legal handling of incident records now, not in month seven.
- Confirm with the external auditor the Type II observation window, the incident population definition, and the evidence they will sample. Everything from the day-7 floor onwards must count.
- Map incident response to the Trust Services Criteria with Compliance: CC7.2–CC7.5 (monitoring, identification, response, recovery), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), A1.2 (availability).
- Define the evidence set and where it is produced automatically: incident records, paging and acknowledgement logs, role assignments, status-page history, reportability decisions, postmortems, action closure with verification, training register, drill records, access reviews.
- Agree retention, confidentiality, legal hold and access rules. Decide with counsel which postmortem content is privileged and how privileged material is segregated **without making the ordinary postmortem secret**.
- Start the obligation matrix with Legal: NYDFS 23 NYCRR 500, state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network contractual windows, cyber-insurer notice.
- Begin monthly evidence sampling from month 1 so control operation is visible long before the audit.
5. Listening tour, resistance map and the written on-call deal (depends on: 1, 3)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written deal signed by the sponsor.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
- State the hard gate publicly: no new mandatory night rotation starts before compensation, training, runbooks and staffing rules are live.
6. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- Record for every service: one accountable team, engineering manager, chat channel, escalation policy, dashboards, runbook link, dependency list, regions and data stores.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record SLO, RTO, RPO, active-active or region-bound, failover method, and the dependencies that make nominal two-region redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
7. Severity standard, declaration rights and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning responders and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and responders paged, comms and scribe if customer-visible, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, credit-generating or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Attach a financial-integrity flag or a security flag to any severity. The flag forces dual control, Legal engagement and the reportability checkpoint without inventing a fifth level.
- Anchor on payments reality alongside error rates: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including a customer report.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
8. Roles, authority, concurrency and handover discipline (depends on: 7)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice to the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1 and SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, Finance and Vendor Management join on defined triggers rather than by invitation.
- Rules: command claimed within 5 minutes and stated in channel; distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- **Concurrency doctrine:** two simultaneous SEV1/SEV2 incidents activate the secondary commander and a surge roster; a designated Multi-Incident Coordinator arbitrates shared resources such as the ledger, the database platform and the deploy freeze.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
9. Three-layer 24x7 coverage: command corps, domain rotations, overnight triage desk (depends on: 5, 6, 8)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving roughly one primary week per person every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement and reconciliation, auth, API edge, Kubernetes platform, data and reporting, partner integrations). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Publish the staffing arithmetic: 10–12 domain rotations of six to eight people plus a 30-person command corps means roughly 100 of 260 engineers carry a night obligation, at about one week in six. Twenty-eight independent rotations would be unstaffable and is therefore rejected on the numbers.
- Teams too small for a fair rotation get headcount, service reassignment, or a time-limited executive exception. **Never a two-person 24x7 rotation.**
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: a human acknowledgement within 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
10. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 5, 9)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour, and a premium for the overnight triage desk.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per responder per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty since day 7 is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons without career penalty, and on-call load is published quarterly by team.
11. Alert quality contract and page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression can never hide an incident.
12. Detection uplift on the money path, validated by incident replay (depends on: 6, 11)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action and a missed-detection review.
13. Change intelligence and deployment safety on the critical path (depends on: 6, 12)
Most of these incidents start with something we changed. Make change the first hypothesis the tooling answers, and make changes safer to reverse.
- Stream every deployment, configuration change, feature-flag flip, schema migration and infrastructure change into the incident timeline with service labels and owners.
- Give the commander an automatic "what changed in the last 60 minutes on the affected journey" panel at declaration time.
- Require Tier 0/1 changes to be progressively delivered with a documented rollback that is tested, time-bounded and executable by the on-call responder without the author.
- Treat ledger schema migrations and settlement-affecting changes as a separate class: dual approval, rehearsed rollback, no deploys inside the settlement window.
- Enforce a change freeze during SEV1 and SEV2, lifted only by the commander and logged.
- Report change-correlated incidents monthly; a rising ratio is a signal to strengthen release safety, not to blame a team.
14. One pager, one incident record, one status page — migrated without a detection gap (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy paging path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access, MFA and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page delivered outside this platform creates no on-call obligation.
15. Escalation ladder and the five-minute command rule (depends on: 8, 9, 14)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts and support entitlements: AWS and database premium support, processors, sponsor banks, card networks.
16. Live execution doctrine and payments safety rules (depends on: 8, 14, 15)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed in advance rather than improvised.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and automatic reopening if impact recurs.
17. Service readiness bar, major-incident playbooks and ledger blast-radius reduction (depends on: 6, 9, 16)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human, and the shared ledger is the largest structural risk in the estate.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as a concentration risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers, stronger regional independence — with board-visible quarterly milestones.
- Enforcement: no readiness sign-off, no night paging alerts unless the manager accepts the gap in writing with an expiry date and a compensating control. Never respond to a gap by turning detection off.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
18. Internal communications protocol (depends on: 8, 14)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, Sales and Security.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change. First internal brief within 10 minutes of SEV1 and 15 of SEV2.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
19. Customer communications, status page and account-manager outreach (depends on: 7, 18)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; monitoring notice after mitigation; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, monitoring, resolution.
- Component-level status page mapped to customer journeys rather than internal service names, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1 and 60 of SEV2, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state observed impact, affected capabilities, any workaround and the next update time; **never speculate on cause, recovery time, data integrity or blame**.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
- Record any legally required restriction or delay of public detail, its approver, and the alternative stakeholder plan.
20. Support and account management as a detection and intake tier (depends on: 7, 12, 19)
Customers detected 40% of incidents first, which means the front line already holds the signal. Turn Support and account managers into an instrumented detection channel rather than a bystander.
- Give Support explicit declaration rights, a one-page trigger card, and a macro that opens an incident candidate directly in the incident platform.
- Automate the clustering rule: five similar tickets or calls in ten minutes auto-creates a triage incident assigned to the duty commander.
- Route credible partner, processor and sponsor-bank notifications into the same declaration path within five minutes.
- Generate the affected-customer list automatically from journey telemetry and the incident record, and push it to Support, CS and the account-manager briefing.
- Train Support and account managers on approved language and prohibit independent technical explanations to customers.
- Measure and publish "signal was in Support before it was in monitoring" as a detection defect, and feed each instance into the detection backlog.
21. Regulatory, partner and legal notification playbook (depends on: 4, 7, 19)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Complete the obligation matrix started in scoping and have counsel validate the triggers, deadlines, channels and submitting authority for each obligation.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with facts considered, decision-maker, timestamp and reassessment trigger.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, card networks, outside counsel, cyber-insurer, critical vendors.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the operational facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly inside the exercise programme, including one full regulator-notification simulation per year.
22. SLA credit workflow, true cost model and incentive guardrails (depends on: 7, 19)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component, including how partial degradation counts.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Guardrail against the perverse incentive: better detection will surface incidents that previously went unbilled, so credits may rise before they fall. Publish this expectation to the executive team in advance, and make it a written rule that no finance or commercial pressure may influence severity or declaration.
23. Blameless postmortem standard and Incident Review Board (depends on: 4, 7, 8)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger; any control failure.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning engineering director is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions where security, privacy or privileged content requires it.
24. Action ownership, reserved capacity and enforcement (depends on: 14, 23)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, accountable manager, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness through tests, telemetry, exercises or production evidence before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
25. Training and certification academy (depends on: 8, 15, 18, 23)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, the 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, contractual clocks, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days of joining a domain.
- Certification is valid 12 months and renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
26. Exercise programme: tabletops, game days, night drills and vendor rehearsals (depends on: 14, 17, 25)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander. First company-wide tabletop within 30 days.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times, plus a concurrency drill with two simultaneous incidents.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down. Rehearse joint escalation with AWS, a processor and a sponsor bank.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
27. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 15, 18, 19, 21, 23, 24)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, escalation and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix, Readiness Bar.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
28. Pilot on the payments critical path (depends on: 10, 12, 13, 14, 17, 25, 27)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement and reconciliation, PostgreSQL platform, Kubernetes platform, API edge, auth and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, action tracking, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The programme lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time in 95% of cases, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
29. Alert noise burn-down campaign (depends on: 11, 14, 28)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → under 1,500 pages a month by day 90, under 500 with noise below 15% by month 6.
30. Metrics, dashboards, review cadence and anti-gaming (depends on: 14, 23, 28)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time from first impact to detect, declare, commander, acknowledge, mitigate, resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors, change-correlated incident ratio.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
31. Wave rollout to all 28 teams with readiness gates (depends on: 28, 30)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the programme office for three weeks; the director signs the gate.
- **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard by team, tier and control gap.
32. Change management, fairness and pager culture (depends on: 5, 10, 28)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
33. SOC 2 evidence by design, internal control test and mock audit (depends on: 4, 27, 30, 31)
The real process is the evidence. Never build a parallel audit process, and never reconstruct records after the fact.
- Maintain the evidence set defined in scoping, produced automatically and indexed: versioned policies and exceptions, catalog records, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Sample incidents monthly from first signal through verified corrective action, deliberately including customer-reported events, downgraded incidents and missed timelines.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling end to end and reporting gaps to the sponsor.
- Run a formal mock audit in month 7 using the populations, evidence requests and interviews the auditor will use: a commander, a random engineer, Support, Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
34. Programme risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps and immediate staffing intervention.
- **A major SEV1 mid-rollout:** the programme lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
35. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 30, 31, 33)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, exercise calendar and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and quarterly board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 2 ---
Proposal ID: 1f1527bb-d992-451c-bbc4-9e61ffdb1f30
Content:
Estimated Complexity: high
Success Metrics: - By day 7, every suspected major incident uses one record, one coordination channel, and a named Incident Commander within 10 minutes.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes.
- By month 3, at least 95% of SEV1 and SEV2 incidents have a named commander within 5 minutes.
- By month 3, at least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and appropriate technical coverage.
- By week 18, all 180 services have an owner, tier, escalation path, and tested coverage model.
- No new mandatory night rotation begins before compensation, training, access, runbooks, and minimum staffing are active.
- Every direct 24x7 technical rotation has at least six qualified responders or an approved, expiring executive exception.
- No responder is routinely primary more often than one week in six or assigned to two simultaneous primary rotations.
- Median impact-to-detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- By day 90, incident replay identifies a current detector for at least 90% of the 31 historical incidents, including its expected detection minute.
- Median time to mitigation falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of applicable SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet the required internal and customer update cadence by month 3.
- Monthly human notification episodes fall from the current 3,400 alert events to no more than 1,500 by day 90 and 500 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without reducing Tier 0 or Tier 1 replay coverage.
- Average after-hours load remains at or below two notification episodes per responder per week; every sustained breach receives a dated remediation plan.
- All six monitoring sources route human pages through the controlled paging platform by week 12, with direct legacy routes retired by week 18.
- All required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 from month 3 onward.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due dates with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Reportability is assessed and recorded within 1 hour for 100% of SEV1 and qualifying SEV2 incidents, including not-reportable decisions.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Approved ledger RTO and RPO, controlled failover, read-only mode, and post-recovery reconciliation are exercised before month 6.
- Quarterly responder surveys reach at least 75% favorable responses for fairness, ownership boundaries, compensation, and sustainability by month 6.
- Monthly contracted availability meets or exceeds 99.95% by month 6 using the contractually authoritative measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- The month-7 mock audit finds no unowned high-risk control gap and at least 95% of sampled incidents contain complete operating evidence.
- The SOC 2 Type II incident-response controls complete external testing without an unresolved material exception.
Steps (28):
1. Charter the program and fund immediate action
Make incident management a **company operating process** within 48 hours. Give one accountable leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Include Engineering, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit in a decision group. The sponsor resolves blocked decisions within 48 hours.
- Approve non-negotiables: one severity model, one human-paging path, one incident record, paid on-call, named service ownership, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, resilience, and incident actions.
- Fund tooling, compensation, training, exercises, and program staff. Compare the cost with the existing $1.3M in annual SLA credits.
- Set the schedule: operating floor by day 7, design during weeks 1–4, pilot during weeks 5–10, rollout during weeks 11–18, control tests in months 4 and 6, and mock audit in month 7.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for the final policy or new tooling. Put a minimum viable process into operation immediately and retain evidence from the first day.
- Publish a one-page interim severity guide, declaration procedure, role card, and communications clock.
- Provide one monitored declaration route through chat, telephone, and the current paging environment.
- Create one channel, bridge, timeline, and incident identifier for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced managers and engineers.
- Require a named commander within 10 minutes. The duty engineering director assumes command if the command page is unclaimed.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Stop uncompensated mandatory after-hours expansion. Pay interim duty under a temporary stipend, retroactive to program launch.
- Hold a 15-minute daily operational review until the permanent process is active.
3. Create the factual, contractual, and control baseline (depends on: 1)
Build one defensible baseline for process design, investment decisions, and SOC 2 testing. Preserve the original data so progress cannot be created by changing definitions.
- Reconstruct all 31 incidents from impact start through detection, declaration, command, mitigation, recovery, communications, credits, and corrective actions.
- Reconstruct the two incidents with unclear command minute by minute.
- Identify the missing signal for every customer-first detection.
- Inventory all six alert sources, 3,400 monthly alert events, duplicates, noisy rules, missing owners, and missing runbooks.
- Record current rotations, unpaid duty, overnight activations, schedule size, and uncovered services.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, 85% noise, $1.3M in credits, and 11 of 64 actions closed.
- Inventory customer-specific availability definitions, notice periods, credit terms, sponsor-bank obligations, and other incident-related contracts.
- Confirm the required SOC 2 Type II observation period and evidence expectations with the auditor during week 1.
4. Publish the on-call fairness contract (depends on: 1)
Treat pager resistance as a legitimate design constraint. Adoption depends on a written agreement that separates command from technical ownership.
- Interview representatives from all 28 teams and from Support, Customer Success, Security, and Operations.
- Separate concerns about unpaid work, sleep loss, unfamiliar systems, noisy alerts, inadequate runbooks, and blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for systems they own or have formally accepted and trained to support, and assisted by a separate Incident Commander.
- State that Platform may temporarily triage unknown ownership but does not inherit another team's service.
- Provide confidential accommodations for health, disability, pregnancy, or caregiving constraints without career penalty.
- Measure baseline trust, fairness, fatigue, and alert confidence. Repeat at days 60 and 120, then quarterly.
5. Build the service catalog and customer-journey map (depends on: 3)
Make a machine-readable catalog the source of truth for routing, impact analysis, status-page components, and control evidence. Every production service must have one accountable owner.
- Record the owning team, manager, business capability, repository, channel, dashboard, runbook, escalation policy, dependencies, regions, and data stores for all 180 services.
- Map initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, reporting, and onboarding to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing systems that can create material contractual impact.
- Classify Tier 2 as deferrable internal or batch systems and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional recovery mode, contractual commitments, and critical vendors for Tier 0 and Tier 1.
- Keep separate but coordinated ownership for the ledger application and PostgreSQL platform.
- Give orphan services an owner or approved decommission date within 30 days. Treat unowned Tier 0 services as release-blocking executive risks.
6. Adopt the severity model and incident modifiers (depends on: 3, 5)
Classify incidents by credible customer, financial, security, regulatory, and contractual harm. Start at the higher plausible severity while scope or integrity remains unknown.
- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; broad loss of a core payment journey; both regions impaired; or processing must be stopped. Immediately page all response roles and the Executive Duty Officer. Open the bridge within 5 minutes, freeze unrelated changes, issue internal notice within 10 minutes, publish applicable customer status within 15 minutes, start legal assessment within 1 hour, and require a postmortem.
- **SEV2 — major:** material payment degradation; settlement deadline at risk; regional impairment with reduced resilience; a critical customer or material cohort unavailable; or an SLA breach is likely. Page command and technical roles immediately. Issue internal notice within 15 minutes, applicable customer status within 30 minutes, and require a postmortem.
- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads. Page only when immediate action can reduce harm.
- **SEV4 — operational event:** no current customer impact and no credible imminent harm. Create a ticket and handle during normal hours.
- Add FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, and VENDOR modifiers. These invoke specialist controls without distorting customer-impact severity.
- Use at least SEV2 posture for unknown impact lasting 15 minutes, credible ledger-integrity risk, or a cross-domain incident without clear ownership.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Permit anyone to declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Publish a decision tree and examples based on the 31 historical incidents.
7. Define roles, authority, and handoffs (depends on: 6)
Separate coordination, communications, recordkeeping, and technical repair. One named person holds command continuously throughout every SEV1 and SEV2.
- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not act as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, executive updates, and coordination with Legal.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- **Subject-Matter Responders:** diagnose and mitigate only systems for which they have ownership, access, training, or a formally accepted support agreement.
- **Executive Duty Officer:** removes organizational obstacles and makes exceptional business decisions without displacing the commander.
- Add Security, Legal, Compliance, Finance, and Vendor Management on modifier-specific triggers.
- Require distinct commander, Communications Lead, scribe, and technical lead for SEV1. Communications and scribe may combine for the first 10 minutes of a bounded SEV2.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, reconciliation, and change evidence for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, including the exact transfer time and unresolved risks.
8. Create sustainable 24x7 coverage (depends on: 4, 5, 7)
Use central command coverage and risk-based technical coverage instead of creating 28 fragile night rotations. Responders do not carry pagers for unfamiliar code.
- Create a 24x7 Incident Command corps of approximately 30 certified people with primary and backup schedules. Two schedules create about 104 weekly assignments a year, or roughly three to four weeks per person annually.
- Create a 24x7 Communications pool of 18–24 trained people from Support, Customer Operations, Engineering Operations, and management.
- Create a similarly sized scribe pool. The backup commander temporarily records the first minutes if a scribe has not joined.
- Maintain a 24x7 Executive Duty Officer schedule and specialist contact paths for Security and Legal.
- Group Tier 0 and Tier 1 systems into roughly 8–12 coherent response domains only where responders share training, access, runbooks, and explicit support acceptance.
- Staff every critical domain with primary and secondary responders and at least six qualified people. Target eight where overnight activation is frequent.
- Give Tier 2 and Tier 3 services business-hours ownership plus tested manager and director escalation.
- Reclassify any lower-tier service capable of causing severe overnight harm rather than hiding the risk behind manager callback.
- Route unknown-owner incidents to Duty Command and Platform temporarily. Record each occurrence as a catalog control defect.
- Prohibit simultaneous primary assignments and two-person 24x7 rotations.
9. Implement compensation and fatigue protections (depends on: 4, 8)
End unpaid on-call before expanding mandatory coverage. Use fixed duty compensation so responders are not rewarded for alert volume.
- Use planning bands of $900–$1,200 per Tier 0/1 primary week and $300–$500 per secondary week.
- Use planning bands of $1,000–$1,300 per Duty Commander week and $400–$700 for Communications or scribe primary duty.
- Pay holiday premiums. Compensate all legally compensable active and waiting time for non-exempt staff, including overtime where required.
- Have HR, Finance, Payroll, and employment counsel approve final bands, tax handling, FLSA classification, New York wage-hour treatment, and schedule constraints within 14 days.
- Provide a protected recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal delivery commitments by about 15% during a primary week.
- Prohibit consecutive primary weeks, on-call during leave, hidden schedule swaps, and primary duty more often than one week in six.
- Allow responders to declare temporary fatigue-related unfitness without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Make active payroll setup, training, access, and readiness hard gates before any new mandatory night rotation starts.
10. Codify the live incident lifecycle (depends on: 6, 7)
Give every incident the same operational flow from first signal through verified recovery. The first objective is limiting customer and financial harm, not proving root cause.
- Use the states Detected, Declared, Triaged, Mitigating, Mitigated, Monitoring, Resolved, and Reviewed.
- Record impact start separately from detection. Use the earliest defensible evidence and revise it transparently when later facts emerge.
- Open with a standard command message: severity, known impact, assigned roles, immediate objective, workstreams, and next update time.
- Freeze unrelated production changes during SEV1 and normally during SEV2. Record every exception.
- Prefer reversible mitigation: rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, partner rerouting, or controlled processing suspension.
- Separate mitigation and diagnosis workstreams when staffing allows.
- Keep decisions in the shared incident record rather than direct messages.
- Require formal command and technical handoffs for shift changes, fatigue, or incidents exceeding four hours.
- For payment incidents, verify backlog handling, duplicate protection, customer state, settlement exposure, and ledger reconciliation before resolution.
- Require a severity-specific stability period and explicit handback to the owning team, Support, and Customer Success.
11. Establish one paging and incident system of record (depends on: 5, 7, 8)
Monitoring tools may remain specialized, but every human page and major-incident record must enter one controlled platform. This provides consistent routing and an audit trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page within two weeks.
- Ingest events from all six monitoring tools before disabling their direct human-paging routes.
- Route pages using the service catalog and deduplicate events belonging to the same symptom.
- Provide one declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, decisions, severity changes, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status publishing, and conference facilities.
- Apply MFA, role-based access, periodic access reviews, and tamper-evident history.
- Keep privileged legal or security material in restricted linked records rather than exposing it in the general timeline.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, the status provider, or an AWS region.
- Retire each legacy paging route only after ownership review, end-to-end testing, and two weeks of verified operation.
12. Enforce the alert-quality contract (depends on: 3, 5)
Treat every human page as a production interface with an owner and a required action. Measure human notification episodes rather than raw monitoring events.
- Require every paging rule to identify the service, owner, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Define an actionable page as one that causes or materially informs a timely intervention or risk decision.
- Define noise as duplicate, non-urgent, unactionable, stale, test-generated, or incorrectly routed notification.
- Page on payment outcomes, error-budget burn, queue age against deadlines, and financial-integrity risk rather than raw CPU, memory, pod, or log thresholds.
- Run new rules in shadow mode for seven days unless a documented emergency exception applies.
- Review repeated no-action pages within two business days.
- Set a page budget of no more than two after-hours notification episodes per responder per week, measured over four weeks.
- Make a sustained budget breach trigger a tuning sprint and block additional non-emergency paging rules.
- Require compensating detection and central approval before suppressing a Tier 0 or Tier 1 rule.
- Never disable an existing critical detector solely because its metadata or runbook is incomplete. Track the gap with a dated remediation owner.
13. Detect payment failures before customers (depends on: 5, 12)
Move detection from infrastructure health to customer journeys and ledger truth. Validate coverage against actual historical failures.
- Define SLIs and internal SLOs for initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, APIs, webhooks, and reporting freshness.
- Set internal objectives with enough headroom to protect the contractual 99.95% availability commitment.
- Run external synthetic transactions through critical journeys at least once per minute from paths independent of the production platform.
- Test each region and expose dependencies that defeat nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed payment value relative to settlement deadlines.
- Add tenant and cohort anomaly detection for high-value customers and critical payment methods.
- Convert credible Support, account-manager, processor, bank, and network reports into incident candidates within five minutes.
- Replay all 31 historical incidents. Record which current detector would fire and at what minute.
- Treat customer-first detection as a mandatory missed-detection review with a tracked action.
14. Implement the detection and escalation ladder (depends on: 7, 8, 11)
Create one time-bound path from the first credible signal to named command and the correct technical owner. Device delivery does not count as acknowledgement.
- Converge automated alerts, engineer observations, support cases, account-manager reports, partner notices, and customer calls on the same declaration path.
- Page the Duty Incident Commander and owning critical-domain primary immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at 5 minutes, manager at 10 minutes, and director at 15 minutes.
- Escalate unclaimed command to the backup commander at 5 minutes. The Executive Duty Officer assumes temporary command at 10 minutes until a certified transfer occurs.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Keep command with the current commander when service ownership remains unclear. Assign a temporary technical lead and record the ownership gap.
- Maintain tested external escalation routes for AWS, database support, processors, sponsor banks, networks, and critical vendors.
- Test the full declaration, acknowledgement, fallback, conference, and status-publishing path weekly.
- Treat missed acknowledgement, failed routing, and unowned incidents as control failures requiring review.
15. Standardize internal communications (depends on: 7, 10, 11)
Give responders one working room and stakeholders one controlled source of truth. Executives must not interrupt the technical command path.
- Maintain one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue the initial internal notice within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within the same initial-notice window.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record every material decision and outbound message in the incident timeline.
- Require the executive team to sign the communication behavior rules.
16. Standardize customer, account, and regulatory communications (depends on: 3, 6, 15)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate observed impact before root cause is known.
- Publish customer status within 15 minutes of a customer-visible SEV1 and within 30 minutes of a customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map public status components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, available workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Have account managers contact affected strategic accounts within 30 minutes for SEV1 and 60 minutes for SEV2 using the approved briefing.
- Offer status subscriptions to all customers. Auto-enroll only where contracts, consent, and applicable communication rules permit.
- Encode customer-specific notice deadlines and channels in the customer record.
- Have Legal and Compliance maintain a counsel-validated matrix covering applicable NYDFS, breach, GLBA or FTC, PCI, money-transmitter, sponsor-bank, network, insurance, and customer obligations.
- Complete and record a reportability assessment within one hour for every SEV1 and every security, privacy, or integrity-related SEV2, including decisions of not reportable.
- Let Legal own regulatory text and submission while the Incident Commander owns operational facts. Record any legally required restriction of public detail and its alternative stakeholder plan.
17. Tie incidents to SLA credits and financial exposure (depends on: 3, 16)
Link the incident record to contractual and financial outcomes. Finance should not discover outages through later credit claims.
- Define the authoritative availability calculation for each contract and customer journey with Legal and Finance.
- Calculate affected customers and minutes from incident scope and journey telemetry.
- Produce a preliminary credit and contractual-exposure estimate within five business days of resolution.
- Record failed payment count, failed or delayed value, settlement exposure, reconciliation breaks, support effort, and engineering effort.
- Establish a documented approval path for proactive credits and claims-based credits.
- Attribute credits and financial harm to recurring failure families.
- Use the quarterly credit analysis to prioritize detection, resilience, and architectural investment.
18. Establish mandatory blameless postmortems (depends on: 6, 7)
Use one learning standard with fixed deadlines. Keep learning separate from disciplinary and misconduct processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact lasting more than two hours, SLA credits, contractual breach, repeated contributing factors, major control failures, and ledger-integrity near misses.
- Produce a factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the owning engineering director accountable for completion. The commander owns response analysis, and the scribe supplies the timeline.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit answers to why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the commander or primary technical responder.
- Describe decisions using the context and information available at the time. Do not name an individual as the root cause.
- Keep HR, misconduct, and personnel matters in separate processes.
- Publish broadly useful findings internally while restricting security, privacy, personnel, and privileged content appropriately.
19. Make corrective actions enforceable risk commitments (depends on: 18)
An action is not complete when its ticket is closed. It is complete when the intended risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within 7 days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Prefer actions that remove hazards or reduce blast radius over vague actions such as retraining or adding monitoring.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of engineering capacity for approved reliability work.
- Escalate overdue high-risk actions to the manager after 7 days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit unrepaired severe conditions to block related releases.
- Verify effectiveness using tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days as complete, re-planned, superseded with evidence, or formally risk-accepted.
20. Set the service-readiness bar and payments playbooks (depends on: 5, 10, 12, 13)
A critical service must be supportable at 3 a.m. before it enters direct overnight coverage. Existing critical detection remains active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and integrity constraints.
- Require responders to demonstrate access and safe execution before independent primary duty.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional impairment, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document split-brain prevention, replay protection, failover, controlled backlog recovery, and post-recovery reconciliation.
- Require executive-approved RTO and RPO for the shared ledger cluster.
- Exercise critical runbooks at least twice a year and after material changes.
- Block new Tier 0 or Tier 1 releases and paging rules when readiness requirements are missing.
- Handle existing gaps using named owners, compensating controls, executive-approved expiry dates, and remediation plans.
- Run a parallel architecture workstream to reduce ledger blast radius, isolate non-critical readers, and strengthen regional independence.
21. Train and certify every response role (depends on: 7, 14, 15, 18, 20)
Command and communication are learned skills. Use paid working time and certify people before independent duty.
- Give all employees a 30-minute module on recognizing impact, declaring incidents, and locating the status page.
- Give responders a half-day module on severity, acknowledgement, escalation, evidence, runbooks, and financial-integrity precautions.
- Train scribes for two hours on timeline quality and fact-versus-hypothesis labeling.
- Train Incident Commanders for two days on delegation, uncertainty, severity, mitigation strategy, fatigue, handoffs, and executive management.
- Require commander candidates to complete a simulated SEV1 and two shadowed incidents or exercises.
- Train Communications Leads for one day on status writing, customer segmentation, legal boundaries, and contractual clocks.
- Require domain responders to demonstrate dashboards, access, rollback, failover, escalation, and relevant playbooks.
- Require two shadow shifts before independent primary duty.
- Renew certification annually through simulation.
- Maintain the training, assessment, and certification register as operational and audit evidence.
- Nominate an incident-management champion in each of the 28 teams.
22. Exercise command, recovery, and tool failure (depends on: 11, 20, 21)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first cross-company command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous operational and security events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors when relevant.
- Conduct at least one unannounced after-hours paging test before the audit and two annually thereafter.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure acknowledgement, command, customer notice, mitigation decision, handoff, and recovery times.
- Create tracked corrective actions for every material exercise finding.
23. Publish the signed policy and control set (depends on: 6, 7, 8, 9, 10, 12, 14, 15, 16, 18, 19)
Convert the design into concise documents that people can use during an incident. The actual operating process must also be the documented and audited process.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, evidence, and exception standards.
- Include one-page cards for severity, roles, authority, escalation, and communication timings inside the incident tool.
- State explicitly that responders support only owned or formally accepted and trained service portfolios.
- Include the compensation structure, fatigue rules, declaration rights, and non-retaliation commitment.
- Obtain approval from the CTO, HR, Legal, Security, Compliance, and Internal Audit.
- Announce the policy at an all-hands and through team briefings.
- Create an exception register with owner, rationale, compensating control, approver, review date, and expiry.
- Version every policy change. Do not rewrite historical records when the process changes.
24. Instrument the scorecard and review forums (depends on: 3, 11, 18, 19)
Measure the process before the pilot so failures become visible immediately. Report medians and 90th percentiles rather than averages alone.
- Measure impact-to-detection, detection-to-declaration, declaration-to-command, acknowledgement, mitigation, recovery, and resolution.
- Split results by severity, service tier, customer journey, region, detection source, and business-hours status.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track notification episodes, actionability, duplicates, after-hours load, missed detection, routing errors, and page-budget breaches.
- Track postmortem timeliness, action age, due-date performance, verified effectiveness, and repeated contributing factors.
- Track journey availability, error-budget burn, failed or delayed value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight work, recovery days, exceptions, sentiment, and responder attrition.
- Hold a weekly Incident Review Board chaired by the Head of Reliability with relevant directors.
- Hold a monthly executive reliability review and a quarterly control and resilience review with Internal Audit.
- Reconcile incident records monthly against support cases, customer complaints, status history, credits, and major operational anomalies to detect under-reporting.
- Use team-level scorecards to direct help and investment. Never penalize an individual for good-faith declaration.
25. Pilot the complete process on the payment path (depends on: 9, 11, 13, 20, 21, 23, 24)
Run a four-to-six-week pilot across the highest-risk journey before expanding. The interim operating floor remains active for the rest of the company.
- Include payment orchestration, ledger application, PostgreSQL platform, API edge, authentication, settlement, reconciliation, Kubernetes platform, and Support intake.
- Include teams with existing on-call experience and teams new to the model.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking together.
- Run legacy and new paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently taking command.
- Correct critical process or tooling defects within 48 hours.
- Exit only after 95% timely command assignment, 95% communications compliance, no unpaid pages, complete required postmortems, tested fallbacks, and at least 50% lower pilot noise.
- Publish the pilot results, defects, and policy changes company-wide.
26. Roll out by risk with readiness gates (depends on: 25)
Expand in controlled waves and finish early enough to accumulate operating evidence before the audit. A calendar date does not override a failed readiness gate.
- Roll out remaining Tier 0 domains first, followed by Tier 1, Tier 2, and Tier 3.
- Use four waves of six to eight teams, each lasting two to three weeks.
- Gate each service on catalog ownership, appropriate coverage, active compensation, trained responders, access, tested escalation, alert quality, runbooks, and a passed tabletop.
- Require at least six responders only for direct 24x7 technical rotations. Use business-hours coverage for lower tiers.
- Give each wave a named coach and director sign-off.
- Reschedule failed gates or use a time-limited executive exception with compensating controls. Do not create silent waivers.
- Disable legacy human-paging routes after verified cutover for each wave.
- Run a quota-based noise reduction sprint in every wave, starting with the highest-volume rules.
- Pair every suppression with a compensating-detection check.
- Publish an internal adoption dashboard by team, service tier, coverage, and control gap.
- Complete critical coverage by approximately week 12 and all 28 teams by week 18.
27. Prove SOC 2 operating effectiveness (depends on: 22, 23, 24, 26)
Generate evidence through normal operation rather than reconstructing it before fieldwork. Test both control design and consistent execution.
- Map controls to the applicable Trust Services Criteria with Compliance and the auditor, including monitoring, incident identification, response, recovery, communications, and availability.
- Retain approved policies, exceptions, service ownership, schedules, compensation activation, access reviews, training, incidents, communications, reportability decisions, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, missed timelines, non-reportable decisions, and exercises in the testing population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the evidence populations and interviews expected from the external auditor.
- Correct deviations through tracked actions with owners and dates. Never edit history to create apparent compliance.
- Verify that evidence retention covers the full auditor-defined observation period.
- Brief commanders, engineers, Support, and Compliance on the actual process without scripting inaccurate answers.
28. Inspect, adapt, and institutionalize ownership (depends on: 26, 27)
Prevent the process from decaying after rollout or the audit. Change it using measured operating evidence rather than opinion.
- Review the policy after 90 days of live operation using severity calibration, page load, missed detection, communication compliance, action closure, fatigue, and survey results.
- Remove steps that create work without reducing risk. Add controls only where incidents, exercises, or evidence show a gap.
- Reassess Tier 0 and Tier 1 classification and domain boundaries every six months.
- Assign permanent owners for policy, catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review compensation bands, rotation burden, accommodations, and staffing annually.
- Report severe incidents, credits, overdue high-risk actions, and ledger concentration risk to the board or risk committee quarterly.
- Maintain the ledger blast-radius program as an executive risk until failover, degraded mode, reconciliation, and regional independence meet approved objectives.
- Evaluate follow-the-sun coverage using one year of actual activation and staffing data.
- Build year-two plans for automated mitigation, safer deployments, graceful degradation, and error-budget release controls.
--- PROPOSAL 3 ---
Proposal ID: f61aca40-a8a9-4e91-a5f3-ac1d03634441
Content:
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of 'not reportable'.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (29):
1. Charter the program, fund it, and start the audit clock
Convert the CEO email into a company operating process with one accountable owner, a budget, and a dated timeline that starts this week.
- Name the CTO as executive sponsor and appoint a Head of Reliability & Incident Management as the single accountable owner with full-time authority over all 28 teams.
- Stand up a three-person program office: program lead, platform engineer, reliability analyst.
- Form an eight-person steering group spanning Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, and HR. It proposes; the sponsor decides within 48 hours. Never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one human-paging path, one incident record, one postmortem format, paid on-call, mandatory action tracking, and named service ownership.
- Approve budget anchored against the $1.3M in SLA credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs, and 15% reserved engineering capacity.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Publish a one-page charter company-wide on day 2. Incident response is a company process, not a per-team preference.
- Timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day interim command floor (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place this week so the next outage has a named commander and the evidence clock starts immediately.
- Publish a one-page interim severity card and one declaration path: a Slack command, a phone number, and the existing pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs from the 12 teams already on-call. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical action items into complete, re-plan, or formally risk-accept within 30 days, prioritising ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline of incidents, alerts, and money lost (depends on: 1)
Rebuild the facts before locking any design. This is both the design input and the frozen 'before' picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines in one signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the written fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written, signed deal.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- State explicitly that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Map compliance, evidence, and audit requirements from day one (depends on: 1)
Design evidence as a by-product of operations, not a reconstruction before the auditor arrives. The interim process in week 1 is already evidence.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for internal and external communication; CC5 for control activities; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure with verification, training register, drill records, reportability decisions including 'not reportable'.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records. Require role-based access, MFA, and periodic access review.
- Version and approve all policy documents from day one: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard, Alert Quality Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has a named owner. A wrong owner recreates the pager objection. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Assign one accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. **Tier 1**: customer-facing but degradable. **Tier 2**: internal or batch. **Tier 3**: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for every Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform. They fail differently and need different hands.
- Orphan services get an owner within 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Missing ownership or runbooks blocks Tier 0/1 releases.
7. Adopt the severity scale, declaration rights, and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix. Anyone may declare. Nobody is penalised for over-declaring.
- **SEV1 (crisis):** money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; deploy freeze; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads; business-hours comms; postmortem if customer-detected, over two hours, a repeat, or credit-generating.
- **SEV4:** no customer impact. Ticket only. Never pages.
- Payments-specific anchors: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident become at least SEV2.
- Only the Incident Commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
8. Define incident roles, authority, dual-control, and handover discipline (depends on: 7)
Solve 'nobody in charge for an hour' by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Command claimed within 5 minutes and stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Staff 24x7 with a three-layer coverage model, not 28 night rotations (depends on: 6, 8)
Do not create 28 night rotations. That is precisely what engineers are rejecting. Centralise coordination in a trained command corps and keep technical ownership local.
- **Layer A — Incident Command corps:** approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. Paired with a Communications Lead pool of approximately 18 from Support, CS, and engineering management, and a scribe pool used as the training entry point.
- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat a Lisbon or APAC follow-the-sun cell as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
10. Approve paid on-call, New York labor compliance, and fatigue safeguards (depends on: 4, 9)
Unpaid on-call in New York is both a retention problem and a wage-hour exposure. Pay must be live in payroll before any new mandatory night rotation starts. Publish the actual numbers, then ask people to sign up.
- Indicative scheme locked by HR, Finance, and employment counsel within 14 days: approximately $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as approximately 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
- Publish on-call load by team quarterly.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
11. Set the alert-quality standard and a hard page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target: 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, noise below 15%, with no loss of Tier 0/1 detection coverage.
12. Detect payment and ledger failures before customers do (depends on: 6, 11)
The goal is blunt: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
- Run the replay test: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish replay coverage as a leading metric.
13. Consolidate to one paging platform, one incident record, one status page (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
14. Codify the five-minute escalation path and live execution doctrine (depends on: 8, 9, 13)
Write one unskippable path from 'something looks wrong' to 'someone is in charge'. The default action is never waiting. If nobody claims command within 5 minutes, the platform assigns it and announces it.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute.
- Separate diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Closure requires a stability observation window and explicit handback.
15. Set the service-readiness bar and write major-incident playbooks (depends on: 6, 9, 12)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human at night. The shared ledger cluster is the single largest structural risk.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction: partitioning, read replicas, isolation of non-critical readers, with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
16. Run one communications clock for internals, customers, and regulators (depends on: 7, 8, 13)
Replace 'whoever is around' with one timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure. State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is 'not reportable'.
- Obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
17. Tie incidents to SLA credits and true financial cost (depends on: 7, 16)
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
18. Make blameless postmortems mandatory with one format and fixed deadlines (depends on: 7, 8)
Replace 'some incidents, various formats' with one mandatory format, fixed deadlines, and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did?
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and are never the commander of the incident under review.
- Publish a searchable library and a quarterly top-five recurring causes analysis, with restricted versions for security or privileged content.
19. Enforce action ownership, reserved capacity, and tracking (depends on: 13, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity. Closing a ticket without evidence does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
- Target: 90% of high-priority actions closed by due date within two quarters.
20. Train and certify every role before independent duty (depends on: 8, 14, 16, 18)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it. The certification register is an audit artefact.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days plus shadowing): command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
21. Rehearse with tabletops, game days, and unannounced drills (depends on: 13, 15, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-mitigation.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
22. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 14, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 10, 12, 13, 15, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse.
- Publish a one-page result to the whole company.
24. Instrument metrics, dashboards, review cadence, and anti-gaming (depends on: 13, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
25. Roll out in risk-ordered waves with readiness gates (depends on: 23, 24)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, never waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave. Give every team its ranked list of noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Review missed detections monthly with the same seriousness as noise.
- Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
26. Run change management, fairness, and pager culture from day one (depends on: 4, 10)
Run this in parallel from day one. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
27. Produce SOC 2 evidence by operating, test internally, and mock-audit (depends on: 22, 24, 25)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence. Documented exceptions beat a claim of perfection.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5, CC2.2/CC2.3, CC5, and A1.2. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including 'not reportable', postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in months 4 and 6, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
28. Maintain the program risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and pre-commit the response. Review monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
29. Inspect at 90 days, lock year-two ownership, and prevent decay (depends on: 24, 25, 27)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 4 ---
Proposal ID: 91ea9817-3664-46dd-a758-888405cea636
Content:
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- By day 7, every suspected major incident uses one record, one channel, and a named commander; Support can declare without engineering confirmation.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of customer-visible SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (26):
1. Charter the program, fund it, and start the audit clock
Turn the CEO email into a chartered company program within 48 hours. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability as full-time owner, with a three-person office: program lead, platform engineer, and analyst.
- Form a small decision group of Engineering, SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit. It proposes. The sponsor decides within 48 hours.
- Lock the non-negotiables now: one severity scale, one human-paging platform, one incident record, one postmortem format, mandatory action tracking, paid on-call, and named service ownership.
- Confirm the SOC 2 Type II observation window with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, exercises, and reserved engineering capacity against the $1.3M in SLA credits.
- Reserve 15% of engineering capacity for detection, runbooks, and incident actions, protected by the sponsor.
- Publish a one-page charter on day 2. Clock: floor by day 7, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for policy, tooling, or the audit. Put a crude but real process in place this week so the next outage already has an owner.
- Publish a one-page interim severity card and one declaration path: chat command, phone number, and current pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline, and one naming convention for every suspected major incident.
- Triage the 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
- Replay one of the two nobody-in-charge incidents as a tabletop within 14 days using this floor process.
3. Rebuild the forensic baseline and incident-replay catalog (depends on: 1)
Rebuild the facts before locking design. This is the design input, the frozen before-picture for the CEO, and the test set for detection work.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact missing signal. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Name the top 50 noisy rules and every rule with no owner or runbook.
- Quantify true cost: credits, failed payment volume, delayed value, reconciliation breaks, and engineering hours lost.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11 of 64 actions closed, on-call in 12 of 28 teams.
- Build a **replay catalog**: for each historical incident, the detector that should now fire, at which minute, and the owner of the gap.
4. Publish the on-call fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders. Do not punish them by making them the company's permanent night watch.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the **fairness contract** in writing: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive exemption path for health, disability, or caregiving.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and journey tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has one owner. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Record owning team, manager, chat channel, escalation policy, dashboards, runbook, dependencies, regions, and data stores for all 180 services.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, refund, report, onboard — to services, data stores, regions, sponsor banks, and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform. They fail differently.
- Orphan services get an owner in 30 days or an approved decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Coverage follows the journey, not the org chart. Small teams that own Tier 0 pieces get headcount, service reassignment, or membership in a domain rotation. Never a two-person 24x7 rota.
6. Lock severity levels, integrity flags, and the incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Classify on actual or credible customer, financial, security, and contractual harm, never on who reported it or how hard the fix looks.
- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 contained: narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages a human.
- Attach an **integrity, security, settlement, or regulatory flag** to any severity. The flag forces dual-control, Legal, and the reportability checkpoint without inventing a fifth level. A one-customer ledger corruption is still a flagged crisis.
- Auto-escalate to at least SEV2: any ledger-cluster event, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident.
- Anyone may declare. Nobody is punished for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- Communications Lead: single voice for the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists. A human validates at SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory text. The commander owns the facts.
- Distinct people for command, comms, and technical lead at SEV1. Comms and scribe may combine only for bounded SEV2.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
- Every role assignment and handover is announced verbally and in writing with the exact time and open risks.
8. Staff 24x7 with command, domains, and a Duty Triage Desk (depends on: 4, 5, 7)
Do not create 28 night rotations. That is what engineers are rejecting. Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified people from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to eight months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate ownership into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- **Duty Triage Desk**: a small paid overnight first-line rotation that owns the first ten minutes of ambiguous, unowned, or low-confidence pages. It verifies, enriches, and applies only the runbook's safe steps, then wakes the owning domain. It never sits on a suspected ledger, payment-halt, or security event. Those page commander and likely domains immediately, in parallel.
- Overnight comms: SEV1 pages a Communications Lead 24x7. Customer-visible SEV2 lets the commander publish the first status from a template; Comms is paged if the incident is still open at 30 minutes or a top-100 account is affected.
- Staffing math: 260 engineers can sustain 10–12 domain rotations, one command corps, and one triage desk. They cannot sustain 28 night rotas. Role exclusivity: nobody is primary on two rotations in the same week. Commanders may also be domain responders in different weeks.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
- Evaluate follow-the-sun as a 12-month option, not a year-one dependency. Seed Layer B from the 12 teams already on-call.
9. Pay on-call, meet New York labor rules, and cap fatigue (depends on: 4, 8)
Unpaid on-call in New York is a retention problem and a wage-hour exposure. Pay must be in payroll before any new mandatory night rotation starts.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, $400–$800 for Comms or triage-desk duty, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- **Hard gate: no mandatory night rotation starts before compensation is live in payroll.** Interim duty is paid retroactively. Budget roughly $0.9M–$1.2M a year, then refine with actual rotation count.
10. Enforce an alert-quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers, then replay the past (depends on: 5, 10)
Stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. A detector is not done until it would have caught the last 12 months.
- Define SLIs and SLOs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner, processor, and sponsor-bank notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a named defect class with a mandatory tracked action.
- Run the **replay test** against the S3 catalog. For each of the 31 incidents, name the detector that would now fire and at which minute. Close gaps the replay exposes before calling detection improved.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One chat command creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy. Unowned pages go to the Duty Triage Desk and Duty Command, and log a catalog defect.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Codify escalation, the five-minute command rule, and vendor incidents (depends on: 7, 8, 12)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting. Payments also fail at processors and banks, which you cannot patch.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Vendor-incident class: processor, sponsor-bank, card-network, or cloud-control-plane failure. Command is still required. Success is time-to-customer-notice, time-to-failover-decision, and queue management, not root cause at the vendor.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Write live execution doctrine, readiness bars, and ledger playbooks (depends on: 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving root cause. A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis. Split diagnosis and mitigation workstreams once staffing allows.
- Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Drain backlogs under control. Complete reconciliation before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date. Existing detection stays on while gaps are repaired.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers, with board-visible milestones and executive-signed RPO/RTO.
15. Run one communications clock for internals, customers, and account managers (depends on: 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for customer-visible SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys, not internal service names. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer record.
- Language rules: state impact and the next update time. **Never speculate** on cause, recovery time, data integrity, or blame.
16. Operationalize regulatory notice, partner clocks, and SLA credits (depends on: 6, 15)
In payments, some incidents start a legal clock at detection. Tie incidents to money so Finance does not learn about outages from invoices.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when not reportable**, with facts, decision-maker, and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families. Track failed payment volume, delayed value, and reconciliation breaks as the true incident cost.
17. Make postmortems mandatory and actions enforceable (depends on: 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth. Eleven of 64 closed is a process nobody enforces.
- Mandatory for every SEV1 and SEV2, any incident a customer detected first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and gap, timeline, response analysis, contributing conditions, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items. Facilitators are trained and are never the commander of the incident under review.
- Every action gets one named owner, priority, due date, expected risk reduction, verification method, and a ticket created automatically. If it is not on the board, it does not exist.
- Classes: containment in 7 days, corrective in 30, strategic in 90. SEV1 recurrence-prevention enters the next sprint ahead of roadmap work. Reserve 15% capacity.
- Overdue ladder: manager at +7 days, director at +14, CTO at +30. Overdue high-risk items need written residual-risk acceptance and can block related releases. Verify effectiveness before closing.
18. Train and certify every response role before independent duty (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Duty Triage Desk, half day: enrichment, safe-step limits, when to wake a domain immediately, when not to delay.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
19. Publish Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17)
Collapse the design into a document people will actually open mid-outage, and make it official. Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
20. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 18, 19)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, Duty Triage Desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
21. Instrument metrics, reviews, and anti-gaming (depends on: 12, 17, 20)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- Response: time from first impact to detect, declare, commander, acknowledge, mitigate, and resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
22. Rehearse with tabletops, game days, and night drills (depends on: 12, 14, 18, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced overnight paging drill to measure real acknowledgement times.
- One combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO, before the audit.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
23. Burn down alert noise and roll out by risk with gates (depends on: 20, 21)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Noise reduction is a quota inside each wave, not a background hope.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy paging paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Give every team its noisiest rules from the baseline. Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed. Pair every suppression with a compensating-detection check.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
- If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
24. Operate the fairness and culture program in parallel (depends on: 4, 9)
Run this from the moment the deal is published. Engineers judge the process on fairness. Executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
25. Prove SOC 2 operating effectiveness before fieldwork (depends on: 19, 21, 22, 23)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (depends on: 21, 23, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, metrics, and the ledger blast-radius workstream, independent of the audit cycle.
- Pre-committed contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 5 ---
Proposal ID: b427ef6b-42c2-41ac-8f9d-992d99ba5e20
Content:
Estimated Complexity: high
Success Metrics: - Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (30):
1. Charter the program and start the audit clock
Turn the CEO email into a chartered company program with one accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
- Confirm the SOC 2 Type II observation window with the auditor; map the future response controls to applicable Trust Services Criteria; define evidence set and retention requirements.
4. Listening tour, resistance map and the written on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (depends on: 5, 7)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 4, 8)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (depends on: 5, 10)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (depends on: 7, 13)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 12, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
- Agree with Legal and Finance the availability measurement method per contract and per component; compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide credit posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain; attribute credits to root-cause families for investment decisions.
18. Blameless postmortem standard and Incident Review Board (depends on: 6, 7, 12)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
19. Action ownership, reserved capacity and enforcement (depends on: 12, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
20. Training and certification academy (depends on: 7, 13, 15, 18)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
21. Exercise programme: tabletops, game days and unannounced drills (depends on: 12, 14, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
22. Publish Incident Management Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
24. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
25. Alert noise burn-down campaign (depends on: 10, 12, 23)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
26. Wave rollout to all 28 teams with readiness gates (depends on: 23, 24, 25)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
27. Change management, fairness and pager culture (depends on: 4, 9, 23)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
28. SOC 2 evidence by design, internal testing and mock audit (depends on: 22, 24, 26)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
29. Risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
30. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 24, 26, 28)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Instructions:
1. Carefully evaluate each proposal considering:
- Clarity and detail of the proposed steps.
- Overall alignment with the Main Objective.
- Quality and relevance of the proposed success metrics.
- Coherence of dependencies between steps (if any).
2. Think deeply about each proposal, analyzing its strengths and weaknesses. Consider how well each proposal addresses the main objective and which one would be most effective.
After your thorough analysis, give the Proposal ID of the selected proposal and a one- or two-sentence justification of your choice.
{"selected_proposal_id": "30f15b26-0ec0-4756-9f4a-6279eb889884", "justification": "Proposal 1 is the most comprehensive plan, with 35 steps that uniquely include a dedicated **change intelligence** step, a **concurrency doctrine** for simultaneous incidents, and an explicit **Support-as-detection-tier** step, while also carrying the most detailed success metrics (including change-correlated incident tracking and top-100 account outreach). Its dependencies are internally consistent and it addresses every element of the main objective with the greatest operational specificity."} |
| grok4.6_voter_4 | grok4.6 · xai/grok-4.6 |
Proposal A · opus5_refine_1 winner | Proposal 1 is the most complete match to the brief: it defines severity, roles, 24x7 staffing, paid on-call, detection, communications, postmortems, and metrics in operational detail, and it puts a live command floor in place on day 7. It uniquely adds a Duty Triage Desk, change intelligence, concurrency rules, and Support as a detection path, which directly attack unclear ownership, customer-first detection, and pager resistance. |
50.6k in · 135 out · 1 min 56 s | show[SYSTEM]
You are an expert and objective evaluator of project plan proposals.
Your task is to select the BEST proposal based on criteria of completeness, clarity, and alignment with the main objective.
Use your internal reasoning processes to thoroughly analyze each proposal, considering all aspects and implications.
Take as much time and space as you need to evaluate each proposal in depth before making your decision.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
Proposals to Evaluate:
--- PROPOSAL 1 ---
Proposal ID: 30f15b26-0ec0-4756-9f4a-6279eb889884
Content:
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90 a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95% internally and externally.
- Top-100 account outreach completed within 30 minutes of SEV1 in 95% of qualifying cases, using the approved briefing pack.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- Change-correlated incidents are identified within 10 minutes of declaration in 90% of cases, and the change-correlated share of incidents declines quarter over quarter.
- SLA credits fall to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, one concurrency drill and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and a tested read-only mode by month 6, with milestones reported to the board quarterly.
Steps (35):
1. Charter the programme: one owner, one mandate, funded, dated against the audit
Turn the CEO email into a chartered company programme with a single accountable owner and authority over all 28 teams. Incident response stops being a per-team preference and becomes a **company operating process**.
- Name the CTO as executive sponsor and a full-time Head of Reliability & Incident Management as accountable owner, supported by a programme office of three: programme lead, incident-platform engineer, reliability analyst.
- Form a decision group (Engineering, Platform/SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables on day one: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, named service ownership, and **paid on-call**.
- Publish the timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal control test month 5, mock audit month 7, SOC 2 fieldwork month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.3M/yr, training and exercises, plus 15% of engineering capacity reserved for reliability work.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence immediately.
- Publish a one-page interim severity card and a single declaration path: one chat command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim 24x7 Duty Incident Commander roster (primary plus backup) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document and one naming convention for every major incident, starting now.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is the design input, the executive narrative and the frozen "before" picture for the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the **detection backlog**.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Correlate incidents with deployments, config changes and feature flags to quantify how many started with a change we made.
- Quantify true cost beyond credits: failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost, churn risk on named accounts.
- Freeze the baselines in a signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Audit, legal and evidence scoping in week one (depends on: 1)
Design evidence as a by-product of operating, never as a reconstruction before fieldwork. Settle the scope and the legal handling of incident records now, not in month seven.
- Confirm with the external auditor the Type II observation window, the incident population definition, and the evidence they will sample. Everything from the day-7 floor onwards must count.
- Map incident response to the Trust Services Criteria with Compliance: CC7.2–CC7.5 (monitoring, identification, response, recovery), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), A1.2 (availability).
- Define the evidence set and where it is produced automatically: incident records, paging and acknowledgement logs, role assignments, status-page history, reportability decisions, postmortems, action closure with verification, training register, drill records, access reviews.
- Agree retention, confidentiality, legal hold and access rules. Decide with counsel which postmortem content is privileged and how privileged material is segregated **without making the ordinary postmortem secret**.
- Start the obligation matrix with Legal: NYDFS 23 NYCRR 500, state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network contractual windows, cyber-insurer notice.
- Begin monthly evidence sampling from month 1 so control operation is visible long before the audit.
5. Listening tour, resistance map and the written on-call deal (depends on: 1, 3)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written deal signed by the sponsor.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
- State the hard gate publicly: no new mandatory night rotation starts before compensation, training, runbooks and staffing rules are live.
6. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- Record for every service: one accountable team, engineering manager, chat channel, escalation policy, dashboards, runbook link, dependency list, regions and data stores.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record SLO, RTO, RPO, active-active or region-bound, failover method, and the dependencies that make nominal two-region redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
7. Severity standard, declaration rights and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning responders and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and responders paged, comms and scribe if customer-visible, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, credit-generating or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Attach a financial-integrity flag or a security flag to any severity. The flag forces dual control, Legal engagement and the reportability checkpoint without inventing a fifth level.
- Anchor on payments reality alongside error rates: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including a customer report.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
8. Roles, authority, concurrency and handover discipline (depends on: 7)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice to the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1 and SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, Finance and Vendor Management join on defined triggers rather than by invitation.
- Rules: command claimed within 5 minutes and stated in channel; distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- **Concurrency doctrine:** two simultaneous SEV1/SEV2 incidents activate the secondary commander and a surge roster; a designated Multi-Incident Coordinator arbitrates shared resources such as the ledger, the database platform and the deploy freeze.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
9. Three-layer 24x7 coverage: command corps, domain rotations, overnight triage desk (depends on: 5, 6, 8)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving roughly one primary week per person every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement and reconciliation, auth, API edge, Kubernetes platform, data and reporting, partner integrations). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Publish the staffing arithmetic: 10–12 domain rotations of six to eight people plus a 30-person command corps means roughly 100 of 260 engineers carry a night obligation, at about one week in six. Twenty-eight independent rotations would be unstaffable and is therefore rejected on the numbers.
- Teams too small for a fair rotation get headcount, service reassignment, or a time-limited executive exception. **Never a two-person 24x7 rotation.**
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: a human acknowledgement within 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
10. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 5, 9)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour, and a premium for the overnight triage desk.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per responder per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty since day 7 is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons without career penalty, and on-call load is published quarterly by team.
11. Alert quality contract and page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression can never hide an incident.
12. Detection uplift on the money path, validated by incident replay (depends on: 6, 11)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action and a missed-detection review.
13. Change intelligence and deployment safety on the critical path (depends on: 6, 12)
Most of these incidents start with something we changed. Make change the first hypothesis the tooling answers, and make changes safer to reverse.
- Stream every deployment, configuration change, feature-flag flip, schema migration and infrastructure change into the incident timeline with service labels and owners.
- Give the commander an automatic "what changed in the last 60 minutes on the affected journey" panel at declaration time.
- Require Tier 0/1 changes to be progressively delivered with a documented rollback that is tested, time-bounded and executable by the on-call responder without the author.
- Treat ledger schema migrations and settlement-affecting changes as a separate class: dual approval, rehearsed rollback, no deploys inside the settlement window.
- Enforce a change freeze during SEV1 and SEV2, lifted only by the commander and logged.
- Report change-correlated incidents monthly; a rising ratio is a signal to strengthen release safety, not to blame a team.
14. One pager, one incident record, one status page — migrated without a detection gap (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy paging path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access, MFA and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page delivered outside this platform creates no on-call obligation.
15. Escalation ladder and the five-minute command rule (depends on: 8, 9, 14)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts and support entitlements: AWS and database premium support, processors, sponsor banks, card networks.
16. Live execution doctrine and payments safety rules (depends on: 8, 14, 15)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed in advance rather than improvised.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and automatic reopening if impact recurs.
17. Service readiness bar, major-incident playbooks and ledger blast-radius reduction (depends on: 6, 9, 16)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human, and the shared ledger is the largest structural risk in the estate.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as a concentration risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers, stronger regional independence — with board-visible quarterly milestones.
- Enforcement: no readiness sign-off, no night paging alerts unless the manager accepts the gap in writing with an expiry date and a compensating control. Never respond to a gap by turning detection off.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
18. Internal communications protocol (depends on: 8, 14)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, Sales and Security.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change. First internal brief within 10 minutes of SEV1 and 15 of SEV2.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
19. Customer communications, status page and account-manager outreach (depends on: 7, 18)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; monitoring notice after mitigation; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, monitoring, resolution.
- Component-level status page mapped to customer journeys rather than internal service names, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1 and 60 of SEV2, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state observed impact, affected capabilities, any workaround and the next update time; **never speculate on cause, recovery time, data integrity or blame**.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
- Record any legally required restriction or delay of public detail, its approver, and the alternative stakeholder plan.
20. Support and account management as a detection and intake tier (depends on: 7, 12, 19)
Customers detected 40% of incidents first, which means the front line already holds the signal. Turn Support and account managers into an instrumented detection channel rather than a bystander.
- Give Support explicit declaration rights, a one-page trigger card, and a macro that opens an incident candidate directly in the incident platform.
- Automate the clustering rule: five similar tickets or calls in ten minutes auto-creates a triage incident assigned to the duty commander.
- Route credible partner, processor and sponsor-bank notifications into the same declaration path within five minutes.
- Generate the affected-customer list automatically from journey telemetry and the incident record, and push it to Support, CS and the account-manager briefing.
- Train Support and account managers on approved language and prohibit independent technical explanations to customers.
- Measure and publish "signal was in Support before it was in monitoring" as a detection defect, and feed each instance into the detection backlog.
21. Regulatory, partner and legal notification playbook (depends on: 4, 7, 19)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Complete the obligation matrix started in scoping and have counsel validate the triggers, deadlines, channels and submitting authority for each obligation.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with facts considered, decision-maker, timestamp and reassessment trigger.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, card networks, outside counsel, cyber-insurer, critical vendors.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the operational facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly inside the exercise programme, including one full regulator-notification simulation per year.
22. SLA credit workflow, true cost model and incentive guardrails (depends on: 7, 19)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component, including how partial degradation counts.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Guardrail against the perverse incentive: better detection will surface incidents that previously went unbilled, so credits may rise before they fall. Publish this expectation to the executive team in advance, and make it a written rule that no finance or commercial pressure may influence severity or declaration.
23. Blameless postmortem standard and Incident Review Board (depends on: 4, 7, 8)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger; any control failure.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning engineering director is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions where security, privacy or privileged content requires it.
24. Action ownership, reserved capacity and enforcement (depends on: 14, 23)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, accountable manager, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness through tests, telemetry, exercises or production evidence before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
25. Training and certification academy (depends on: 8, 15, 18, 23)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, the 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, contractual clocks, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days of joining a domain.
- Certification is valid 12 months and renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
26. Exercise programme: tabletops, game days, night drills and vendor rehearsals (depends on: 14, 17, 25)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander. First company-wide tabletop within 30 days.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times, plus a concurrency drill with two simultaneous incidents.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down. Rehearse joint escalation with AWS, a processor and a sponsor bank.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
27. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 15, 18, 19, 21, 23, 24)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, escalation and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix, Readiness Bar.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
28. Pilot on the payments critical path (depends on: 10, 12, 13, 14, 17, 25, 27)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement and reconciliation, PostgreSQL platform, Kubernetes platform, API edge, auth and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, action tracking, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The programme lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time in 95% of cases, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
29. Alert noise burn-down campaign (depends on: 11, 14, 28)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → under 1,500 pages a month by day 90, under 500 with noise below 15% by month 6.
30. Metrics, dashboards, review cadence and anti-gaming (depends on: 14, 23, 28)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time from first impact to detect, declare, commander, acknowledge, mitigate, resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors, change-correlated incident ratio.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
31. Wave rollout to all 28 teams with readiness gates (depends on: 28, 30)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the programme office for three weeks; the director signs the gate.
- **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard by team, tier and control gap.
32. Change management, fairness and pager culture (depends on: 5, 10, 28)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
33. SOC 2 evidence by design, internal control test and mock audit (depends on: 4, 27, 30, 31)
The real process is the evidence. Never build a parallel audit process, and never reconstruct records after the fact.
- Maintain the evidence set defined in scoping, produced automatically and indexed: versioned policies and exceptions, catalog records, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Sample incidents monthly from first signal through verified corrective action, deliberately including customer-reported events, downgraded incidents and missed timelines.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling end to end and reporting gaps to the sponsor.
- Run a formal mock audit in month 7 using the populations, evidence requests and interviews the auditor will use: a commander, a random engineer, Support, Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
34. Programme risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps and immediate staffing intervention.
- **A major SEV1 mid-rollout:** the programme lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
35. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 30, 31, 33)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, exercise calendar and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and quarterly board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 2 ---
Proposal ID: 1f1527bb-d992-451c-bbc4-9e61ffdb1f30
Content:
Estimated Complexity: high
Success Metrics: - By day 7, every suspected major incident uses one record, one coordination channel, and a named Incident Commander within 10 minutes.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes.
- By month 3, at least 95% of SEV1 and SEV2 incidents have a named commander within 5 minutes.
- By month 3, at least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and appropriate technical coverage.
- By week 18, all 180 services have an owner, tier, escalation path, and tested coverage model.
- No new mandatory night rotation begins before compensation, training, access, runbooks, and minimum staffing are active.
- Every direct 24x7 technical rotation has at least six qualified responders or an approved, expiring executive exception.
- No responder is routinely primary more often than one week in six or assigned to two simultaneous primary rotations.
- Median impact-to-detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- By day 90, incident replay identifies a current detector for at least 90% of the 31 historical incidents, including its expected detection minute.
- Median time to mitigation falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of applicable SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet the required internal and customer update cadence by month 3.
- Monthly human notification episodes fall from the current 3,400 alert events to no more than 1,500 by day 90 and 500 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without reducing Tier 0 or Tier 1 replay coverage.
- Average after-hours load remains at or below two notification episodes per responder per week; every sustained breach receives a dated remediation plan.
- All six monitoring sources route human pages through the controlled paging platform by week 12, with direct legacy routes retired by week 18.
- All required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 from month 3 onward.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due dates with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Reportability is assessed and recorded within 1 hour for 100% of SEV1 and qualifying SEV2 incidents, including not-reportable decisions.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Approved ledger RTO and RPO, controlled failover, read-only mode, and post-recovery reconciliation are exercised before month 6.
- Quarterly responder surveys reach at least 75% favorable responses for fairness, ownership boundaries, compensation, and sustainability by month 6.
- Monthly contracted availability meets or exceeds 99.95% by month 6 using the contractually authoritative measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- The month-7 mock audit finds no unowned high-risk control gap and at least 95% of sampled incidents contain complete operating evidence.
- The SOC 2 Type II incident-response controls complete external testing without an unresolved material exception.
Steps (28):
1. Charter the program and fund immediate action
Make incident management a **company operating process** within 48 hours. Give one accountable leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Include Engineering, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit in a decision group. The sponsor resolves blocked decisions within 48 hours.
- Approve non-negotiables: one severity model, one human-paging path, one incident record, paid on-call, named service ownership, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, resilience, and incident actions.
- Fund tooling, compensation, training, exercises, and program staff. Compare the cost with the existing $1.3M in annual SLA credits.
- Set the schedule: operating floor by day 7, design during weeks 1–4, pilot during weeks 5–10, rollout during weeks 11–18, control tests in months 4 and 6, and mock audit in month 7.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for the final policy or new tooling. Put a minimum viable process into operation immediately and retain evidence from the first day.
- Publish a one-page interim severity guide, declaration procedure, role card, and communications clock.
- Provide one monitored declaration route through chat, telephone, and the current paging environment.
- Create one channel, bridge, timeline, and incident identifier for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced managers and engineers.
- Require a named commander within 10 minutes. The duty engineering director assumes command if the command page is unclaimed.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Stop uncompensated mandatory after-hours expansion. Pay interim duty under a temporary stipend, retroactive to program launch.
- Hold a 15-minute daily operational review until the permanent process is active.
3. Create the factual, contractual, and control baseline (depends on: 1)
Build one defensible baseline for process design, investment decisions, and SOC 2 testing. Preserve the original data so progress cannot be created by changing definitions.
- Reconstruct all 31 incidents from impact start through detection, declaration, command, mitigation, recovery, communications, credits, and corrective actions.
- Reconstruct the two incidents with unclear command minute by minute.
- Identify the missing signal for every customer-first detection.
- Inventory all six alert sources, 3,400 monthly alert events, duplicates, noisy rules, missing owners, and missing runbooks.
- Record current rotations, unpaid duty, overnight activations, schedule size, and uncovered services.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, 85% noise, $1.3M in credits, and 11 of 64 actions closed.
- Inventory customer-specific availability definitions, notice periods, credit terms, sponsor-bank obligations, and other incident-related contracts.
- Confirm the required SOC 2 Type II observation period and evidence expectations with the auditor during week 1.
4. Publish the on-call fairness contract (depends on: 1)
Treat pager resistance as a legitimate design constraint. Adoption depends on a written agreement that separates command from technical ownership.
- Interview representatives from all 28 teams and from Support, Customer Success, Security, and Operations.
- Separate concerns about unpaid work, sleep loss, unfamiliar systems, noisy alerts, inadequate runbooks, and blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for systems they own or have formally accepted and trained to support, and assisted by a separate Incident Commander.
- State that Platform may temporarily triage unknown ownership but does not inherit another team's service.
- Provide confidential accommodations for health, disability, pregnancy, or caregiving constraints without career penalty.
- Measure baseline trust, fairness, fatigue, and alert confidence. Repeat at days 60 and 120, then quarterly.
5. Build the service catalog and customer-journey map (depends on: 3)
Make a machine-readable catalog the source of truth for routing, impact analysis, status-page components, and control evidence. Every production service must have one accountable owner.
- Record the owning team, manager, business capability, repository, channel, dashboard, runbook, escalation policy, dependencies, regions, and data stores for all 180 services.
- Map initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, reporting, and onboarding to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing systems that can create material contractual impact.
- Classify Tier 2 as deferrable internal or batch systems and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional recovery mode, contractual commitments, and critical vendors for Tier 0 and Tier 1.
- Keep separate but coordinated ownership for the ledger application and PostgreSQL platform.
- Give orphan services an owner or approved decommission date within 30 days. Treat unowned Tier 0 services as release-blocking executive risks.
6. Adopt the severity model and incident modifiers (depends on: 3, 5)
Classify incidents by credible customer, financial, security, regulatory, and contractual harm. Start at the higher plausible severity while scope or integrity remains unknown.
- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; broad loss of a core payment journey; both regions impaired; or processing must be stopped. Immediately page all response roles and the Executive Duty Officer. Open the bridge within 5 minutes, freeze unrelated changes, issue internal notice within 10 minutes, publish applicable customer status within 15 minutes, start legal assessment within 1 hour, and require a postmortem.
- **SEV2 — major:** material payment degradation; settlement deadline at risk; regional impairment with reduced resilience; a critical customer or material cohort unavailable; or an SLA breach is likely. Page command and technical roles immediately. Issue internal notice within 15 minutes, applicable customer status within 30 minutes, and require a postmortem.
- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads. Page only when immediate action can reduce harm.
- **SEV4 — operational event:** no current customer impact and no credible imminent harm. Create a ticket and handle during normal hours.
- Add FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, and VENDOR modifiers. These invoke specialist controls without distorting customer-impact severity.
- Use at least SEV2 posture for unknown impact lasting 15 minutes, credible ledger-integrity risk, or a cross-domain incident without clear ownership.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Permit anyone to declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Publish a decision tree and examples based on the 31 historical incidents.
7. Define roles, authority, and handoffs (depends on: 6)
Separate coordination, communications, recordkeeping, and technical repair. One named person holds command continuously throughout every SEV1 and SEV2.
- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not act as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, executive updates, and coordination with Legal.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- **Subject-Matter Responders:** diagnose and mitigate only systems for which they have ownership, access, training, or a formally accepted support agreement.
- **Executive Duty Officer:** removes organizational obstacles and makes exceptional business decisions without displacing the commander.
- Add Security, Legal, Compliance, Finance, and Vendor Management on modifier-specific triggers.
- Require distinct commander, Communications Lead, scribe, and technical lead for SEV1. Communications and scribe may combine for the first 10 minutes of a bounded SEV2.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, reconciliation, and change evidence for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, including the exact transfer time and unresolved risks.
8. Create sustainable 24x7 coverage (depends on: 4, 5, 7)
Use central command coverage and risk-based technical coverage instead of creating 28 fragile night rotations. Responders do not carry pagers for unfamiliar code.
- Create a 24x7 Incident Command corps of approximately 30 certified people with primary and backup schedules. Two schedules create about 104 weekly assignments a year, or roughly three to four weeks per person annually.
- Create a 24x7 Communications pool of 18–24 trained people from Support, Customer Operations, Engineering Operations, and management.
- Create a similarly sized scribe pool. The backup commander temporarily records the first minutes if a scribe has not joined.
- Maintain a 24x7 Executive Duty Officer schedule and specialist contact paths for Security and Legal.
- Group Tier 0 and Tier 1 systems into roughly 8–12 coherent response domains only where responders share training, access, runbooks, and explicit support acceptance.
- Staff every critical domain with primary and secondary responders and at least six qualified people. Target eight where overnight activation is frequent.
- Give Tier 2 and Tier 3 services business-hours ownership plus tested manager and director escalation.
- Reclassify any lower-tier service capable of causing severe overnight harm rather than hiding the risk behind manager callback.
- Route unknown-owner incidents to Duty Command and Platform temporarily. Record each occurrence as a catalog control defect.
- Prohibit simultaneous primary assignments and two-person 24x7 rotations.
9. Implement compensation and fatigue protections (depends on: 4, 8)
End unpaid on-call before expanding mandatory coverage. Use fixed duty compensation so responders are not rewarded for alert volume.
- Use planning bands of $900–$1,200 per Tier 0/1 primary week and $300–$500 per secondary week.
- Use planning bands of $1,000–$1,300 per Duty Commander week and $400–$700 for Communications or scribe primary duty.
- Pay holiday premiums. Compensate all legally compensable active and waiting time for non-exempt staff, including overtime where required.
- Have HR, Finance, Payroll, and employment counsel approve final bands, tax handling, FLSA classification, New York wage-hour treatment, and schedule constraints within 14 days.
- Provide a protected recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal delivery commitments by about 15% during a primary week.
- Prohibit consecutive primary weeks, on-call during leave, hidden schedule swaps, and primary duty more often than one week in six.
- Allow responders to declare temporary fatigue-related unfitness without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Make active payroll setup, training, access, and readiness hard gates before any new mandatory night rotation starts.
10. Codify the live incident lifecycle (depends on: 6, 7)
Give every incident the same operational flow from first signal through verified recovery. The first objective is limiting customer and financial harm, not proving root cause.
- Use the states Detected, Declared, Triaged, Mitigating, Mitigated, Monitoring, Resolved, and Reviewed.
- Record impact start separately from detection. Use the earliest defensible evidence and revise it transparently when later facts emerge.
- Open with a standard command message: severity, known impact, assigned roles, immediate objective, workstreams, and next update time.
- Freeze unrelated production changes during SEV1 and normally during SEV2. Record every exception.
- Prefer reversible mitigation: rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, partner rerouting, or controlled processing suspension.
- Separate mitigation and diagnosis workstreams when staffing allows.
- Keep decisions in the shared incident record rather than direct messages.
- Require formal command and technical handoffs for shift changes, fatigue, or incidents exceeding four hours.
- For payment incidents, verify backlog handling, duplicate protection, customer state, settlement exposure, and ledger reconciliation before resolution.
- Require a severity-specific stability period and explicit handback to the owning team, Support, and Customer Success.
11. Establish one paging and incident system of record (depends on: 5, 7, 8)
Monitoring tools may remain specialized, but every human page and major-incident record must enter one controlled platform. This provides consistent routing and an audit trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page within two weeks.
- Ingest events from all six monitoring tools before disabling their direct human-paging routes.
- Route pages using the service catalog and deduplicate events belonging to the same symptom.
- Provide one declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, decisions, severity changes, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status publishing, and conference facilities.
- Apply MFA, role-based access, periodic access reviews, and tamper-evident history.
- Keep privileged legal or security material in restricted linked records rather than exposing it in the general timeline.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, the status provider, or an AWS region.
- Retire each legacy paging route only after ownership review, end-to-end testing, and two weeks of verified operation.
12. Enforce the alert-quality contract (depends on: 3, 5)
Treat every human page as a production interface with an owner and a required action. Measure human notification episodes rather than raw monitoring events.
- Require every paging rule to identify the service, owner, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Define an actionable page as one that causes or materially informs a timely intervention or risk decision.
- Define noise as duplicate, non-urgent, unactionable, stale, test-generated, or incorrectly routed notification.
- Page on payment outcomes, error-budget burn, queue age against deadlines, and financial-integrity risk rather than raw CPU, memory, pod, or log thresholds.
- Run new rules in shadow mode for seven days unless a documented emergency exception applies.
- Review repeated no-action pages within two business days.
- Set a page budget of no more than two after-hours notification episodes per responder per week, measured over four weeks.
- Make a sustained budget breach trigger a tuning sprint and block additional non-emergency paging rules.
- Require compensating detection and central approval before suppressing a Tier 0 or Tier 1 rule.
- Never disable an existing critical detector solely because its metadata or runbook is incomplete. Track the gap with a dated remediation owner.
13. Detect payment failures before customers (depends on: 5, 12)
Move detection from infrastructure health to customer journeys and ledger truth. Validate coverage against actual historical failures.
- Define SLIs and internal SLOs for initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, APIs, webhooks, and reporting freshness.
- Set internal objectives with enough headroom to protect the contractual 99.95% availability commitment.
- Run external synthetic transactions through critical journeys at least once per minute from paths independent of the production platform.
- Test each region and expose dependencies that defeat nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed payment value relative to settlement deadlines.
- Add tenant and cohort anomaly detection for high-value customers and critical payment methods.
- Convert credible Support, account-manager, processor, bank, and network reports into incident candidates within five minutes.
- Replay all 31 historical incidents. Record which current detector would fire and at what minute.
- Treat customer-first detection as a mandatory missed-detection review with a tracked action.
14. Implement the detection and escalation ladder (depends on: 7, 8, 11)
Create one time-bound path from the first credible signal to named command and the correct technical owner. Device delivery does not count as acknowledgement.
- Converge automated alerts, engineer observations, support cases, account-manager reports, partner notices, and customer calls on the same declaration path.
- Page the Duty Incident Commander and owning critical-domain primary immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at 5 minutes, manager at 10 minutes, and director at 15 minutes.
- Escalate unclaimed command to the backup commander at 5 minutes. The Executive Duty Officer assumes temporary command at 10 minutes until a certified transfer occurs.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Keep command with the current commander when service ownership remains unclear. Assign a temporary technical lead and record the ownership gap.
- Maintain tested external escalation routes for AWS, database support, processors, sponsor banks, networks, and critical vendors.
- Test the full declaration, acknowledgement, fallback, conference, and status-publishing path weekly.
- Treat missed acknowledgement, failed routing, and unowned incidents as control failures requiring review.
15. Standardize internal communications (depends on: 7, 10, 11)
Give responders one working room and stakeholders one controlled source of truth. Executives must not interrupt the technical command path.
- Maintain one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue the initial internal notice within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within the same initial-notice window.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record every material decision and outbound message in the incident timeline.
- Require the executive team to sign the communication behavior rules.
16. Standardize customer, account, and regulatory communications (depends on: 3, 6, 15)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate observed impact before root cause is known.
- Publish customer status within 15 minutes of a customer-visible SEV1 and within 30 minutes of a customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map public status components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, available workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Have account managers contact affected strategic accounts within 30 minutes for SEV1 and 60 minutes for SEV2 using the approved briefing.
- Offer status subscriptions to all customers. Auto-enroll only where contracts, consent, and applicable communication rules permit.
- Encode customer-specific notice deadlines and channels in the customer record.
- Have Legal and Compliance maintain a counsel-validated matrix covering applicable NYDFS, breach, GLBA or FTC, PCI, money-transmitter, sponsor-bank, network, insurance, and customer obligations.
- Complete and record a reportability assessment within one hour for every SEV1 and every security, privacy, or integrity-related SEV2, including decisions of not reportable.
- Let Legal own regulatory text and submission while the Incident Commander owns operational facts. Record any legally required restriction of public detail and its alternative stakeholder plan.
17. Tie incidents to SLA credits and financial exposure (depends on: 3, 16)
Link the incident record to contractual and financial outcomes. Finance should not discover outages through later credit claims.
- Define the authoritative availability calculation for each contract and customer journey with Legal and Finance.
- Calculate affected customers and minutes from incident scope and journey telemetry.
- Produce a preliminary credit and contractual-exposure estimate within five business days of resolution.
- Record failed payment count, failed or delayed value, settlement exposure, reconciliation breaks, support effort, and engineering effort.
- Establish a documented approval path for proactive credits and claims-based credits.
- Attribute credits and financial harm to recurring failure families.
- Use the quarterly credit analysis to prioritize detection, resilience, and architectural investment.
18. Establish mandatory blameless postmortems (depends on: 6, 7)
Use one learning standard with fixed deadlines. Keep learning separate from disciplinary and misconduct processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact lasting more than two hours, SLA credits, contractual breach, repeated contributing factors, major control failures, and ledger-integrity near misses.
- Produce a factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the owning engineering director accountable for completion. The commander owns response analysis, and the scribe supplies the timeline.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit answers to why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the commander or primary technical responder.
- Describe decisions using the context and information available at the time. Do not name an individual as the root cause.
- Keep HR, misconduct, and personnel matters in separate processes.
- Publish broadly useful findings internally while restricting security, privacy, personnel, and privileged content appropriately.
19. Make corrective actions enforceable risk commitments (depends on: 18)
An action is not complete when its ticket is closed. It is complete when the intended risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within 7 days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Prefer actions that remove hazards or reduce blast radius over vague actions such as retraining or adding monitoring.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of engineering capacity for approved reliability work.
- Escalate overdue high-risk actions to the manager after 7 days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit unrepaired severe conditions to block related releases.
- Verify effectiveness using tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days as complete, re-planned, superseded with evidence, or formally risk-accepted.
20. Set the service-readiness bar and payments playbooks (depends on: 5, 10, 12, 13)
A critical service must be supportable at 3 a.m. before it enters direct overnight coverage. Existing critical detection remains active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and integrity constraints.
- Require responders to demonstrate access and safe execution before independent primary duty.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional impairment, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document split-brain prevention, replay protection, failover, controlled backlog recovery, and post-recovery reconciliation.
- Require executive-approved RTO and RPO for the shared ledger cluster.
- Exercise critical runbooks at least twice a year and after material changes.
- Block new Tier 0 or Tier 1 releases and paging rules when readiness requirements are missing.
- Handle existing gaps using named owners, compensating controls, executive-approved expiry dates, and remediation plans.
- Run a parallel architecture workstream to reduce ledger blast radius, isolate non-critical readers, and strengthen regional independence.
21. Train and certify every response role (depends on: 7, 14, 15, 18, 20)
Command and communication are learned skills. Use paid working time and certify people before independent duty.
- Give all employees a 30-minute module on recognizing impact, declaring incidents, and locating the status page.
- Give responders a half-day module on severity, acknowledgement, escalation, evidence, runbooks, and financial-integrity precautions.
- Train scribes for two hours on timeline quality and fact-versus-hypothesis labeling.
- Train Incident Commanders for two days on delegation, uncertainty, severity, mitigation strategy, fatigue, handoffs, and executive management.
- Require commander candidates to complete a simulated SEV1 and two shadowed incidents or exercises.
- Train Communications Leads for one day on status writing, customer segmentation, legal boundaries, and contractual clocks.
- Require domain responders to demonstrate dashboards, access, rollback, failover, escalation, and relevant playbooks.
- Require two shadow shifts before independent primary duty.
- Renew certification annually through simulation.
- Maintain the training, assessment, and certification register as operational and audit evidence.
- Nominate an incident-management champion in each of the 28 teams.
22. Exercise command, recovery, and tool failure (depends on: 11, 20, 21)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first cross-company command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous operational and security events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors when relevant.
- Conduct at least one unannounced after-hours paging test before the audit and two annually thereafter.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure acknowledgement, command, customer notice, mitigation decision, handoff, and recovery times.
- Create tracked corrective actions for every material exercise finding.
23. Publish the signed policy and control set (depends on: 6, 7, 8, 9, 10, 12, 14, 15, 16, 18, 19)
Convert the design into concise documents that people can use during an incident. The actual operating process must also be the documented and audited process.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, evidence, and exception standards.
- Include one-page cards for severity, roles, authority, escalation, and communication timings inside the incident tool.
- State explicitly that responders support only owned or formally accepted and trained service portfolios.
- Include the compensation structure, fatigue rules, declaration rights, and non-retaliation commitment.
- Obtain approval from the CTO, HR, Legal, Security, Compliance, and Internal Audit.
- Announce the policy at an all-hands and through team briefings.
- Create an exception register with owner, rationale, compensating control, approver, review date, and expiry.
- Version every policy change. Do not rewrite historical records when the process changes.
24. Instrument the scorecard and review forums (depends on: 3, 11, 18, 19)
Measure the process before the pilot so failures become visible immediately. Report medians and 90th percentiles rather than averages alone.
- Measure impact-to-detection, detection-to-declaration, declaration-to-command, acknowledgement, mitigation, recovery, and resolution.
- Split results by severity, service tier, customer journey, region, detection source, and business-hours status.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track notification episodes, actionability, duplicates, after-hours load, missed detection, routing errors, and page-budget breaches.
- Track postmortem timeliness, action age, due-date performance, verified effectiveness, and repeated contributing factors.
- Track journey availability, error-budget burn, failed or delayed value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight work, recovery days, exceptions, sentiment, and responder attrition.
- Hold a weekly Incident Review Board chaired by the Head of Reliability with relevant directors.
- Hold a monthly executive reliability review and a quarterly control and resilience review with Internal Audit.
- Reconcile incident records monthly against support cases, customer complaints, status history, credits, and major operational anomalies to detect under-reporting.
- Use team-level scorecards to direct help and investment. Never penalize an individual for good-faith declaration.
25. Pilot the complete process on the payment path (depends on: 9, 11, 13, 20, 21, 23, 24)
Run a four-to-six-week pilot across the highest-risk journey before expanding. The interim operating floor remains active for the rest of the company.
- Include payment orchestration, ledger application, PostgreSQL platform, API edge, authentication, settlement, reconciliation, Kubernetes platform, and Support intake.
- Include teams with existing on-call experience and teams new to the model.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking together.
- Run legacy and new paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently taking command.
- Correct critical process or tooling defects within 48 hours.
- Exit only after 95% timely command assignment, 95% communications compliance, no unpaid pages, complete required postmortems, tested fallbacks, and at least 50% lower pilot noise.
- Publish the pilot results, defects, and policy changes company-wide.
26. Roll out by risk with readiness gates (depends on: 25)
Expand in controlled waves and finish early enough to accumulate operating evidence before the audit. A calendar date does not override a failed readiness gate.
- Roll out remaining Tier 0 domains first, followed by Tier 1, Tier 2, and Tier 3.
- Use four waves of six to eight teams, each lasting two to three weeks.
- Gate each service on catalog ownership, appropriate coverage, active compensation, trained responders, access, tested escalation, alert quality, runbooks, and a passed tabletop.
- Require at least six responders only for direct 24x7 technical rotations. Use business-hours coverage for lower tiers.
- Give each wave a named coach and director sign-off.
- Reschedule failed gates or use a time-limited executive exception with compensating controls. Do not create silent waivers.
- Disable legacy human-paging routes after verified cutover for each wave.
- Run a quota-based noise reduction sprint in every wave, starting with the highest-volume rules.
- Pair every suppression with a compensating-detection check.
- Publish an internal adoption dashboard by team, service tier, coverage, and control gap.
- Complete critical coverage by approximately week 12 and all 28 teams by week 18.
27. Prove SOC 2 operating effectiveness (depends on: 22, 23, 24, 26)
Generate evidence through normal operation rather than reconstructing it before fieldwork. Test both control design and consistent execution.
- Map controls to the applicable Trust Services Criteria with Compliance and the auditor, including monitoring, incident identification, response, recovery, communications, and availability.
- Retain approved policies, exceptions, service ownership, schedules, compensation activation, access reviews, training, incidents, communications, reportability decisions, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, missed timelines, non-reportable decisions, and exercises in the testing population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the evidence populations and interviews expected from the external auditor.
- Correct deviations through tracked actions with owners and dates. Never edit history to create apparent compliance.
- Verify that evidence retention covers the full auditor-defined observation period.
- Brief commanders, engineers, Support, and Compliance on the actual process without scripting inaccurate answers.
28. Inspect, adapt, and institutionalize ownership (depends on: 26, 27)
Prevent the process from decaying after rollout or the audit. Change it using measured operating evidence rather than opinion.
- Review the policy after 90 days of live operation using severity calibration, page load, missed detection, communication compliance, action closure, fatigue, and survey results.
- Remove steps that create work without reducing risk. Add controls only where incidents, exercises, or evidence show a gap.
- Reassess Tier 0 and Tier 1 classification and domain boundaries every six months.
- Assign permanent owners for policy, catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review compensation bands, rotation burden, accommodations, and staffing annually.
- Report severe incidents, credits, overdue high-risk actions, and ledger concentration risk to the board or risk committee quarterly.
- Maintain the ledger blast-radius program as an executive risk until failover, degraded mode, reconciliation, and regional independence meet approved objectives.
- Evaluate follow-the-sun coverage using one year of actual activation and staffing data.
- Build year-two plans for automated mitigation, safer deployments, graceful degradation, and error-budget release controls.
--- PROPOSAL 3 ---
Proposal ID: f61aca40-a8a9-4e91-a5f3-ac1d03634441
Content:
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of 'not reportable'.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (29):
1. Charter the program, fund it, and start the audit clock
Convert the CEO email into a company operating process with one accountable owner, a budget, and a dated timeline that starts this week.
- Name the CTO as executive sponsor and appoint a Head of Reliability & Incident Management as the single accountable owner with full-time authority over all 28 teams.
- Stand up a three-person program office: program lead, platform engineer, reliability analyst.
- Form an eight-person steering group spanning Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, and HR. It proposes; the sponsor decides within 48 hours. Never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one human-paging path, one incident record, one postmortem format, paid on-call, mandatory action tracking, and named service ownership.
- Approve budget anchored against the $1.3M in SLA credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs, and 15% reserved engineering capacity.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Publish a one-page charter company-wide on day 2. Incident response is a company process, not a per-team preference.
- Timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day interim command floor (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place this week so the next outage has a named commander and the evidence clock starts immediately.
- Publish a one-page interim severity card and one declaration path: a Slack command, a phone number, and the existing pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs from the 12 teams already on-call. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical action items into complete, re-plan, or formally risk-accept within 30 days, prioritising ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline of incidents, alerts, and money lost (depends on: 1)
Rebuild the facts before locking any design. This is both the design input and the frozen 'before' picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines in one signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the written fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written, signed deal.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- State explicitly that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Map compliance, evidence, and audit requirements from day one (depends on: 1)
Design evidence as a by-product of operations, not a reconstruction before the auditor arrives. The interim process in week 1 is already evidence.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for internal and external communication; CC5 for control activities; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure with verification, training register, drill records, reportability decisions including 'not reportable'.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records. Require role-based access, MFA, and periodic access review.
- Version and approve all policy documents from day one: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard, Alert Quality Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has a named owner. A wrong owner recreates the pager objection. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Assign one accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. **Tier 1**: customer-facing but degradable. **Tier 2**: internal or batch. **Tier 3**: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for every Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform. They fail differently and need different hands.
- Orphan services get an owner within 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Missing ownership or runbooks blocks Tier 0/1 releases.
7. Adopt the severity scale, declaration rights, and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix. Anyone may declare. Nobody is penalised for over-declaring.
- **SEV1 (crisis):** money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; deploy freeze; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads; business-hours comms; postmortem if customer-detected, over two hours, a repeat, or credit-generating.
- **SEV4:** no customer impact. Ticket only. Never pages.
- Payments-specific anchors: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident become at least SEV2.
- Only the Incident Commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
8. Define incident roles, authority, dual-control, and handover discipline (depends on: 7)
Solve 'nobody in charge for an hour' by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Command claimed within 5 minutes and stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Staff 24x7 with a three-layer coverage model, not 28 night rotations (depends on: 6, 8)
Do not create 28 night rotations. That is precisely what engineers are rejecting. Centralise coordination in a trained command corps and keep technical ownership local.
- **Layer A — Incident Command corps:** approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. Paired with a Communications Lead pool of approximately 18 from Support, CS, and engineering management, and a scribe pool used as the training entry point.
- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat a Lisbon or APAC follow-the-sun cell as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
10. Approve paid on-call, New York labor compliance, and fatigue safeguards (depends on: 4, 9)
Unpaid on-call in New York is both a retention problem and a wage-hour exposure. Pay must be live in payroll before any new mandatory night rotation starts. Publish the actual numbers, then ask people to sign up.
- Indicative scheme locked by HR, Finance, and employment counsel within 14 days: approximately $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as approximately 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
- Publish on-call load by team quarterly.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
11. Set the alert-quality standard and a hard page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target: 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, noise below 15%, with no loss of Tier 0/1 detection coverage.
12. Detect payment and ledger failures before customers do (depends on: 6, 11)
The goal is blunt: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
- Run the replay test: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish replay coverage as a leading metric.
13. Consolidate to one paging platform, one incident record, one status page (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
14. Codify the five-minute escalation path and live execution doctrine (depends on: 8, 9, 13)
Write one unskippable path from 'something looks wrong' to 'someone is in charge'. The default action is never waiting. If nobody claims command within 5 minutes, the platform assigns it and announces it.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute.
- Separate diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Closure requires a stability observation window and explicit handback.
15. Set the service-readiness bar and write major-incident playbooks (depends on: 6, 9, 12)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human at night. The shared ledger cluster is the single largest structural risk.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction: partitioning, read replicas, isolation of non-critical readers, with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
16. Run one communications clock for internals, customers, and regulators (depends on: 7, 8, 13)
Replace 'whoever is around' with one timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure. State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is 'not reportable'.
- Obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
17. Tie incidents to SLA credits and true financial cost (depends on: 7, 16)
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
18. Make blameless postmortems mandatory with one format and fixed deadlines (depends on: 7, 8)
Replace 'some incidents, various formats' with one mandatory format, fixed deadlines, and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did?
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and are never the commander of the incident under review.
- Publish a searchable library and a quarterly top-five recurring causes analysis, with restricted versions for security or privileged content.
19. Enforce action ownership, reserved capacity, and tracking (depends on: 13, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity. Closing a ticket without evidence does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
- Target: 90% of high-priority actions closed by due date within two quarters.
20. Train and certify every role before independent duty (depends on: 8, 14, 16, 18)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it. The certification register is an audit artefact.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days plus shadowing): command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
21. Rehearse with tabletops, game days, and unannounced drills (depends on: 13, 15, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-mitigation.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
22. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 14, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 10, 12, 13, 15, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse.
- Publish a one-page result to the whole company.
24. Instrument metrics, dashboards, review cadence, and anti-gaming (depends on: 13, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
25. Roll out in risk-ordered waves with readiness gates (depends on: 23, 24)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, never waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave. Give every team its ranked list of noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Review missed detections monthly with the same seriousness as noise.
- Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
26. Run change management, fairness, and pager culture from day one (depends on: 4, 10)
Run this in parallel from day one. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
27. Produce SOC 2 evidence by operating, test internally, and mock-audit (depends on: 22, 24, 25)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence. Documented exceptions beat a claim of perfection.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5, CC2.2/CC2.3, CC5, and A1.2. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including 'not reportable', postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in months 4 and 6, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
28. Maintain the program risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and pre-commit the response. Review monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
29. Inspect at 90 days, lock year-two ownership, and prevent decay (depends on: 24, 25, 27)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 4 ---
Proposal ID: 91ea9817-3664-46dd-a758-888405cea636
Content:
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- By day 7, every suspected major incident uses one record, one channel, and a named commander; Support can declare without engineering confirmation.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of customer-visible SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (26):
1. Charter the program, fund it, and start the audit clock
Turn the CEO email into a chartered company program within 48 hours. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability as full-time owner, with a three-person office: program lead, platform engineer, and analyst.
- Form a small decision group of Engineering, SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit. It proposes. The sponsor decides within 48 hours.
- Lock the non-negotiables now: one severity scale, one human-paging platform, one incident record, one postmortem format, mandatory action tracking, paid on-call, and named service ownership.
- Confirm the SOC 2 Type II observation window with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, exercises, and reserved engineering capacity against the $1.3M in SLA credits.
- Reserve 15% of engineering capacity for detection, runbooks, and incident actions, protected by the sponsor.
- Publish a one-page charter on day 2. Clock: floor by day 7, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for policy, tooling, or the audit. Put a crude but real process in place this week so the next outage already has an owner.
- Publish a one-page interim severity card and one declaration path: chat command, phone number, and current pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline, and one naming convention for every suspected major incident.
- Triage the 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
- Replay one of the two nobody-in-charge incidents as a tabletop within 14 days using this floor process.
3. Rebuild the forensic baseline and incident-replay catalog (depends on: 1)
Rebuild the facts before locking design. This is the design input, the frozen before-picture for the CEO, and the test set for detection work.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact missing signal. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Name the top 50 noisy rules and every rule with no owner or runbook.
- Quantify true cost: credits, failed payment volume, delayed value, reconciliation breaks, and engineering hours lost.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11 of 64 actions closed, on-call in 12 of 28 teams.
- Build a **replay catalog**: for each historical incident, the detector that should now fire, at which minute, and the owner of the gap.
4. Publish the on-call fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders. Do not punish them by making them the company's permanent night watch.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the **fairness contract** in writing: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive exemption path for health, disability, or caregiving.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and journey tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has one owner. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Record owning team, manager, chat channel, escalation policy, dashboards, runbook, dependencies, regions, and data stores for all 180 services.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, refund, report, onboard — to services, data stores, regions, sponsor banks, and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform. They fail differently.
- Orphan services get an owner in 30 days or an approved decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Coverage follows the journey, not the org chart. Small teams that own Tier 0 pieces get headcount, service reassignment, or membership in a domain rotation. Never a two-person 24x7 rota.
6. Lock severity levels, integrity flags, and the incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Classify on actual or credible customer, financial, security, and contractual harm, never on who reported it or how hard the fix looks.
- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 contained: narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages a human.
- Attach an **integrity, security, settlement, or regulatory flag** to any severity. The flag forces dual-control, Legal, and the reportability checkpoint without inventing a fifth level. A one-customer ledger corruption is still a flagged crisis.
- Auto-escalate to at least SEV2: any ledger-cluster event, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident.
- Anyone may declare. Nobody is punished for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- Communications Lead: single voice for the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists. A human validates at SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory text. The commander owns the facts.
- Distinct people for command, comms, and technical lead at SEV1. Comms and scribe may combine only for bounded SEV2.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
- Every role assignment and handover is announced verbally and in writing with the exact time and open risks.
8. Staff 24x7 with command, domains, and a Duty Triage Desk (depends on: 4, 5, 7)
Do not create 28 night rotations. That is what engineers are rejecting. Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified people from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to eight months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate ownership into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- **Duty Triage Desk**: a small paid overnight first-line rotation that owns the first ten minutes of ambiguous, unowned, or low-confidence pages. It verifies, enriches, and applies only the runbook's safe steps, then wakes the owning domain. It never sits on a suspected ledger, payment-halt, or security event. Those page commander and likely domains immediately, in parallel.
- Overnight comms: SEV1 pages a Communications Lead 24x7. Customer-visible SEV2 lets the commander publish the first status from a template; Comms is paged if the incident is still open at 30 minutes or a top-100 account is affected.
- Staffing math: 260 engineers can sustain 10–12 domain rotations, one command corps, and one triage desk. They cannot sustain 28 night rotas. Role exclusivity: nobody is primary on two rotations in the same week. Commanders may also be domain responders in different weeks.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
- Evaluate follow-the-sun as a 12-month option, not a year-one dependency. Seed Layer B from the 12 teams already on-call.
9. Pay on-call, meet New York labor rules, and cap fatigue (depends on: 4, 8)
Unpaid on-call in New York is a retention problem and a wage-hour exposure. Pay must be in payroll before any new mandatory night rotation starts.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, $400–$800 for Comms or triage-desk duty, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- **Hard gate: no mandatory night rotation starts before compensation is live in payroll.** Interim duty is paid retroactively. Budget roughly $0.9M–$1.2M a year, then refine with actual rotation count.
10. Enforce an alert-quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers, then replay the past (depends on: 5, 10)
Stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. A detector is not done until it would have caught the last 12 months.
- Define SLIs and SLOs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner, processor, and sponsor-bank notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a named defect class with a mandatory tracked action.
- Run the **replay test** against the S3 catalog. For each of the 31 incidents, name the detector that would now fire and at which minute. Close gaps the replay exposes before calling detection improved.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One chat command creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy. Unowned pages go to the Duty Triage Desk and Duty Command, and log a catalog defect.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Codify escalation, the five-minute command rule, and vendor incidents (depends on: 7, 8, 12)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting. Payments also fail at processors and banks, which you cannot patch.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Vendor-incident class: processor, sponsor-bank, card-network, or cloud-control-plane failure. Command is still required. Success is time-to-customer-notice, time-to-failover-decision, and queue management, not root cause at the vendor.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Write live execution doctrine, readiness bars, and ledger playbooks (depends on: 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving root cause. A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis. Split diagnosis and mitigation workstreams once staffing allows.
- Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Drain backlogs under control. Complete reconciliation before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date. Existing detection stays on while gaps are repaired.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers, with board-visible milestones and executive-signed RPO/RTO.
15. Run one communications clock for internals, customers, and account managers (depends on: 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for customer-visible SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys, not internal service names. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer record.
- Language rules: state impact and the next update time. **Never speculate** on cause, recovery time, data integrity, or blame.
16. Operationalize regulatory notice, partner clocks, and SLA credits (depends on: 6, 15)
In payments, some incidents start a legal clock at detection. Tie incidents to money so Finance does not learn about outages from invoices.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when not reportable**, with facts, decision-maker, and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families. Track failed payment volume, delayed value, and reconciliation breaks as the true incident cost.
17. Make postmortems mandatory and actions enforceable (depends on: 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth. Eleven of 64 closed is a process nobody enforces.
- Mandatory for every SEV1 and SEV2, any incident a customer detected first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and gap, timeline, response analysis, contributing conditions, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items. Facilitators are trained and are never the commander of the incident under review.
- Every action gets one named owner, priority, due date, expected risk reduction, verification method, and a ticket created automatically. If it is not on the board, it does not exist.
- Classes: containment in 7 days, corrective in 30, strategic in 90. SEV1 recurrence-prevention enters the next sprint ahead of roadmap work. Reserve 15% capacity.
- Overdue ladder: manager at +7 days, director at +14, CTO at +30. Overdue high-risk items need written residual-risk acceptance and can block related releases. Verify effectiveness before closing.
18. Train and certify every response role before independent duty (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Duty Triage Desk, half day: enrichment, safe-step limits, when to wake a domain immediately, when not to delay.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
19. Publish Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17)
Collapse the design into a document people will actually open mid-outage, and make it official. Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
20. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 18, 19)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, Duty Triage Desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
21. Instrument metrics, reviews, and anti-gaming (depends on: 12, 17, 20)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- Response: time from first impact to detect, declare, commander, acknowledge, mitigate, and resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
22. Rehearse with tabletops, game days, and night drills (depends on: 12, 14, 18, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced overnight paging drill to measure real acknowledgement times.
- One combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO, before the audit.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
23. Burn down alert noise and roll out by risk with gates (depends on: 20, 21)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Noise reduction is a quota inside each wave, not a background hope.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy paging paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Give every team its noisiest rules from the baseline. Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed. Pair every suppression with a compensating-detection check.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
- If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
24. Operate the fairness and culture program in parallel (depends on: 4, 9)
Run this from the moment the deal is published. Engineers judge the process on fairness. Executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
25. Prove SOC 2 operating effectiveness before fieldwork (depends on: 19, 21, 22, 23)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (depends on: 21, 23, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, metrics, and the ledger blast-radius workstream, independent of the audit cycle.
- Pre-committed contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 5 ---
Proposal ID: b427ef6b-42c2-41ac-8f9d-992d99ba5e20
Content:
Estimated Complexity: high
Success Metrics: - Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (30):
1. Charter the program and start the audit clock
Turn the CEO email into a chartered company program with one accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
- Confirm the SOC 2 Type II observation window with the auditor; map the future response controls to applicable Trust Services Criteria; define evidence set and retention requirements.
4. Listening tour, resistance map and the written on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (depends on: 5, 7)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 4, 8)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (depends on: 5, 10)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (depends on: 7, 13)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 12, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
- Agree with Legal and Finance the availability measurement method per contract and per component; compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide credit posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain; attribute credits to root-cause families for investment decisions.
18. Blameless postmortem standard and Incident Review Board (depends on: 6, 7, 12)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
19. Action ownership, reserved capacity and enforcement (depends on: 12, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
20. Training and certification academy (depends on: 7, 13, 15, 18)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
21. Exercise programme: tabletops, game days and unannounced drills (depends on: 12, 14, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
22. Publish Incident Management Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
24. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
25. Alert noise burn-down campaign (depends on: 10, 12, 23)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
26. Wave rollout to all 28 teams with readiness gates (depends on: 23, 24, 25)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
27. Change management, fairness and pager culture (depends on: 4, 9, 23)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
28. SOC 2 evidence by design, internal testing and mock audit (depends on: 22, 24, 26)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
29. Risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
30. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 24, 26, 28)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Instructions:
1. Carefully evaluate each proposal considering:
- Clarity and detail of the proposed steps.
- Overall alignment with the Main Objective.
- Quality and relevance of the proposed success metrics.
- Coherence of dependencies between steps (if any).
2. Think deeply about each proposal, analyzing its strengths and weaknesses. Consider how well each proposal addresses the main objective and which one would be most effective.
After your thorough analysis, give the Proposal ID of the selected proposal and a one- or two-sentence justification of your choice.
{"selected_proposal_id": "30f15b26-0ec0-4756-9f4a-6279eb889884", "justification": "Proposal 1 is the most complete match to the brief: it defines severity, roles, 24x7 staffing, paid on-call, detection, communications, postmortems, and metrics in operational detail, and it puts a live command floor in place on day 7.\n\nIt uniquely adds a **Duty Triage Desk**, change intelligence, concurrency rules, and Support as a detection path, which directly attack unclear ownership, customer-first detection, and pager resistance."} |
| deepseek-v4-pro_voter_5 | deepseek-v4-pro · deepseek/deepseek-v4-pro |
Proposal A · opus5_refine_1 winner | Proposal 1 is the most comprehensive and detailed, explicitly defining every required element—severity, roles, 24x7 coverage, paid on-call, alert quality, detection/replay, communications, postmortems, metrics, and SOC 2 readiness—with clear dependencies and measurable success metrics. It uniquely adds high-value components such as change intelligence, SLA credit guardrails, and a Duty Triage Desk, which directly target the stated pain points. |
50.2k in · 4.2k out · 1 min 10 s | show[SYSTEM]
You are an expert and objective evaluator of project plan proposals.
Your task is to select the BEST proposal based on criteria of completeness, clarity, and alignment with the main objective.
Use your internal reasoning processes to thoroughly analyze each proposal, considering all aspects and implications.
Take as much time and space as you need to evaluate each proposal in depth before making your decision.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
Proposals to Evaluate:
--- PROPOSAL 1 ---
Proposal ID: 30f15b26-0ec0-4756-9f4a-6279eb889884
Content:
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90 a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95% internally and externally.
- Top-100 account outreach completed within 30 minutes of SEV1 in 95% of qualifying cases, using the approved briefing pack.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- Change-correlated incidents are identified within 10 minutes of declaration in 90% of cases, and the change-correlated share of incidents declines quarter over quarter.
- SLA credits fall to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, one concurrency drill and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and a tested read-only mode by month 6, with milestones reported to the board quarterly.
Steps (35):
1. Charter the programme: one owner, one mandate, funded, dated against the audit
Turn the CEO email into a chartered company programme with a single accountable owner and authority over all 28 teams. Incident response stops being a per-team preference and becomes a **company operating process**.
- Name the CTO as executive sponsor and a full-time Head of Reliability & Incident Management as accountable owner, supported by a programme office of three: programme lead, incident-platform engineer, reliability analyst.
- Form a decision group (Engineering, Platform/SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables on day one: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, named service ownership, and **paid on-call**.
- Publish the timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal control test month 5, mock audit month 7, SOC 2 fieldwork month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.3M/yr, training and exercises, plus 15% of engineering capacity reserved for reliability work.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence immediately.
- Publish a one-page interim severity card and a single declaration path: one chat command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim 24x7 Duty Incident Commander roster (primary plus backup) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document and one naming convention for every major incident, starting now.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is the design input, the executive narrative and the frozen "before" picture for the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the **detection backlog**.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Correlate incidents with deployments, config changes and feature flags to quantify how many started with a change we made.
- Quantify true cost beyond credits: failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost, churn risk on named accounts.
- Freeze the baselines in a signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Audit, legal and evidence scoping in week one (depends on: 1)
Design evidence as a by-product of operating, never as a reconstruction before fieldwork. Settle the scope and the legal handling of incident records now, not in month seven.
- Confirm with the external auditor the Type II observation window, the incident population definition, and the evidence they will sample. Everything from the day-7 floor onwards must count.
- Map incident response to the Trust Services Criteria with Compliance: CC7.2–CC7.5 (monitoring, identification, response, recovery), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), A1.2 (availability).
- Define the evidence set and where it is produced automatically: incident records, paging and acknowledgement logs, role assignments, status-page history, reportability decisions, postmortems, action closure with verification, training register, drill records, access reviews.
- Agree retention, confidentiality, legal hold and access rules. Decide with counsel which postmortem content is privileged and how privileged material is segregated **without making the ordinary postmortem secret**.
- Start the obligation matrix with Legal: NYDFS 23 NYCRR 500, state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network contractual windows, cyber-insurer notice.
- Begin monthly evidence sampling from month 1 so control operation is visible long before the audit.
5. Listening tour, resistance map and the written on-call deal (depends on: 1, 3)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written deal signed by the sponsor.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
- State the hard gate publicly: no new mandatory night rotation starts before compensation, training, runbooks and staffing rules are live.
6. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- Record for every service: one accountable team, engineering manager, chat channel, escalation policy, dashboards, runbook link, dependency list, regions and data stores.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record SLO, RTO, RPO, active-active or region-bound, failover method, and the dependencies that make nominal two-region redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
7. Severity standard, declaration rights and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning responders and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and responders paged, comms and scribe if customer-visible, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, credit-generating or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Attach a financial-integrity flag or a security flag to any severity. The flag forces dual control, Legal engagement and the reportability checkpoint without inventing a fifth level.
- Anchor on payments reality alongside error rates: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including a customer report.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
8. Roles, authority, concurrency and handover discipline (depends on: 7)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice to the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1 and SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, Finance and Vendor Management join on defined triggers rather than by invitation.
- Rules: command claimed within 5 minutes and stated in channel; distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- **Concurrency doctrine:** two simultaneous SEV1/SEV2 incidents activate the secondary commander and a surge roster; a designated Multi-Incident Coordinator arbitrates shared resources such as the ledger, the database platform and the deploy freeze.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
9. Three-layer 24x7 coverage: command corps, domain rotations, overnight triage desk (depends on: 5, 6, 8)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving roughly one primary week per person every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement and reconciliation, auth, API edge, Kubernetes platform, data and reporting, partner integrations). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Publish the staffing arithmetic: 10–12 domain rotations of six to eight people plus a 30-person command corps means roughly 100 of 260 engineers carry a night obligation, at about one week in six. Twenty-eight independent rotations would be unstaffable and is therefore rejected on the numbers.
- Teams too small for a fair rotation get headcount, service reassignment, or a time-limited executive exception. **Never a two-person 24x7 rotation.**
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: a human acknowledgement within 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
10. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 5, 9)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour, and a premium for the overnight triage desk.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per responder per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty since day 7 is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons without career penalty, and on-call load is published quarterly by team.
11. Alert quality contract and page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression can never hide an incident.
12. Detection uplift on the money path, validated by incident replay (depends on: 6, 11)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action and a missed-detection review.
13. Change intelligence and deployment safety on the critical path (depends on: 6, 12)
Most of these incidents start with something we changed. Make change the first hypothesis the tooling answers, and make changes safer to reverse.
- Stream every deployment, configuration change, feature-flag flip, schema migration and infrastructure change into the incident timeline with service labels and owners.
- Give the commander an automatic "what changed in the last 60 minutes on the affected journey" panel at declaration time.
- Require Tier 0/1 changes to be progressively delivered with a documented rollback that is tested, time-bounded and executable by the on-call responder without the author.
- Treat ledger schema migrations and settlement-affecting changes as a separate class: dual approval, rehearsed rollback, no deploys inside the settlement window.
- Enforce a change freeze during SEV1 and SEV2, lifted only by the commander and logged.
- Report change-correlated incidents monthly; a rising ratio is a signal to strengthen release safety, not to blame a team.
14. One pager, one incident record, one status page — migrated without a detection gap (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy paging path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access, MFA and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page delivered outside this platform creates no on-call obligation.
15. Escalation ladder and the five-minute command rule (depends on: 8, 9, 14)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts and support entitlements: AWS and database premium support, processors, sponsor banks, card networks.
16. Live execution doctrine and payments safety rules (depends on: 8, 14, 15)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed in advance rather than improvised.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and automatic reopening if impact recurs.
17. Service readiness bar, major-incident playbooks and ledger blast-radius reduction (depends on: 6, 9, 16)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human, and the shared ledger is the largest structural risk in the estate.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as a concentration risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers, stronger regional independence — with board-visible quarterly milestones.
- Enforcement: no readiness sign-off, no night paging alerts unless the manager accepts the gap in writing with an expiry date and a compensating control. Never respond to a gap by turning detection off.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
18. Internal communications protocol (depends on: 8, 14)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, Sales and Security.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change. First internal brief within 10 minutes of SEV1 and 15 of SEV2.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
19. Customer communications, status page and account-manager outreach (depends on: 7, 18)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; monitoring notice after mitigation; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, monitoring, resolution.
- Component-level status page mapped to customer journeys rather than internal service names, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1 and 60 of SEV2, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state observed impact, affected capabilities, any workaround and the next update time; **never speculate on cause, recovery time, data integrity or blame**.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
- Record any legally required restriction or delay of public detail, its approver, and the alternative stakeholder plan.
20. Support and account management as a detection and intake tier (depends on: 7, 12, 19)
Customers detected 40% of incidents first, which means the front line already holds the signal. Turn Support and account managers into an instrumented detection channel rather than a bystander.
- Give Support explicit declaration rights, a one-page trigger card, and a macro that opens an incident candidate directly in the incident platform.
- Automate the clustering rule: five similar tickets or calls in ten minutes auto-creates a triage incident assigned to the duty commander.
- Route credible partner, processor and sponsor-bank notifications into the same declaration path within five minutes.
- Generate the affected-customer list automatically from journey telemetry and the incident record, and push it to Support, CS and the account-manager briefing.
- Train Support and account managers on approved language and prohibit independent technical explanations to customers.
- Measure and publish "signal was in Support before it was in monitoring" as a detection defect, and feed each instance into the detection backlog.
21. Regulatory, partner and legal notification playbook (depends on: 4, 7, 19)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Complete the obligation matrix started in scoping and have counsel validate the triggers, deadlines, channels and submitting authority for each obligation.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with facts considered, decision-maker, timestamp and reassessment trigger.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, card networks, outside counsel, cyber-insurer, critical vendors.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the operational facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly inside the exercise programme, including one full regulator-notification simulation per year.
22. SLA credit workflow, true cost model and incentive guardrails (depends on: 7, 19)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component, including how partial degradation counts.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Guardrail against the perverse incentive: better detection will surface incidents that previously went unbilled, so credits may rise before they fall. Publish this expectation to the executive team in advance, and make it a written rule that no finance or commercial pressure may influence severity or declaration.
23. Blameless postmortem standard and Incident Review Board (depends on: 4, 7, 8)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger; any control failure.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning engineering director is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions where security, privacy or privileged content requires it.
24. Action ownership, reserved capacity and enforcement (depends on: 14, 23)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, accountable manager, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness through tests, telemetry, exercises or production evidence before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
25. Training and certification academy (depends on: 8, 15, 18, 23)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, the 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, contractual clocks, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days of joining a domain.
- Certification is valid 12 months and renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
26. Exercise programme: tabletops, game days, night drills and vendor rehearsals (depends on: 14, 17, 25)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander. First company-wide tabletop within 30 days.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times, plus a concurrency drill with two simultaneous incidents.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down. Rehearse joint escalation with AWS, a processor and a sponsor bank.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
27. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 15, 18, 19, 21, 23, 24)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, escalation and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix, Readiness Bar.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
28. Pilot on the payments critical path (depends on: 10, 12, 13, 14, 17, 25, 27)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement and reconciliation, PostgreSQL platform, Kubernetes platform, API edge, auth and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, action tracking, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The programme lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time in 95% of cases, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
29. Alert noise burn-down campaign (depends on: 11, 14, 28)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → under 1,500 pages a month by day 90, under 500 with noise below 15% by month 6.
30. Metrics, dashboards, review cadence and anti-gaming (depends on: 14, 23, 28)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time from first impact to detect, declare, commander, acknowledge, mitigate, resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors, change-correlated incident ratio.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
31. Wave rollout to all 28 teams with readiness gates (depends on: 28, 30)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the programme office for three weeks; the director signs the gate.
- **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard by team, tier and control gap.
32. Change management, fairness and pager culture (depends on: 5, 10, 28)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
33. SOC 2 evidence by design, internal control test and mock audit (depends on: 4, 27, 30, 31)
The real process is the evidence. Never build a parallel audit process, and never reconstruct records after the fact.
- Maintain the evidence set defined in scoping, produced automatically and indexed: versioned policies and exceptions, catalog records, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Sample incidents monthly from first signal through verified corrective action, deliberately including customer-reported events, downgraded incidents and missed timelines.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling end to end and reporting gaps to the sponsor.
- Run a formal mock audit in month 7 using the populations, evidence requests and interviews the auditor will use: a commander, a random engineer, Support, Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
34. Programme risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps and immediate staffing intervention.
- **A major SEV1 mid-rollout:** the programme lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
35. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 30, 31, 33)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, exercise calendar and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and quarterly board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 2 ---
Proposal ID: 1f1527bb-d992-451c-bbc4-9e61ffdb1f30
Content:
Estimated Complexity: high
Success Metrics: - By day 7, every suspected major incident uses one record, one coordination channel, and a named Incident Commander within 10 minutes.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes.
- By month 3, at least 95% of SEV1 and SEV2 incidents have a named commander within 5 minutes.
- By month 3, at least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and appropriate technical coverage.
- By week 18, all 180 services have an owner, tier, escalation path, and tested coverage model.
- No new mandatory night rotation begins before compensation, training, access, runbooks, and minimum staffing are active.
- Every direct 24x7 technical rotation has at least six qualified responders or an approved, expiring executive exception.
- No responder is routinely primary more often than one week in six or assigned to two simultaneous primary rotations.
- Median impact-to-detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- By day 90, incident replay identifies a current detector for at least 90% of the 31 historical incidents, including its expected detection minute.
- Median time to mitigation falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of applicable SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet the required internal and customer update cadence by month 3.
- Monthly human notification episodes fall from the current 3,400 alert events to no more than 1,500 by day 90 and 500 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without reducing Tier 0 or Tier 1 replay coverage.
- Average after-hours load remains at or below two notification episodes per responder per week; every sustained breach receives a dated remediation plan.
- All six monitoring sources route human pages through the controlled paging platform by week 12, with direct legacy routes retired by week 18.
- All required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 from month 3 onward.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due dates with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Reportability is assessed and recorded within 1 hour for 100% of SEV1 and qualifying SEV2 incidents, including not-reportable decisions.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Approved ledger RTO and RPO, controlled failover, read-only mode, and post-recovery reconciliation are exercised before month 6.
- Quarterly responder surveys reach at least 75% favorable responses for fairness, ownership boundaries, compensation, and sustainability by month 6.
- Monthly contracted availability meets or exceeds 99.95% by month 6 using the contractually authoritative measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- The month-7 mock audit finds no unowned high-risk control gap and at least 95% of sampled incidents contain complete operating evidence.
- The SOC 2 Type II incident-response controls complete external testing without an unresolved material exception.
Steps (28):
1. Charter the program and fund immediate action
Make incident management a **company operating process** within 48 hours. Give one accountable leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Include Engineering, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit in a decision group. The sponsor resolves blocked decisions within 48 hours.
- Approve non-negotiables: one severity model, one human-paging path, one incident record, paid on-call, named service ownership, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, resilience, and incident actions.
- Fund tooling, compensation, training, exercises, and program staff. Compare the cost with the existing $1.3M in annual SLA credits.
- Set the schedule: operating floor by day 7, design during weeks 1–4, pilot during weeks 5–10, rollout during weeks 11–18, control tests in months 4 and 6, and mock audit in month 7.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for the final policy or new tooling. Put a minimum viable process into operation immediately and retain evidence from the first day.
- Publish a one-page interim severity guide, declaration procedure, role card, and communications clock.
- Provide one monitored declaration route through chat, telephone, and the current paging environment.
- Create one channel, bridge, timeline, and incident identifier for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced managers and engineers.
- Require a named commander within 10 minutes. The duty engineering director assumes command if the command page is unclaimed.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Stop uncompensated mandatory after-hours expansion. Pay interim duty under a temporary stipend, retroactive to program launch.
- Hold a 15-minute daily operational review until the permanent process is active.
3. Create the factual, contractual, and control baseline (depends on: 1)
Build one defensible baseline for process design, investment decisions, and SOC 2 testing. Preserve the original data so progress cannot be created by changing definitions.
- Reconstruct all 31 incidents from impact start through detection, declaration, command, mitigation, recovery, communications, credits, and corrective actions.
- Reconstruct the two incidents with unclear command minute by minute.
- Identify the missing signal for every customer-first detection.
- Inventory all six alert sources, 3,400 monthly alert events, duplicates, noisy rules, missing owners, and missing runbooks.
- Record current rotations, unpaid duty, overnight activations, schedule size, and uncovered services.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, 85% noise, $1.3M in credits, and 11 of 64 actions closed.
- Inventory customer-specific availability definitions, notice periods, credit terms, sponsor-bank obligations, and other incident-related contracts.
- Confirm the required SOC 2 Type II observation period and evidence expectations with the auditor during week 1.
4. Publish the on-call fairness contract (depends on: 1)
Treat pager resistance as a legitimate design constraint. Adoption depends on a written agreement that separates command from technical ownership.
- Interview representatives from all 28 teams and from Support, Customer Success, Security, and Operations.
- Separate concerns about unpaid work, sleep loss, unfamiliar systems, noisy alerts, inadequate runbooks, and blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for systems they own or have formally accepted and trained to support, and assisted by a separate Incident Commander.
- State that Platform may temporarily triage unknown ownership but does not inherit another team's service.
- Provide confidential accommodations for health, disability, pregnancy, or caregiving constraints without career penalty.
- Measure baseline trust, fairness, fatigue, and alert confidence. Repeat at days 60 and 120, then quarterly.
5. Build the service catalog and customer-journey map (depends on: 3)
Make a machine-readable catalog the source of truth for routing, impact analysis, status-page components, and control evidence. Every production service must have one accountable owner.
- Record the owning team, manager, business capability, repository, channel, dashboard, runbook, escalation policy, dependencies, regions, and data stores for all 180 services.
- Map initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, reporting, and onboarding to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing systems that can create material contractual impact.
- Classify Tier 2 as deferrable internal or batch systems and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional recovery mode, contractual commitments, and critical vendors for Tier 0 and Tier 1.
- Keep separate but coordinated ownership for the ledger application and PostgreSQL platform.
- Give orphan services an owner or approved decommission date within 30 days. Treat unowned Tier 0 services as release-blocking executive risks.
6. Adopt the severity model and incident modifiers (depends on: 3, 5)
Classify incidents by credible customer, financial, security, regulatory, and contractual harm. Start at the higher plausible severity while scope or integrity remains unknown.
- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; broad loss of a core payment journey; both regions impaired; or processing must be stopped. Immediately page all response roles and the Executive Duty Officer. Open the bridge within 5 minutes, freeze unrelated changes, issue internal notice within 10 minutes, publish applicable customer status within 15 minutes, start legal assessment within 1 hour, and require a postmortem.
- **SEV2 — major:** material payment degradation; settlement deadline at risk; regional impairment with reduced resilience; a critical customer or material cohort unavailable; or an SLA breach is likely. Page command and technical roles immediately. Issue internal notice within 15 minutes, applicable customer status within 30 minutes, and require a postmortem.
- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads. Page only when immediate action can reduce harm.
- **SEV4 — operational event:** no current customer impact and no credible imminent harm. Create a ticket and handle during normal hours.
- Add FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, and VENDOR modifiers. These invoke specialist controls without distorting customer-impact severity.
- Use at least SEV2 posture for unknown impact lasting 15 minutes, credible ledger-integrity risk, or a cross-domain incident without clear ownership.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Permit anyone to declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Publish a decision tree and examples based on the 31 historical incidents.
7. Define roles, authority, and handoffs (depends on: 6)
Separate coordination, communications, recordkeeping, and technical repair. One named person holds command continuously throughout every SEV1 and SEV2.
- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not act as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, executive updates, and coordination with Legal.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- **Subject-Matter Responders:** diagnose and mitigate only systems for which they have ownership, access, training, or a formally accepted support agreement.
- **Executive Duty Officer:** removes organizational obstacles and makes exceptional business decisions without displacing the commander.
- Add Security, Legal, Compliance, Finance, and Vendor Management on modifier-specific triggers.
- Require distinct commander, Communications Lead, scribe, and technical lead for SEV1. Communications and scribe may combine for the first 10 minutes of a bounded SEV2.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, reconciliation, and change evidence for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, including the exact transfer time and unresolved risks.
8. Create sustainable 24x7 coverage (depends on: 4, 5, 7)
Use central command coverage and risk-based technical coverage instead of creating 28 fragile night rotations. Responders do not carry pagers for unfamiliar code.
- Create a 24x7 Incident Command corps of approximately 30 certified people with primary and backup schedules. Two schedules create about 104 weekly assignments a year, or roughly three to four weeks per person annually.
- Create a 24x7 Communications pool of 18–24 trained people from Support, Customer Operations, Engineering Operations, and management.
- Create a similarly sized scribe pool. The backup commander temporarily records the first minutes if a scribe has not joined.
- Maintain a 24x7 Executive Duty Officer schedule and specialist contact paths for Security and Legal.
- Group Tier 0 and Tier 1 systems into roughly 8–12 coherent response domains only where responders share training, access, runbooks, and explicit support acceptance.
- Staff every critical domain with primary and secondary responders and at least six qualified people. Target eight where overnight activation is frequent.
- Give Tier 2 and Tier 3 services business-hours ownership plus tested manager and director escalation.
- Reclassify any lower-tier service capable of causing severe overnight harm rather than hiding the risk behind manager callback.
- Route unknown-owner incidents to Duty Command and Platform temporarily. Record each occurrence as a catalog control defect.
- Prohibit simultaneous primary assignments and two-person 24x7 rotations.
9. Implement compensation and fatigue protections (depends on: 4, 8)
End unpaid on-call before expanding mandatory coverage. Use fixed duty compensation so responders are not rewarded for alert volume.
- Use planning bands of $900–$1,200 per Tier 0/1 primary week and $300–$500 per secondary week.
- Use planning bands of $1,000–$1,300 per Duty Commander week and $400–$700 for Communications or scribe primary duty.
- Pay holiday premiums. Compensate all legally compensable active and waiting time for non-exempt staff, including overtime where required.
- Have HR, Finance, Payroll, and employment counsel approve final bands, tax handling, FLSA classification, New York wage-hour treatment, and schedule constraints within 14 days.
- Provide a protected recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal delivery commitments by about 15% during a primary week.
- Prohibit consecutive primary weeks, on-call during leave, hidden schedule swaps, and primary duty more often than one week in six.
- Allow responders to declare temporary fatigue-related unfitness without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Make active payroll setup, training, access, and readiness hard gates before any new mandatory night rotation starts.
10. Codify the live incident lifecycle (depends on: 6, 7)
Give every incident the same operational flow from first signal through verified recovery. The first objective is limiting customer and financial harm, not proving root cause.
- Use the states Detected, Declared, Triaged, Mitigating, Mitigated, Monitoring, Resolved, and Reviewed.
- Record impact start separately from detection. Use the earliest defensible evidence and revise it transparently when later facts emerge.
- Open with a standard command message: severity, known impact, assigned roles, immediate objective, workstreams, and next update time.
- Freeze unrelated production changes during SEV1 and normally during SEV2. Record every exception.
- Prefer reversible mitigation: rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, partner rerouting, or controlled processing suspension.
- Separate mitigation and diagnosis workstreams when staffing allows.
- Keep decisions in the shared incident record rather than direct messages.
- Require formal command and technical handoffs for shift changes, fatigue, or incidents exceeding four hours.
- For payment incidents, verify backlog handling, duplicate protection, customer state, settlement exposure, and ledger reconciliation before resolution.
- Require a severity-specific stability period and explicit handback to the owning team, Support, and Customer Success.
11. Establish one paging and incident system of record (depends on: 5, 7, 8)
Monitoring tools may remain specialized, but every human page and major-incident record must enter one controlled platform. This provides consistent routing and an audit trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page within two weeks.
- Ingest events from all six monitoring tools before disabling their direct human-paging routes.
- Route pages using the service catalog and deduplicate events belonging to the same symptom.
- Provide one declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, decisions, severity changes, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status publishing, and conference facilities.
- Apply MFA, role-based access, periodic access reviews, and tamper-evident history.
- Keep privileged legal or security material in restricted linked records rather than exposing it in the general timeline.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, the status provider, or an AWS region.
- Retire each legacy paging route only after ownership review, end-to-end testing, and two weeks of verified operation.
12. Enforce the alert-quality contract (depends on: 3, 5)
Treat every human page as a production interface with an owner and a required action. Measure human notification episodes rather than raw monitoring events.
- Require every paging rule to identify the service, owner, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Define an actionable page as one that causes or materially informs a timely intervention or risk decision.
- Define noise as duplicate, non-urgent, unactionable, stale, test-generated, or incorrectly routed notification.
- Page on payment outcomes, error-budget burn, queue age against deadlines, and financial-integrity risk rather than raw CPU, memory, pod, or log thresholds.
- Run new rules in shadow mode for seven days unless a documented emergency exception applies.
- Review repeated no-action pages within two business days.
- Set a page budget of no more than two after-hours notification episodes per responder per week, measured over four weeks.
- Make a sustained budget breach trigger a tuning sprint and block additional non-emergency paging rules.
- Require compensating detection and central approval before suppressing a Tier 0 or Tier 1 rule.
- Never disable an existing critical detector solely because its metadata or runbook is incomplete. Track the gap with a dated remediation owner.
13. Detect payment failures before customers (depends on: 5, 12)
Move detection from infrastructure health to customer journeys and ledger truth. Validate coverage against actual historical failures.
- Define SLIs and internal SLOs for initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, APIs, webhooks, and reporting freshness.
- Set internal objectives with enough headroom to protect the contractual 99.95% availability commitment.
- Run external synthetic transactions through critical journeys at least once per minute from paths independent of the production platform.
- Test each region and expose dependencies that defeat nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed payment value relative to settlement deadlines.
- Add tenant and cohort anomaly detection for high-value customers and critical payment methods.
- Convert credible Support, account-manager, processor, bank, and network reports into incident candidates within five minutes.
- Replay all 31 historical incidents. Record which current detector would fire and at what minute.
- Treat customer-first detection as a mandatory missed-detection review with a tracked action.
14. Implement the detection and escalation ladder (depends on: 7, 8, 11)
Create one time-bound path from the first credible signal to named command and the correct technical owner. Device delivery does not count as acknowledgement.
- Converge automated alerts, engineer observations, support cases, account-manager reports, partner notices, and customer calls on the same declaration path.
- Page the Duty Incident Commander and owning critical-domain primary immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at 5 minutes, manager at 10 minutes, and director at 15 minutes.
- Escalate unclaimed command to the backup commander at 5 minutes. The Executive Duty Officer assumes temporary command at 10 minutes until a certified transfer occurs.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Keep command with the current commander when service ownership remains unclear. Assign a temporary technical lead and record the ownership gap.
- Maintain tested external escalation routes for AWS, database support, processors, sponsor banks, networks, and critical vendors.
- Test the full declaration, acknowledgement, fallback, conference, and status-publishing path weekly.
- Treat missed acknowledgement, failed routing, and unowned incidents as control failures requiring review.
15. Standardize internal communications (depends on: 7, 10, 11)
Give responders one working room and stakeholders one controlled source of truth. Executives must not interrupt the technical command path.
- Maintain one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue the initial internal notice within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within the same initial-notice window.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record every material decision and outbound message in the incident timeline.
- Require the executive team to sign the communication behavior rules.
16. Standardize customer, account, and regulatory communications (depends on: 3, 6, 15)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate observed impact before root cause is known.
- Publish customer status within 15 minutes of a customer-visible SEV1 and within 30 minutes of a customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map public status components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, available workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Have account managers contact affected strategic accounts within 30 minutes for SEV1 and 60 minutes for SEV2 using the approved briefing.
- Offer status subscriptions to all customers. Auto-enroll only where contracts, consent, and applicable communication rules permit.
- Encode customer-specific notice deadlines and channels in the customer record.
- Have Legal and Compliance maintain a counsel-validated matrix covering applicable NYDFS, breach, GLBA or FTC, PCI, money-transmitter, sponsor-bank, network, insurance, and customer obligations.
- Complete and record a reportability assessment within one hour for every SEV1 and every security, privacy, or integrity-related SEV2, including decisions of not reportable.
- Let Legal own regulatory text and submission while the Incident Commander owns operational facts. Record any legally required restriction of public detail and its alternative stakeholder plan.
17. Tie incidents to SLA credits and financial exposure (depends on: 3, 16)
Link the incident record to contractual and financial outcomes. Finance should not discover outages through later credit claims.
- Define the authoritative availability calculation for each contract and customer journey with Legal and Finance.
- Calculate affected customers and minutes from incident scope and journey telemetry.
- Produce a preliminary credit and contractual-exposure estimate within five business days of resolution.
- Record failed payment count, failed or delayed value, settlement exposure, reconciliation breaks, support effort, and engineering effort.
- Establish a documented approval path for proactive credits and claims-based credits.
- Attribute credits and financial harm to recurring failure families.
- Use the quarterly credit analysis to prioritize detection, resilience, and architectural investment.
18. Establish mandatory blameless postmortems (depends on: 6, 7)
Use one learning standard with fixed deadlines. Keep learning separate from disciplinary and misconduct processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact lasting more than two hours, SLA credits, contractual breach, repeated contributing factors, major control failures, and ledger-integrity near misses.
- Produce a factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the owning engineering director accountable for completion. The commander owns response analysis, and the scribe supplies the timeline.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit answers to why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the commander or primary technical responder.
- Describe decisions using the context and information available at the time. Do not name an individual as the root cause.
- Keep HR, misconduct, and personnel matters in separate processes.
- Publish broadly useful findings internally while restricting security, privacy, personnel, and privileged content appropriately.
19. Make corrective actions enforceable risk commitments (depends on: 18)
An action is not complete when its ticket is closed. It is complete when the intended risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within 7 days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Prefer actions that remove hazards or reduce blast radius over vague actions such as retraining or adding monitoring.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of engineering capacity for approved reliability work.
- Escalate overdue high-risk actions to the manager after 7 days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit unrepaired severe conditions to block related releases.
- Verify effectiveness using tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days as complete, re-planned, superseded with evidence, or formally risk-accepted.
20. Set the service-readiness bar and payments playbooks (depends on: 5, 10, 12, 13)
A critical service must be supportable at 3 a.m. before it enters direct overnight coverage. Existing critical detection remains active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and integrity constraints.
- Require responders to demonstrate access and safe execution before independent primary duty.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional impairment, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document split-brain prevention, replay protection, failover, controlled backlog recovery, and post-recovery reconciliation.
- Require executive-approved RTO and RPO for the shared ledger cluster.
- Exercise critical runbooks at least twice a year and after material changes.
- Block new Tier 0 or Tier 1 releases and paging rules when readiness requirements are missing.
- Handle existing gaps using named owners, compensating controls, executive-approved expiry dates, and remediation plans.
- Run a parallel architecture workstream to reduce ledger blast radius, isolate non-critical readers, and strengthen regional independence.
21. Train and certify every response role (depends on: 7, 14, 15, 18, 20)
Command and communication are learned skills. Use paid working time and certify people before independent duty.
- Give all employees a 30-minute module on recognizing impact, declaring incidents, and locating the status page.
- Give responders a half-day module on severity, acknowledgement, escalation, evidence, runbooks, and financial-integrity precautions.
- Train scribes for two hours on timeline quality and fact-versus-hypothesis labeling.
- Train Incident Commanders for two days on delegation, uncertainty, severity, mitigation strategy, fatigue, handoffs, and executive management.
- Require commander candidates to complete a simulated SEV1 and two shadowed incidents or exercises.
- Train Communications Leads for one day on status writing, customer segmentation, legal boundaries, and contractual clocks.
- Require domain responders to demonstrate dashboards, access, rollback, failover, escalation, and relevant playbooks.
- Require two shadow shifts before independent primary duty.
- Renew certification annually through simulation.
- Maintain the training, assessment, and certification register as operational and audit evidence.
- Nominate an incident-management champion in each of the 28 teams.
22. Exercise command, recovery, and tool failure (depends on: 11, 20, 21)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first cross-company command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous operational and security events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors when relevant.
- Conduct at least one unannounced after-hours paging test before the audit and two annually thereafter.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure acknowledgement, command, customer notice, mitigation decision, handoff, and recovery times.
- Create tracked corrective actions for every material exercise finding.
23. Publish the signed policy and control set (depends on: 6, 7, 8, 9, 10, 12, 14, 15, 16, 18, 19)
Convert the design into concise documents that people can use during an incident. The actual operating process must also be the documented and audited process.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, evidence, and exception standards.
- Include one-page cards for severity, roles, authority, escalation, and communication timings inside the incident tool.
- State explicitly that responders support only owned or formally accepted and trained service portfolios.
- Include the compensation structure, fatigue rules, declaration rights, and non-retaliation commitment.
- Obtain approval from the CTO, HR, Legal, Security, Compliance, and Internal Audit.
- Announce the policy at an all-hands and through team briefings.
- Create an exception register with owner, rationale, compensating control, approver, review date, and expiry.
- Version every policy change. Do not rewrite historical records when the process changes.
24. Instrument the scorecard and review forums (depends on: 3, 11, 18, 19)
Measure the process before the pilot so failures become visible immediately. Report medians and 90th percentiles rather than averages alone.
- Measure impact-to-detection, detection-to-declaration, declaration-to-command, acknowledgement, mitigation, recovery, and resolution.
- Split results by severity, service tier, customer journey, region, detection source, and business-hours status.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track notification episodes, actionability, duplicates, after-hours load, missed detection, routing errors, and page-budget breaches.
- Track postmortem timeliness, action age, due-date performance, verified effectiveness, and repeated contributing factors.
- Track journey availability, error-budget burn, failed or delayed value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight work, recovery days, exceptions, sentiment, and responder attrition.
- Hold a weekly Incident Review Board chaired by the Head of Reliability with relevant directors.
- Hold a monthly executive reliability review and a quarterly control and resilience review with Internal Audit.
- Reconcile incident records monthly against support cases, customer complaints, status history, credits, and major operational anomalies to detect under-reporting.
- Use team-level scorecards to direct help and investment. Never penalize an individual for good-faith declaration.
25. Pilot the complete process on the payment path (depends on: 9, 11, 13, 20, 21, 23, 24)
Run a four-to-six-week pilot across the highest-risk journey before expanding. The interim operating floor remains active for the rest of the company.
- Include payment orchestration, ledger application, PostgreSQL platform, API edge, authentication, settlement, reconciliation, Kubernetes platform, and Support intake.
- Include teams with existing on-call experience and teams new to the model.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking together.
- Run legacy and new paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently taking command.
- Correct critical process or tooling defects within 48 hours.
- Exit only after 95% timely command assignment, 95% communications compliance, no unpaid pages, complete required postmortems, tested fallbacks, and at least 50% lower pilot noise.
- Publish the pilot results, defects, and policy changes company-wide.
26. Roll out by risk with readiness gates (depends on: 25)
Expand in controlled waves and finish early enough to accumulate operating evidence before the audit. A calendar date does not override a failed readiness gate.
- Roll out remaining Tier 0 domains first, followed by Tier 1, Tier 2, and Tier 3.
- Use four waves of six to eight teams, each lasting two to three weeks.
- Gate each service on catalog ownership, appropriate coverage, active compensation, trained responders, access, tested escalation, alert quality, runbooks, and a passed tabletop.
- Require at least six responders only for direct 24x7 technical rotations. Use business-hours coverage for lower tiers.
- Give each wave a named coach and director sign-off.
- Reschedule failed gates or use a time-limited executive exception with compensating controls. Do not create silent waivers.
- Disable legacy human-paging routes after verified cutover for each wave.
- Run a quota-based noise reduction sprint in every wave, starting with the highest-volume rules.
- Pair every suppression with a compensating-detection check.
- Publish an internal adoption dashboard by team, service tier, coverage, and control gap.
- Complete critical coverage by approximately week 12 and all 28 teams by week 18.
27. Prove SOC 2 operating effectiveness (depends on: 22, 23, 24, 26)
Generate evidence through normal operation rather than reconstructing it before fieldwork. Test both control design and consistent execution.
- Map controls to the applicable Trust Services Criteria with Compliance and the auditor, including monitoring, incident identification, response, recovery, communications, and availability.
- Retain approved policies, exceptions, service ownership, schedules, compensation activation, access reviews, training, incidents, communications, reportability decisions, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, missed timelines, non-reportable decisions, and exercises in the testing population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the evidence populations and interviews expected from the external auditor.
- Correct deviations through tracked actions with owners and dates. Never edit history to create apparent compliance.
- Verify that evidence retention covers the full auditor-defined observation period.
- Brief commanders, engineers, Support, and Compliance on the actual process without scripting inaccurate answers.
28. Inspect, adapt, and institutionalize ownership (depends on: 26, 27)
Prevent the process from decaying after rollout or the audit. Change it using measured operating evidence rather than opinion.
- Review the policy after 90 days of live operation using severity calibration, page load, missed detection, communication compliance, action closure, fatigue, and survey results.
- Remove steps that create work without reducing risk. Add controls only where incidents, exercises, or evidence show a gap.
- Reassess Tier 0 and Tier 1 classification and domain boundaries every six months.
- Assign permanent owners for policy, catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review compensation bands, rotation burden, accommodations, and staffing annually.
- Report severe incidents, credits, overdue high-risk actions, and ledger concentration risk to the board or risk committee quarterly.
- Maintain the ledger blast-radius program as an executive risk until failover, degraded mode, reconciliation, and regional independence meet approved objectives.
- Evaluate follow-the-sun coverage using one year of actual activation and staffing data.
- Build year-two plans for automated mitigation, safer deployments, graceful degradation, and error-budget release controls.
--- PROPOSAL 3 ---
Proposal ID: f61aca40-a8a9-4e91-a5f3-ac1d03634441
Content:
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of 'not reportable'.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (29):
1. Charter the program, fund it, and start the audit clock
Convert the CEO email into a company operating process with one accountable owner, a budget, and a dated timeline that starts this week.
- Name the CTO as executive sponsor and appoint a Head of Reliability & Incident Management as the single accountable owner with full-time authority over all 28 teams.
- Stand up a three-person program office: program lead, platform engineer, reliability analyst.
- Form an eight-person steering group spanning Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, and HR. It proposes; the sponsor decides within 48 hours. Never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one human-paging path, one incident record, one postmortem format, paid on-call, mandatory action tracking, and named service ownership.
- Approve budget anchored against the $1.3M in SLA credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs, and 15% reserved engineering capacity.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Publish a one-page charter company-wide on day 2. Incident response is a company process, not a per-team preference.
- Timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day interim command floor (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place this week so the next outage has a named commander and the evidence clock starts immediately.
- Publish a one-page interim severity card and one declaration path: a Slack command, a phone number, and the existing pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs from the 12 teams already on-call. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical action items into complete, re-plan, or formally risk-accept within 30 days, prioritising ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline of incidents, alerts, and money lost (depends on: 1)
Rebuild the facts before locking any design. This is both the design input and the frozen 'before' picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines in one signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the written fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written, signed deal.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- State explicitly that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Map compliance, evidence, and audit requirements from day one (depends on: 1)
Design evidence as a by-product of operations, not a reconstruction before the auditor arrives. The interim process in week 1 is already evidence.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for internal and external communication; CC5 for control activities; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure with verification, training register, drill records, reportability decisions including 'not reportable'.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records. Require role-based access, MFA, and periodic access review.
- Version and approve all policy documents from day one: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard, Alert Quality Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has a named owner. A wrong owner recreates the pager objection. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Assign one accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. **Tier 1**: customer-facing but degradable. **Tier 2**: internal or batch. **Tier 3**: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for every Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform. They fail differently and need different hands.
- Orphan services get an owner within 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Missing ownership or runbooks blocks Tier 0/1 releases.
7. Adopt the severity scale, declaration rights, and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix. Anyone may declare. Nobody is penalised for over-declaring.
- **SEV1 (crisis):** money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; deploy freeze; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads; business-hours comms; postmortem if customer-detected, over two hours, a repeat, or credit-generating.
- **SEV4:** no customer impact. Ticket only. Never pages.
- Payments-specific anchors: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident become at least SEV2.
- Only the Incident Commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
8. Define incident roles, authority, dual-control, and handover discipline (depends on: 7)
Solve 'nobody in charge for an hour' by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Command claimed within 5 minutes and stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Staff 24x7 with a three-layer coverage model, not 28 night rotations (depends on: 6, 8)
Do not create 28 night rotations. That is precisely what engineers are rejecting. Centralise coordination in a trained command corps and keep technical ownership local.
- **Layer A — Incident Command corps:** approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. Paired with a Communications Lead pool of approximately 18 from Support, CS, and engineering management, and a scribe pool used as the training entry point.
- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat a Lisbon or APAC follow-the-sun cell as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
10. Approve paid on-call, New York labor compliance, and fatigue safeguards (depends on: 4, 9)
Unpaid on-call in New York is both a retention problem and a wage-hour exposure. Pay must be live in payroll before any new mandatory night rotation starts. Publish the actual numbers, then ask people to sign up.
- Indicative scheme locked by HR, Finance, and employment counsel within 14 days: approximately $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as approximately 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
- Publish on-call load by team quarterly.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
11. Set the alert-quality standard and a hard page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target: 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, noise below 15%, with no loss of Tier 0/1 detection coverage.
12. Detect payment and ledger failures before customers do (depends on: 6, 11)
The goal is blunt: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
- Run the replay test: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish replay coverage as a leading metric.
13. Consolidate to one paging platform, one incident record, one status page (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
14. Codify the five-minute escalation path and live execution doctrine (depends on: 8, 9, 13)
Write one unskippable path from 'something looks wrong' to 'someone is in charge'. The default action is never waiting. If nobody claims command within 5 minutes, the platform assigns it and announces it.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute.
- Separate diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Closure requires a stability observation window and explicit handback.
15. Set the service-readiness bar and write major-incident playbooks (depends on: 6, 9, 12)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human at night. The shared ledger cluster is the single largest structural risk.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction: partitioning, read replicas, isolation of non-critical readers, with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
16. Run one communications clock for internals, customers, and regulators (depends on: 7, 8, 13)
Replace 'whoever is around' with one timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure. State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is 'not reportable'.
- Obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
17. Tie incidents to SLA credits and true financial cost (depends on: 7, 16)
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
18. Make blameless postmortems mandatory with one format and fixed deadlines (depends on: 7, 8)
Replace 'some incidents, various formats' with one mandatory format, fixed deadlines, and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did?
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and are never the commander of the incident under review.
- Publish a searchable library and a quarterly top-five recurring causes analysis, with restricted versions for security or privileged content.
19. Enforce action ownership, reserved capacity, and tracking (depends on: 13, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity. Closing a ticket without evidence does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
- Target: 90% of high-priority actions closed by due date within two quarters.
20. Train and certify every role before independent duty (depends on: 8, 14, 16, 18)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it. The certification register is an audit artefact.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days plus shadowing): command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
21. Rehearse with tabletops, game days, and unannounced drills (depends on: 13, 15, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-mitigation.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
22. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 14, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 10, 12, 13, 15, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse.
- Publish a one-page result to the whole company.
24. Instrument metrics, dashboards, review cadence, and anti-gaming (depends on: 13, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
25. Roll out in risk-ordered waves with readiness gates (depends on: 23, 24)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, never waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave. Give every team its ranked list of noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Review missed detections monthly with the same seriousness as noise.
- Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
26. Run change management, fairness, and pager culture from day one (depends on: 4, 10)
Run this in parallel from day one. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
27. Produce SOC 2 evidence by operating, test internally, and mock-audit (depends on: 22, 24, 25)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence. Documented exceptions beat a claim of perfection.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5, CC2.2/CC2.3, CC5, and A1.2. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including 'not reportable', postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in months 4 and 6, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
28. Maintain the program risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and pre-commit the response. Review monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
29. Inspect at 90 days, lock year-two ownership, and prevent decay (depends on: 24, 25, 27)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 4 ---
Proposal ID: 91ea9817-3664-46dd-a758-888405cea636
Content:
Estimated Complexity: high
Success Metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- By day 7, every suspected major incident uses one record, one channel, and a named commander; Support can declare without engineering confirmation.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of customer-visible SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (26):
1. Charter the program, fund it, and start the audit clock
Turn the CEO email into a chartered company program within 48 hours. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability as full-time owner, with a three-person office: program lead, platform engineer, and analyst.
- Form a small decision group of Engineering, SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit. It proposes. The sponsor decides within 48 hours.
- Lock the non-negotiables now: one severity scale, one human-paging platform, one incident record, one postmortem format, mandatory action tracking, paid on-call, and named service ownership.
- Confirm the SOC 2 Type II observation window with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, exercises, and reserved engineering capacity against the $1.3M in SLA credits.
- Reserve 15% of engineering capacity for detection, runbooks, and incident actions, protected by the sponsor.
- Publish a one-page charter on day 2. Clock: floor by day 7, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for policy, tooling, or the audit. Put a crude but real process in place this week so the next outage already has an owner.
- Publish a one-page interim severity card and one declaration path: chat command, phone number, and current pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline, and one naming convention for every suspected major incident.
- Triage the 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
- Replay one of the two nobody-in-charge incidents as a tabletop within 14 days using this floor process.
3. Rebuild the forensic baseline and incident-replay catalog (depends on: 1)
Rebuild the facts before locking design. This is the design input, the frozen before-picture for the CEO, and the test set for detection work.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact missing signal. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Name the top 50 noisy rules and every rule with no owner or runbook.
- Quantify true cost: credits, failed payment volume, delayed value, reconciliation breaks, and engineering hours lost.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11 of 64 actions closed, on-call in 12 of 28 teams.
- Build a **replay catalog**: for each historical incident, the detector that should now fire, at which minute, and the owner of the gap.
4. Publish the on-call fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders. Do not punish them by making them the company's permanent night watch.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the **fairness contract** in writing: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive exemption path for health, disability, or caregiving.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and journey tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has one owner. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Record owning team, manager, chat channel, escalation policy, dashboards, runbook, dependencies, regions, and data stores for all 180 services.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, refund, report, onboard — to services, data stores, regions, sponsor banks, and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform. They fail differently.
- Orphan services get an owner in 30 days or an approved decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Coverage follows the journey, not the org chart. Small teams that own Tier 0 pieces get headcount, service reassignment, or membership in a domain rotation. Never a two-person 24x7 rota.
6. Lock severity levels, integrity flags, and the incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Classify on actual or credible customer, financial, security, and contractual harm, never on who reported it or how hard the fix looks.
- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 contained: narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages a human.
- Attach an **integrity, security, settlement, or regulatory flag** to any severity. The flag forces dual-control, Legal, and the reportability checkpoint without inventing a fifth level. A one-customer ledger corruption is still a flagged crisis.
- Auto-escalate to at least SEV2: any ledger-cluster event, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident.
- Anyone may declare. Nobody is punished for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- Communications Lead: single voice for the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists. A human validates at SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory text. The commander owns the facts.
- Distinct people for command, comms, and technical lead at SEV1. Comms and scribe may combine only for bounded SEV2.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
- Every role assignment and handover is announced verbally and in writing with the exact time and open risks.
8. Staff 24x7 with command, domains, and a Duty Triage Desk (depends on: 4, 5, 7)
Do not create 28 night rotations. That is what engineers are rejecting. Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified people from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to eight months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate ownership into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- **Duty Triage Desk**: a small paid overnight first-line rotation that owns the first ten minutes of ambiguous, unowned, or low-confidence pages. It verifies, enriches, and applies only the runbook's safe steps, then wakes the owning domain. It never sits on a suspected ledger, payment-halt, or security event. Those page commander and likely domains immediately, in parallel.
- Overnight comms: SEV1 pages a Communications Lead 24x7. Customer-visible SEV2 lets the commander publish the first status from a template; Comms is paged if the incident is still open at 30 minutes or a top-100 account is affected.
- Staffing math: 260 engineers can sustain 10–12 domain rotations, one command corps, and one triage desk. They cannot sustain 28 night rotas. Role exclusivity: nobody is primary on two rotations in the same week. Commanders may also be domain responders in different weeks.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
- Evaluate follow-the-sun as a 12-month option, not a year-one dependency. Seed Layer B from the 12 teams already on-call.
9. Pay on-call, meet New York labor rules, and cap fatigue (depends on: 4, 8)
Unpaid on-call in New York is a retention problem and a wage-hour exposure. Pay must be in payroll before any new mandatory night rotation starts.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, $400–$800 for Comms or triage-desk duty, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- **Hard gate: no mandatory night rotation starts before compensation is live in payroll.** Interim duty is paid retroactively. Budget roughly $0.9M–$1.2M a year, then refine with actual rotation count.
10. Enforce an alert-quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers, then replay the past (depends on: 5, 10)
Stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. A detector is not done until it would have caught the last 12 months.
- Define SLIs and SLOs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner, processor, and sponsor-bank notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a named defect class with a mandatory tracked action.
- Run the **replay test** against the S3 catalog. For each of the 31 incidents, name the detector that would now fire and at which minute. Close gaps the replay exposes before calling detection improved.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One chat command creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy. Unowned pages go to the Duty Triage Desk and Duty Command, and log a catalog defect.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Codify escalation, the five-minute command rule, and vendor incidents (depends on: 7, 8, 12)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting. Payments also fail at processors and banks, which you cannot patch.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Vendor-incident class: processor, sponsor-bank, card-network, or cloud-control-plane failure. Command is still required. Success is time-to-customer-notice, time-to-failover-decision, and queue management, not root cause at the vendor.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Write live execution doctrine, readiness bars, and ledger playbooks (depends on: 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving root cause. A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis. Split diagnosis and mitigation workstreams once staffing allows.
- Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Drain backlogs under control. Complete reconciliation before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date. Existing detection stays on while gaps are repaired.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers, with board-visible milestones and executive-signed RPO/RTO.
15. Run one communications clock for internals, customers, and account managers (depends on: 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for customer-visible SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys, not internal service names. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer record.
- Language rules: state impact and the next update time. **Never speculate** on cause, recovery time, data integrity, or blame.
16. Operationalize regulatory notice, partner clocks, and SLA credits (depends on: 6, 15)
In payments, some incidents start a legal clock at detection. Tie incidents to money so Finance does not learn about outages from invoices.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when not reportable**, with facts, decision-maker, and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families. Track failed payment volume, delayed value, and reconciliation breaks as the true incident cost.
17. Make postmortems mandatory and actions enforceable (depends on: 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth. Eleven of 64 closed is a process nobody enforces.
- Mandatory for every SEV1 and SEV2, any incident a customer detected first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and gap, timeline, response analysis, contributing conditions, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items. Facilitators are trained and are never the commander of the incident under review.
- Every action gets one named owner, priority, due date, expected risk reduction, verification method, and a ticket created automatically. If it is not on the board, it does not exist.
- Classes: containment in 7 days, corrective in 30, strategic in 90. SEV1 recurrence-prevention enters the next sprint ahead of roadmap work. Reserve 15% capacity.
- Overdue ladder: manager at +7 days, director at +14, CTO at +30. Overdue high-risk items need written residual-risk acceptance and can block related releases. Verify effectiveness before closing.
18. Train and certify every response role before independent duty (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Duty Triage Desk, half day: enrichment, safe-step limits, when to wake a domain immediately, when not to delay.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
19. Publish Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17)
Collapse the design into a document people will actually open mid-outage, and make it official. Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
20. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 18, 19)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, Duty Triage Desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
21. Instrument metrics, reviews, and anti-gaming (depends on: 12, 17, 20)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- Response: time from first impact to detect, declare, commander, acknowledge, mitigate, and resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
22. Rehearse with tabletops, game days, and night drills (depends on: 12, 14, 18, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced overnight paging drill to measure real acknowledgement times.
- One combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO, before the audit.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
23. Burn down alert noise and roll out by risk with gates (depends on: 20, 21)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Noise reduction is a quota inside each wave, not a background hope.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy paging paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Give every team its noisiest rules from the baseline. Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed. Pair every suppression with a compensating-detection check.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
- If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
24. Operate the fairness and culture program in parallel (depends on: 4, 9)
Run this from the moment the deal is published. Engineers judge the process on fairness. Executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
25. Prove SOC 2 operating effectiveness before fieldwork (depends on: 19, 21, 22, 23)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (depends on: 21, 23, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, metrics, and the ledger blast-radius workstream, independent of the audit cycle.
- Pre-committed contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 5 ---
Proposal ID: b427ef6b-42c2-41ac-8f9d-992d99ba5e20
Content:
Estimated Complexity: high
Success Metrics: - Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (30):
1. Charter the program and start the audit clock
Turn the CEO email into a chartered company program with one accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
- Confirm the SOC 2 Type II observation window with the auditor; map the future response controls to applicable Trust Services Criteria; define evidence set and retention requirements.
4. Listening tour, resistance map and the written on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (depends on: 5, 7)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 4, 8)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (depends on: 5, 10)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (depends on: 7, 13)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 12, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
- Agree with Legal and Finance the availability measurement method per contract and per component; compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide credit posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain; attribute credits to root-cause families for investment decisions.
18. Blameless postmortem standard and Incident Review Board (depends on: 6, 7, 12)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
19. Action ownership, reserved capacity and enforcement (depends on: 12, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
20. Training and certification academy (depends on: 7, 13, 15, 18)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
21. Exercise programme: tabletops, game days and unannounced drills (depends on: 12, 14, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
22. Publish Incident Management Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
24. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
25. Alert noise burn-down campaign (depends on: 10, 12, 23)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
26. Wave rollout to all 28 teams with readiness gates (depends on: 23, 24, 25)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
27. Change management, fairness and pager culture (depends on: 4, 9, 23)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
28. SOC 2 evidence by design, internal testing and mock audit (depends on: 22, 24, 26)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
29. Risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
30. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 24, 26, 28)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Instructions:
1. Carefully evaluate each proposal considering:
- Clarity and detail of the proposed steps.
- Overall alignment with the Main Objective.
- Quality and relevance of the proposed success metrics.
- Coherence of dependencies between steps (if any).
2. Think deeply about each proposal, analyzing its strengths and weaknesses. Consider how well each proposal addresses the main objective and which one would be most effective.
After your thorough analysis, give the Proposal ID of the selected proposal and a one- or two-sentence justification of your choice.
Answer with a single JSON object, and nothing else, that conforms to this JSON schema:
{"description": "What a voting agent must answer with.", "properties": {"selected_proposal_id": {"title": "Selected Proposal Id", "type": "string", "enum": ["30f15b26-0ec0-4756-9f4a-6279eb889884", "1f1527bb-d992-451c-bbc4-9e61ffdb1f30", "f61aca40-a8a9-4e91-a5f3-ac1d03634441", "91ea9817-3664-46dd-a758-888405cea636", "b427ef6b-42c2-41ac-8f9d-992d99ba5e20"]}, "justification": {"title": "Justification", "type": "string"}}, "required": ["selected_proposal_id", "justification"], "title": "VoteAnswer", "type": "object", "additionalProperties": false}
{"selected_proposal_id": "30f15b26-0ec0-4756-9f4a-6279eb889884", "justification": "Proposal 1 is the most comprehensive and detailed, explicitly defining every required element—severity, roles, 24x7 coverage, paid on-call, alert quality, detection/replay, communications, postmortems, metrics, and SOC 2 readiness—with clear dependencies and measurable success metrics. It uniquely adds high-value components such as change intelligence, SLA credit guardrails, and a Duty Triage Desk, which directly target the stated pain points."} |
Selected plan
1. Charter the programme: one owner, one mandate, funded, dated against the audit
Turn the CEO email into a chartered company programme with a single accountable owner and authority over all 28 teams. Incident response stops being a per-team preference and becomes a company operating process.
- Name the CTO as executive sponsor and a full-time Head of Reliability & Incident Management as accountable owner, supported by a programme office of three: programme lead, incident-platform engineer, reliability analyst.
- Form a decision group (Engineering, Platform/SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables on day one: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, named service ownership, and paid on-call.
- Publish the timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal control test month 5, mock audit month 7, SOC 2 fieldwork month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.3M/yr, training and exercises, plus 15% of engineering capacity reserved for reliability work.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (after 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence immediately.
- Publish a one-page interim severity card and a single declaration path: one chat command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim 24x7 Duty Incident Commander roster (primary plus backup) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document and one naming convention for every major incident, starting now.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (after 1)
Rebuild the facts before designing anything. This is the design input, the executive narrative and the frozen "before" picture for the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Correlate incidents with deployments, config changes and feature flags to quantify how many started with a change we made.
- Quantify true cost beyond credits: failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost, churn risk on named accounts.
- Freeze the baselines in a signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Audit, legal and evidence scoping in week one (after 1)
Design evidence as a by-product of operating, never as a reconstruction before fieldwork. Settle the scope and the legal handling of incident records now, not in month seven.
- Confirm with the external auditor the Type II observation window, the incident population definition, and the evidence they will sample. Everything from the day-7 floor onwards must count.
- Map incident response to the Trust Services Criteria with Compliance: CC7.2–CC7.5 (monitoring, identification, response, recovery), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), A1.2 (availability).
- Define the evidence set and where it is produced automatically: incident records, paging and acknowledgement logs, role assignments, status-page history, reportability decisions, postmortems, action closure with verification, training register, drill records, access reviews.
- Agree retention, confidentiality, legal hold and access rules. Decide with counsel which postmortem content is privileged and how privileged material is segregated without making the ordinary postmortem secret.
- Start the obligation matrix with Legal: NYDFS 23 NYCRR 500, state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network contractual windows, cyber-insurer notice.
- Begin monthly evidence sampling from month 1 so control operation is visible long before the audit.
5. Listening tour, resistance map and the written on-call deal (after 1, 3)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written deal signed by the sponsor.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
- State the hard gate publicly: no new mandatory night rotation starts before compensation, training, runbooks and staffing rules are live.
6. Service catalog, ownership, tiering and customer-journey map (after 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- Record for every service: one accountable team, engineering manager, chat channel, escalation policy, dashboards, runbook link, dependency list, regions and data stores.
- Tier by business impact, not technology. Tier 0: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record SLO, RTO, RPO, active-active or region-bound, failover method, and the dependencies that make nominal two-region redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
7. Severity standard, declaration rights and incident lifecycle (after 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- SEV1 (crisis): money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning responders and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment started within 1 hour; mandatory postmortem.
- SEV2 (critical): material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and responders paged, comms and scribe if customer-visible, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- SEV3 (contained): narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, credit-generating or a repeat.
- SEV4: no customer impact. Ticket only; never pages a human.
- Attach a financial-integrity flag or a security flag to any severity. The flag forces dual control, Legal engagement and the reportability checkpoint without inventing a fifth level.
- Anchor on payments reality alongside error rates: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring. Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including a customer report.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
8. Roles, authority, concurrency and handover discipline (after 7)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- Incident Commander: owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend. Retains command when a VP joins.
- Communications Lead: the single voice to the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1 and SEV2.
- Subject-Matter Responders: engineers of the owning team. They mitigate; they do not run the room.
- Executive Duty Officer (SEV1): removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, Finance and Vendor Management join on defined triggers rather than by invitation.
- Rules: command claimed within 5 minutes and stated in channel; distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Concurrency doctrine: two simultaneous SEV1/SEV2 incidents activate the secondary commander and a surge roster; a designated Multi-Incident Coordinator arbitrates shared resources such as the ledger, the database platform and the deploy freeze.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation or privileged-access rules.
9. Three-layer 24x7 coverage: command corps, domain rotations, overnight triage desk (after 5, 6, 8)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- Layer A — Incident Command corps: ~30 certified commanders drawn from across the 28 teams plus managers, giving roughly one primary week per person every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- Layer B — domain responder rotations: consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement and reconciliation, auth, API edge, Kubernetes platform, data and reporting, partner integrations). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- Layer C — everyone else: business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- Duty Triage Desk: a small paid overnight first-line rotation owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Publish the staffing arithmetic: 10–12 domain rotations of six to eight people plus a 30-person command corps means roughly 100 of 260 engineers carry a night obligation, at about one week in six. Twenty-eight independent rotations would be unstaffable and is therefore rejected on the numbers.
- Teams too small for a fair rotation get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: a human acknowledgement within 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
10. Paid on-call, New York labour compliance and fatigue safeguards (after 5, 9)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour, and a premium for the overnight triage desk.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per responder per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Hard gate: no mandatory night rotation starts before its compensation is live in payroll. Interim duty since day 7 is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons without career penalty, and on-call load is published quarterly by team.
11. Alert quality contract and page budget (after 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a page budget of two out-of-hours pages per responder per week. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression can never hide an incident.
12. Detection uplift on the money path, validated by incident replay (after 6, 11)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Run the replay test: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action and a missed-detection review.
13. Change intelligence and deployment safety on the critical path (after 6, 12)
Most of these incidents start with something we changed. Make change the first hypothesis the tooling answers, and make changes safer to reverse.
- Stream every deployment, configuration change, feature-flag flip, schema migration and infrastructure change into the incident timeline with service labels and owners.
- Give the commander an automatic "what changed in the last 60 minutes on the affected journey" panel at declaration time.
- Require Tier 0/1 changes to be progressively delivered with a documented rollback that is tested, time-bounded and executable by the on-call responder without the author.
- Treat ledger schema migrations and settlement-affecting changes as a separate class: dual approval, rehearsed rollback, no deploys inside the settlement window.
- Enforce a change freeze during SEV1 and SEV2, lifted only by the commander and logged.
- Report change-correlated incidents monthly; a rising ratio is a signal to strengthen release safety, not to blame a team.
14. One pager, one incident record, one status page — migrated without a detection gap (after 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy paging path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access, MFA and periodic access review.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page delivered outside this platform creates no on-call obligation.
15. Escalation ladder and the five-minute command rule (after 8, 9, 14)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts and support entitlements: AWS and database premium support, processors, sponsor banks, card networks.
16. Live execution doctrine and payments safety rules (after 8, 14, 15)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- Payments guards: protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed in advance rather than improvised.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and automatic reopening if impact recurs.
17. Service readiness bar, major-incident playbooks and ledger blast-radius reduction (after 6, 9, 16)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human, and the shared ledger is the largest structural risk in the estate.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as a concentration risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers, stronger regional independence — with board-visible quarterly milestones.
- Enforcement: no readiness sign-off, no night paging alerts unless the manager accepts the gap in writing with an expiry date and a compensating control. Never respond to a gap by turning detection off.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
18. Internal communications protocol (after 8, 14)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, Sales and Security.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change. First internal brief within 10 minutes of SEV1 and 15 of SEV2.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
19. Customer communications, status page and account-manager outreach (after 7, 18) from P3 · round 1 step 14
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; monitoring notice after mitigation; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, monitoring, resolution.
- Component-level status page mapped to customer journeys rather than internal service names, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1 and 60 of SEV2, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state observed impact, affected capabilities, any workaround and the next update time; never speculate on cause, recovery time, data integrity or blame.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
- Record any legally required restriction or delay of public detail, its approver, and the alternative stakeholder plan.
20. Support and account management as a detection and intake tier (after 7, 12, 19)
Customers detected 40% of incidents first, which means the front line already holds the signal. Turn Support and account managers into an instrumented detection channel rather than a bystander.
- Give Support explicit declaration rights, a one-page trigger card, and a macro that opens an incident candidate directly in the incident platform.
- Automate the clustering rule: five similar tickets or calls in ten minutes auto-creates a triage incident assigned to the duty commander.
- Route credible partner, processor and sponsor-bank notifications into the same declaration path within five minutes.
- Generate the affected-customer list automatically from journey telemetry and the incident record, and push it to Support, CS and the account-manager briefing.
- Train Support and account managers on approved language and prohibit independent technical explanations to customers.
- Measure and publish "signal was in Support before it was in monitoring" as a detection defect, and feed each instance into the detection backlog.
21. Regulatory, partner and legal notification playbook (after 4, 7, 19)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Complete the obligation matrix started in scoping and have counsel validate the triggers, deadlines, channels and submitting authority for each obligation.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded even when the answer is "not reportable", with facts considered, decision-maker, timestamp and reassessment trigger.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, card networks, outside counsel, cyber-insurer, critical vendors.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the operational facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly inside the exercise programme, including one full regulator-notification simulation per year.
22. SLA credit workflow, true cost model and incentive guardrails (after 7, 19)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component, including how partial degradation counts.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Guardrail against the perverse incentive: better detection will surface incidents that previously went unbilled, so credits may rise before they fall. Publish this expectation to the executive team in advance, and make it a written rule that no finance or commercial pressure may influence severity or declaration.
23. Blameless postmortem standard and Incident Review Board (after 4, 7, 8)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger; any control failure.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning engineering director is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did?
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions where security, privacy or privileged content requires it.
24. Action ownership, reserved capacity and enforcement (after 14, 23)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, accountable manager, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness through tests, telemetry, exercises or production evidence before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
25. Training and certification academy (after 8, 15, 18, 23)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days plus shadowing): command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, the 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Comms Lead (1 day): status writing, customer tiering, contractual clocks, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days of joining a domain.
- Certification is valid 12 months and renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
26. Exercise programme: tabletops, game days, night drills and vendor rehearsals (after 14, 17, 25)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander. First company-wide tabletop within 30 days.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced overnight paging drill to measure real acknowledgement times, plus a concurrency drill with two simultaneous incidents.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down. Rehearse joint escalation with AWS, a processor and a sponsor bank.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
27. Publish Incident Management Policy v1 and the exception register (after 7, 8, 9, 10, 11, 15, 18, 19, 21, 23, 24) from P4 · round 0 step 14
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, escalation and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix, Readiness Bar.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
28. Pilot on the payments critical path (after 10, 12, 13, 14, 17, 25, 27) from P4 · round 0 step 22
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement and reconciliation, PostgreSQL platform, Kubernetes platform, API edge, auth and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, action tracking, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The programme lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time in 95% of cases, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
29. Alert noise burn-down campaign (after 11, 14, 28) from P4 · round 0 step 23
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → under 1,500 pages a month by day 90, under 500 with noise below 15% by month 6.
30. Metrics, dashboards, review cadence and anti-gaming (after 14, 23, 28)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- Response: time from first impact to detect, declare, commander, acknowledge, mitigate, resolve — split by severity, tier, journey, region and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors, change-correlated incident ratio.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
31. Wave rollout to all 28 teams with readiness gates (after 28, 30)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the programme office for three weeks; the director signs the gate.
- A failed gate is rescheduled, never waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard by team, tier and control gap.
32. Change management, fairness and pager culture (after 5, 10, 28)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
33. SOC 2 evidence by design, internal control test and mock audit (after 4, 27, 30, 31)
The real process is the evidence. Never build a parallel audit process, and never reconstruct records after the fact.
- Maintain the evidence set defined in scoping, produced automatically and indexed: versioned policies and exceptions, catalog records, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Sample incidents monthly from first signal through verified corrective action, deliberately including customer-reported events, downgraded incidents and missed timelines.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling end to end and reporting gaps to the sponsor.
- Run a formal mock audit in month 7 using the populations, evidence requests and interviews the auditor will use: a commander, a random engineer, Support, Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
34. Programme risk register and pre-committed contingencies (after 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps and immediate staffing intervention.
- A major SEV1 mid-rollout: the programme lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
35. Ninety-day inspect-and-adapt, then year-two sustainability (after 30, 31, 33)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, exercise calendar and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to leading indicators: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and quarterly board reporting as permanent commitments owned by the Head of Reliability.
- A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90 a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95% internally and externally.
- Top-100 account outreach completed within 30 minutes of SEV1 in 95% of qualifying cases, using the approved briefing pack.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- Change-correlated incidents are identified within 10 minutes of declaration in 90% of cases, and the change-correlated share of incidents declines quarter over quarter.
- SLA credits fall to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, one concurrency drill and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and a tested read-only mode by month 6, with milestones reported to the board quarterly.
Evolution analysis
The analysis in brief
- Outcome: P1 wins, and the vote agrees with my ranking. P1 is the most complete plan (35 steps) and uniquely carries change-intelligence/deploy-safety (step 13) against the 3 h 10 min MTTM, Support/account managers as an instrumented detection tier (step 20) against 40% customer-first detection, and a concurrency doctrine with a Multi-Incident Coordinator (step 8); P2 is the disciplined runner-up (defines "actionable page" and "noise" in step 12, tests controls twice in months 4 and 6, company-wide by week 18).
- The final plans are clearly better than any round-0 draft, gaining a day-7 operating floor, staffing arithmetic (~30-person command corps + 10–12 domain rotations + business-hours Layer C instead of "28 teams on 24x7"), mitigate-vs-resolve and reconciliation semantics for the ledger, the 31-incident replay test as proof detection improved, shadow-mode/never-silently-disable rails on noise cutting, and dollar pay bands with a payroll hard gate.
- Convergence was genuine in round 0, imitation afterwards. All five independently found the same skeleton, then P3 imported 21 of 32 steps from P1's round-1 plan and P5 copied P1's round-2 plan step for step; only two exchanges were earned — P4 carving ledger/payment-halt/security pages out of P1's Duty Triage Desk (step 8) and P2 generalising P4's integrity flag into five modifiers (step 6).
- Criticism was nearly absent; disagreements ended by fiat. The severity argument (SEV0 tier vs four levels vs flags vs modifiers) closed with three unreconciled answers and nobody testing them against the 31 re-coded incidents; P2's caution against merging small teams into domains was ignored, not rebutted; P1's 11-predecessor join at step 27 and its growth to 35 steps went unchallenged.
- Problem 1: two of five agents became copyists. P3 and P5 added nothing the others lacked from round 2 on, wasting 40% of the parallelism and inflating apparent consensus, so the final vote discriminated mostly on length and two or three unique steps.
- Problem 2: metrics became boilerplate carrying an overclaim. "$1,000 primary week", "2 out-of-hours pages/responder", "30 commanders / 18 comms leads" appear near-verbatim in four plans, and P1, P3, P4 and P5 all promise "SOC 2 passes with zero exceptions" — an outcome the company does not control; only P2 says "no unresolved material exception".
- Problem 3: nothing was ever costed or feasibility-tested. Every plan reserves 10–15% of 260 engineers (~39 FTE) plus $0.9–1.3M compensation plus tooling and a three-person office against $1.3M of credits without doing the subtraction, and regressions spread as easily as improvements (P1's cut from two control tests to one at month 5 was copied by P4 and P5).
- Fixes that would help most: (1) assign non-converging adversarial roles from round 1 — a ≤12-step minimum-viable plan, a cost/feasibility attacker, and a defender of a structurally different coverage model such as a managed overnight NOC or a dedicated 12-person SRE first line; (2) require a written critique artefact each round naming three flaws by proposal and step number plus a visible dissent register, so options like P2's numeric thresholds (>10% failures for 5 min = SEV1) cannot be dropped silently; (3) mandate one costing table per plan and ban shared metrics, capping each plan at ~10 metrics derived from its own design with a falsification condition and an owner.
At a glance
- Analyst's first choice, blind to the vote: proposal 1 · opus5_refine_1
- The vote: proposal 1 · opus5_refine_1 — the analyst agrees
- Against the initial proposals: better than every initial proposal
- Rounds: round 1 converging · round 2 converging · round 3 converging
- The process: 10 problems observed, 10 suggestions
The final round, in the analyst's words
- P1 — the most complete plan (35 steps): the only one with change-intelligence/deploy-safety (step 13), Support and account managers as an instrumented detection tier (step 20), a concurrency doctrine with a Multi-Incident Coordinator (step 8) and a written warning that credits may rise before they fall (step 22).
- P2 — the most rigorous on measurement and audit: defines "actionable page" and "noise" and counts notification episodes (step 12), five severity modifiers instead of a fifth level (step 6), two independent control tests in months 4 and 6, company-wide by week 18, and the only honest audit claim ("no unresolved material exception"); the only plan with no programme risk register.
- P3 — a competent merge of P1's architecture with P2's day-one compliance step; complete but contributes nothing of its own and carries two internal inconsistencies (month-5 vs months-4-and-6 testing; "five months before the report date" against a week-24 rollout).
- P4 — the most compact serious plan (26 steps) with the two sharpest operational rules in the field: a Duty Triage Desk that never holds ledger/payment-halt/security pages (step 8) and a named owner for the 3 a.m. status page; weakened by bundling (steps 14, 23) and by burying contingencies in the final step.
- P5 — complete and correct but almost entirely derivative of P1's earlier draft, with two overloaded steps (14, 17) and a metric that contradicts its own day-7 floor.
- All five share: day-7 operating floor, forensic re-coding of the 31 incidents, written fairness contract, service catalog with Tier 0–3, SEV1–SEV4 with ledger auto-escalation, IC/comms/scribe/SME with dual control preserved, ~30-person command corps plus 10–12 domain rotations plus business-hours Layer C, ~$1,000 primary week with a payroll hard gate, two out-of-hours pages/responder/week budget, synthetic money-path probes plus the 31-incident replay test, one paging platform, 15/30-minute status-page clocks, 3/5/10-day blameless postmortems with verified action closure, pilot then gated waves, month-7 mock audit, 90-day inspect-and-adapt.
Raw prompts and responses of every analysis call
[ROUND 0]
[SYSTEM]
You are an expert reviewer of multi-agent planning processes.
Several LLM agents drafted plans for a task, refined them over a number of rounds while seeing each other's proposals, and finally voted for the best one.
Be exhaustive but precise: name concrete steps, ideas and metrics, never generalities. Judge plans by their fitness for the task as stated, their realism, their completeness, the soundness of their order and dependencies, how measurable their success is and how they handle things going wrong.
You are an impartial evaluator, not a chronicler: assess the proposals and the process on their merits, never rationalise what happened or assume that the outcome was right.
After your analysis, answer in the requested structure.
Every text field you write will be read by a busy person who skims. Make it easy to skim: short sentences and short paragraphs; when you name several things, prefer a list to a paragraph, with sub-items when an item has parts, but keep a single fact as a sentence; lead with the point and then the evidence; name proposals and steps by number (P2, step 4); no preamble, no repetition of the question, no closing summary; bold at most one key phrase per item or paragraph. Text fields accept Markdown: a blank line between paragraphs, "- " for lists, **bold**.
[HUMAN]
Task given to the agents: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
This is round 0: every agent wrote its plan independently, without seeing the others.
--- PROPOSAL 1 (agent opus5_initial_1, anthropic/claude-opus-5) ---
Estimated complexity: high
Success metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 9 months.
- Customer-first detection falls from 40% of incidents to under 10% within 6 months and under 5% within 12.
- Median time to mitigate falls from 3 h 10 min to under 60 minutes within 12 months.
- An Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes.
- Status page updated within 15 minutes of SEV1 declaration and 30 minutes of SEV2 in 95% of cases.
- Monthly alert volume falls from 3,400 to under 500 pages, with actionability above 75%; out-of-hours pages under 2 per person per week.
- All six legacy alerting tools consolidated into one paging platform, legacy paging paths disabled, by week 16.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by their due date.
- SLA credits fall from $1.3M to under $400k in the first 12 months.
- Customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar.
- All 28 teams onboarded by week 24, with 24x7 rotations of 6+ certified responders for every Tier 0/1 team.
- 30+ certified Incident Commanders and 20+ certified Communications Leads, giving 24x7 primary and secondary command cover.
- Paid on-call policy approved by HR, Legal and Finance and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among engineers on rotation.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked action items.
Steps (29):
1. Program charter, executive mandate and funding
Convert the CEO's frustration into a named program with one accountable owner, a budget and a deadline that is earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, the 28 teams will opt out.
- Appoint a single Incident Management Program Lead (Head of Reliability/SRE) with direct exec sponsorship from CTO and CEO.
- Form a steering group: CTO, VP Eng, Head of Support/CS, CISO/Compliance, Legal, Finance (SLA credits), HR (on-call pay).
- Set the non-negotiables: one severity scale, one paging tool, one postmortem format, mandatory action tracking, paid on-call.
- Fix the timeline: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready by week 28 (four weeks of buffer before the audit).
- Approve budget lines: tooling (~$150–250k/yr), on-call compensation (~$600k–1M/yr), 2–3 dedicated program FTEs. Anchor it against $1.3M of credits plus incident cost.
2. Forensic baseline of the 31 incidents and the alert estate (depends on: 1)
Before designing anything, rebuild the facts. Re-open all 31 incidents and profile the 3,400 monthly alerts so every later design decision is evidence-based.
This also creates the "before" picture the exec and the auditor will compare against.
- Re-code each incident: trigger, service, detection source (customer vs monitor), timestamps for detect/acknowledge/declare/mitigate/resolve, who led, credits paid, root cause family.
- Quantify the 40% customer-first detections: which signal was missing in each case.
- Classify the two "nobody in charge" incidents minute by minute; use them as the burning-platform story.
- Audit the six alerting tools: volume per tool, per team, per alert rule; identify the top 50 rules that produce most of the 85% noise; find rules with no owner and no runbook.
- Baseline the numbers formally: MTTD 22 min, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed. Freeze them as the reference line.
3. Stakeholder listening tour and resistance map (depends on: 1)
Engineer pushback against "carrying a pager for other teams' code" is the main delivery risk. Treat it as a design input, not an attitude problem.
Run structured interviews across all 28 teams, plus Support, CS and Sales, in two weeks.
- Test the real objection: is it unpaid work, night sleep, unfamiliar code, poor runbooks, or fear of blame? Each has a different fix.
- Collect current informal practices — the 12 teams already on-call are the pilot candidates and the source of veterans.
- Document the promise that answers the objection: **you are only paged for services your team owns**, plus a trained commander who runs the incident and pulls in others.
- Map influencers and blockers by team; recruit 10–15 credible engineers as a design working group so the process is co-authored, not imposed.
- Survey baseline sentiment (trust in alerts, willingness to be on-call, burnout) to re-measure at 6 and 12 months.
4. Service ownership catalog and criticality tiering (depends on: 2)
You cannot page the right person across 180 services until each service has a named owning team. This is the foundation of both on-call fairness and severity mapping.
Build a machine-readable catalog (Backstage or equivalent) that is the single source of truth for routing.
- One owning team per service, a named engineering manager, a Slack channel, a paging escalation policy, a dependency list.
- Tier services by business impact: Tier 0 (money movement, ledger, auth, shared PostgreSQL cluster), Tier 1 (customer-facing but degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map each Tier 0/1 service to the customer-visible capability it supports (payment initiation, settlement, reporting, onboarding).
- Flag orphan services and cross-team shared components; force an ownership decision for each within 30 days, or schedule decommissioning.
- Publish coverage gaps to the steering group: any Tier 0 service without an owner is an executive escalation.
5. Severity scale and declaration criteria (depends on: 2, 4)
Define a five-level scale with objective, payments-specific triggers so declaration is a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade.
Each level triggers a fixed bundle of response, comms and postmortem obligations.
- **SEV1**: money movement stopped or incorrect, ledger integrity in doubt, data breach, full region loss, >10% of customers impacted. Triggers: immediate 24x7 page of IC + comms + exec, bridge within 5 min, status page within 15 min, mandatory postmortem, regulator assessment.
- **SEV2**: severe degradation, settlement at risk of missing a window, single large/strategic customer fully down, SLA breach likely. Triggers: IC paged, status page within 30 min, mandatory postmortem.
- **SEV3**: partial or workaround-available degradation, no credit exposure. Team-led, business-hours comms, postmortem optional but encouraged.
- **SEV4/5**: minor or internal-only; ticket-tracked, no paging.
- Add auto-escalation rules: any SEV3 open >2 h, or any incident touching the shared ledger cluster, becomes SEV2 automatically. Include a severity decision tree and 12 worked examples drawn from the 31 real incidents.
6. Incident roles, decision authority and handover rules (depends on: 5)
Solve the "nobody in charge for an hour" failure by making command explicit, transferable and logged.
Define five roles with written responsibilities, entry criteria and explicit authority.
- **Incident Commander**: owns the incident, not the fix. Authority to declare severity, pull any engineer, approve customer-impacting mitigations, invoke failover and authorise spend. The IC never types in the terminal.
- **Communications Lead**: owns status page, internal updates, account-manager briefings and the exec summary. Single voice to customers.
- **Scribe**: maintains the timeline, decisions and open questions; feeds the postmortem and the audit evidence trail.
- **Subject-Matter Responders**: engineers from owning teams; they investigate and remediate, and report to the IC.
- **Executive Liaison** (SEV1 only): shields the IC from exec questions and owns regulator/board escalation.
- Rules: the IC role is assumed within 5 minutes of declaration, stated explicitly in the channel ("I am IC"), and any handover is announced and logged. Roles may be combined below SEV2; never at SEV1.
7. 24x7 incident command coverage model (depends on: 6, 4)
Command is staffed by a small, trained, cross-team pool — not by 28 teams individually. This is what makes 24x7 realistic in one New York time zone.
Design a central rotation that scales with a growing certified pool.
- Create a **Duty Incident Commander** rotation of 25–35 certified volunteers (target ~1 per team, plus managers and senior engineers), giving each person roughly one week per 6–8 months.
- Pair with a Duty Comms Lead rotation (Support/CS leads plus engineering managers, ~15–20 people) and a Scribe pool (rotating, lowest barrier, used as the training entry point).
- Coverage: primary + secondary IC at all times; hard 5-minute acknowledgement SLA with automatic failover to secondary, then to the on-call engineering director.
- Night coverage options to evaluate in writing: US-only rotation with paid night stipend now; a Lisbon/Dublin or APAC follow-the-sun cell as a 12-month option; a 24x7 NOC-style triage desk for first-line detection.
- Eligibility: certification required (S21); commanders are volunteers with manager approval and can step out with 30 days' notice.
8. Team on-call structure, rotations and routing rules (depends on: 4, 6)
Rebuild team on-call around the principle that answers the pushback: **you are only paged for code your team owns**.
Apply a tiered obligation so 28 teams are not treated identically.
- Tier 0/1 owning teams (expected 14–18 teams): 24x7 primary + secondary, minimum 6 people per rotation, one-week shifts, handover Wednesday mornings.
- Tier 2/3 teams: business-hours on-call with a best-effort out-of-hours escalation path, no night paging.
- Platform/Infrastructure and Database teams: 24x7, since they own the shared PostgreSQL ledger cluster and the Kubernetes/regional layer.
- Rotations under 6 people are merged across teams or backfilled by hiring; no rotation of fewer than 4 is approved.
- Routing: every page resolves through the service catalog to the owning team's escalation policy; cross-team pages are made by the IC, never by an alert.
- Guardrails: maximum one week in four, no on-call in the week after a SEV1 you led, protected recovery time after any night page, and a per-person page budget (see S10).
9. On-call compensation, labour compliance and fairness policy (depends on: 8, 3)
Unpaid on-call is both a retention risk and a legal exposure in New York. Paying for it is the fastest way to convert resistance into participation.
Design the scheme with HR, Legal, Finance and Payroll, and publish it before asking anyone to sign up.
- Base stipend per week on rotation, differentiated by tier: e.g. $800–1,200 for 24x7 Tier 0/1, $300–500 for business-hours rotations, with premiums for holidays and weekends.
- Per-incident payment for out-of-hours activation (e.g. $150 per night page plus hourly beyond one hour) and guaranteed time-off-in-lieu after night work.
- Separate Duty IC stipend, since command is a distinct and heavier burden.
- Verify FLSA exempt/non-exempt treatment, NY State wage rules and overtime exposure for non-exempt staff; document the legal review.
- Budget and model the annual cost; get board/CFO approval as a line item, benchmarked against $1.3M of credits.
- Add non-cash elements: on-call time counted as delivery load (teams reduce sprint commitment by ~15%), incident leadership recognised in promotion criteria, and a public quarterly report of on-call load per team.
10. Alert quality standard and page budget (depends on: 2, 4)
3,400 alerts a month at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a contractual condition of paging someone.
Publish a standard, then enforce it mechanically.
- Every paging alert must have: a named owning team, a documented customer impact, a runbook link, a tested threshold, and a severity mapping. Alerts failing this are demoted to ticket or deleted.
- Page only on symptoms that affect customers (SLO burn rate, error budget, queue depth against settlement deadlines); cause-based CPU/memory alerts become dashboards or tickets.
- Set a **page budget**: maximum 2 out-of-hours pages per person per week. Breach triggers a mandatory alert-tuning sprint for the owning team and blocks new alert creation.
- Auto-quarantine: any alert that fires more than 5 times a month without action, or has >70% no-action acknowledgements, is silenced automatically and returned to its owner.
- Monthly alert review per team: kill, tune, or keep, with the numbers on screen.
- Target: 3,400 → under 500 pages a month, with actionability above 75% within six months.
11. Detection uplift: SLOs, synthetic journeys and ledger assurance (depends on: 10, 4)
The goal is to stop customers telling you first. Detection must be driven by customer-visible outcomes, not host metrics.
Instrument the money path end to end and alert on it.
- Define SLOs for each Tier 0/1 customer capability: payment initiation success rate, authorisation latency, settlement file timeliness, API availability and reporting freshness. Tie them to the 99.95% contractual SLA with a stricter internal target.
- Deploy synthetic transactions from outside the platform, in both regions, every 60 seconds, covering the full payment lifecycle including a small real-value canary flow where feasible.
- Add ledger assurance checks: continuous double-entry balance reconciliation, replication lag and failover-readiness alarms on the shared PostgreSQL cluster, and settlement-window countdown alerts.
- Build per-customer anomaly detection for the top 100 accounts (volume drop, error spike) so a single-tenant outage is detected before the account manager calls.
- Create an "inbound signal" bridge: any support ticket or account-manager report matching impact keywords auto-creates a triage incident within 5 minutes.
- Track every incident's detection source; make "customer detected first" a reviewed defect with its own follow-up action.
12. Tool consolidation and incident platform implementation (depends on: 5, 7, 8, 10)
Collapse six alerting tools into one paging and incident platform so there is a single queue, a single timeline and a single audit record.
Run a short, time-boxed selection and migrate within the pilot window.
- Select an integrated stack: paging/on-call scheduling plus an incident management layer (e.g. PagerDuty + incident.io/FireHydrant, or a single vendor) and a hosted status page.
- Implement one-command declaration in Slack (`/incident declare`) that creates the channel and bridge, pages the Duty IC, sets severity, opens the timeline and starts the clock.
- Migrate all monitoring sources to route into the one platform; decommission direct paging from the legacy six and block new integrations that bypass it.
- Automate the evidence trail: timestamps, role assignments, severity changes, comms sent, and postmortem linkage exported for SOC 2.
- Integrate with the service catalog for routing, Jira for actions, Salesforce/CS tooling for affected-customer lists, and Zoom/Slack Huddle for the bridge.
- Hard requirement: the platform must work when AWS in one region is down — verify out-of-band paging (SMS/phone) and a printed/offline fallback runbook.
13. Detection-to-escalation path and the five-minute command rule (depends on: 6, 7, 12)
Write the single path from "something looks wrong" to "someone is in charge", and make it impossible to skip.
The design target is detection to commander in under five minutes, any hour.
- Entry points: automated alert, engineer observation, support ticket, account manager, partner bank, customer-facing SEV hotline. All converge on the same declaration command.
- Anyone in the company may declare up to SEV2; nobody is punished for over-declaring. Publish that rule in writing and repeat it.
- Auto-page ladder: Duty IC (5 min) → secondary IC (5 min) → on-call Director (10 min) → CTO. Same ladder for the owning team's responder.
- Cross-team pull: the IC can page any team's on-call directly, with a 10-minute acknowledgement obligation. This is the reciprocal commitment that makes single-team ownership viable.
- Explicit takeover protocol: if no one claims IC within 5 minutes, the platform assigns it and announces it; the assignee cannot decline, only hand over.
- Define standing severity triggers for immediate regional failover, ledger read-only mode and partner-bank notification, with pre-authorised decision rights so the IC does not wait for an executive.
14. Internal communications protocol (depends on: 6, 12)
Standardise the internal channel so responders, executives and support see the same picture without interrupting the IC.
Separate the working channel from the audience channel.
- One incident channel per incident (auto-created), one bridge, and a read-only broadcast channel for executives, Support and Sales.
- Update cadence by severity: SEV1 every 30 minutes even if nothing has changed; SEV2 every 60 minutes; SEV3 at state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, ETA or next update time, current IC and Comms Lead.
- Exec briefing rule: executives ask questions only to the Executive Liaison; the IC is not interrupted. Publish this as a behavioural expectation signed by the exec team.
- Support/CS enablement: a live affected-customer list and a holding statement within 15 minutes of SEV1/SEV2 so the front line is never guessing.
- Handover protocol for incidents beyond 4 hours: formal IC handover checklist, fatigue rule, and staffing of a second shift.
15. Customer communications and status page policy (depends on: 5, 14)
Customers currently learn of outages from their own monitoring and hear from whoever happens to be around. Replace that with a timed, owned, pre-approved process.
The Comms Lead is the single author; templates remove the need to write under pressure.
- Timing commitments: status page posted within 15 minutes of SEV1 declaration and 30 minutes for SEV2; updates every 30/60 minutes; resolution notice within 30 minutes of mitigation; customer-facing summary within 5 business days for SEV1.
- Pre-approve 12–15 templates with Legal and Comms (degradation, delay in settlement, API errors, security event, third-party failure) so nothing needs legal review mid-incident.
- Subscription-based status page with per-component granularity mapped to the customer capabilities from S11, plus an email/webhook/RSS feed.
- Tiered outreach: top 100 accounts get a direct named call or email from their account manager within 30 minutes of SEV1, with a briefing pack from the Comms Lead; long tail gets the status page and a proactive email.
- Rules of language: state impact and next update time, never speculate on cause, never assign blame to a vendor before facts are confirmed.
- Run a quarterly customer-perception check with the top accounts on whether comms were timely and useful.
16. Regulatory, partner and legal notification playbook (depends on: 5, 15)
In payments, some incidents are reportable and the clock starts at detection. Build the assessment into the process so it is never an afterthought.
Work with Legal, Compliance and the CISO to produce a decision tree and contact matrix.
- Map obligations: NYDFS Part 500 (72-hour cybersecurity event notification), state breach laws, GLBA/FTC Safeguards, PCI DSS if card data is in scope, sponsor-bank and card-network contractual notice windows, and any FinCEN/OFAC implications.
- Add a mandatory regulatory-assessment checkpoint to every SEV1 and every security-related SEV2, owned by the Executive Liaison, completed within 2 hours of declaration and recorded even when the answer is "not reportable".
- Build the contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel, with 24x7 numbers and named backups.
- Pre-draft notification letters and hold them under legal privilege review.
- Check customer contracts for bespoke notification SLAs (often 1–4 hours for enterprise accounts) and encode them in the customer tiering.
- Test the playbook once per quarter as part of the simulation programme.
17. SLA credit and financial impact workflow (depends on: 5, 15)
Link incidents to money so severity, credits and prioritisation stay consistent — and so Finance stops being surprised.
Make credit calculation an automated output of the incident record, not a negotiation.
- Define the availability measurement method per contract, per component, and agree it with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and the per-capability telemetry; generate a proposed credit schedule within 5 business days of resolution.
- Decide the posture: proactive credits for the top tier (reputation upside) versus claims-based for the rest; document the approval chain.
- Track credits per incident and per root-cause family; feed a quarterly report showing which reliability investments would have prevented which credits.
- Set a target: reduce credits from $1.3M to under $400k in year one, and use that delta as the ongoing business case.
18. Postmortem standard and blameless review forum (depends on: 5, 6)
Replace "some incidents, various formats" with a mandatory, single-format, blameless process with fixed deadlines.
The discipline is in the deadlines and the forum, not in the template.
- Mandatory for: every SEV1 and SEV2, every incident where a customer detected it first, every incident over 2 hours, every repeat of a known cause, and every near-miss involving the ledger. Optional but templated for SEV3.
- Fixed timeline: draft within 3 business days, peer review within 5, published company-wide within 10. The IC owns delivery; the owning team's manager is accountable.
- One template: timeline, customer and financial impact, detection analysis (why not sooner), response analysis (why mitigation took as long as it did), contributing factors, what went well, action items with owner and due date.
- Blameless rules in writing: describe systems and decisions in the context available at the time; no individual named as a cause; HR and management commit that postmortems are never used in performance reviews.
- Weekly 60-minute Incident Review Board: reviews all postmortems from the prior week, challenges quality, ratifies severity, and approves or rejects action items. Attendance by engineering directors is mandatory.
- Publish a searchable postmortem library and a quarterly "top five recurring causes" analysis.
19. Action item ownership and tracking system (depends on: 18, 12)
11 of 64 actions closed is the single clearest symptom of a process nobody enforces. Give actions the same status as customer commitments.
Track them where engineering work already lives, with visible escalation.
- Every action gets: a named individual owner (not a team), a priority class, a due date and a Jira ticket auto-created from the postmortem.
- Priority classes with hard SLAs: P0 prevents recurrence of a SEV1, due in 30 days; P1 in 60 days; P2 in 90 days. P0s are committed into the next sprint before any roadmap work.
- Capacity rule: teams reserve a standing 15–20% of sprint capacity for reliability and incident actions. Without reserved capacity, the actions will not land.
- Escalation ladder for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Monthly reporting of closure rate by team in the engineering leadership review; include it in manager performance objectives.
- Target: 90% of P0/P1 actions closed on time within two quarters.
20. Runbooks, major-incident playbooks and the on-call readiness bar (depends on: 4, 8)
Nobody can respond well to unfamiliar systems at 3 a.m. without runbooks — and poor runbooks are a real part of the pager resistance.
Define a minimum readiness bar that a service must meet before it is allowed to page anyone.
- Readiness checklist per Tier 0/1 service: current architecture diagram, dependency map, dashboard link, alert-to-runbook mapping, rollback procedure, feature-flag kill switches, escalation contacts, and a data-loss/latency impact statement.
- Write major-incident playbooks for the top failure modes derived from S2: shared PostgreSQL ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank/third-party outage, settlement-window breach, and suspected security compromise.
- Prioritise the shared ledger: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and a clearly stated data-loss tolerance (RPO/RTO) signed off by the exec.
- Runbooks must be tested at least twice a year in a drill; untested runbooks are marked stale in the catalog.
- Enforcement: a service without readiness sign-off cannot create paging alerts, and the gap is reported to its director.
21. Training, certification and the commander academy (depends on: 6, 13, 14, 18)
Command is a skill, not a title. Build a certification path so 24x7 coverage is staffed by people who have practised.
Use a tiered curriculum with real assessment.
- **Scribe** (2 hours): timeline discipline and tooling. The entry point for everyone.
- **Responder** (half a day): severity scale, declaration, escalation, runbook use, comms hygiene. Mandatory for every engineer joining an on-call rotation.
- **Incident Commander** (two days plus shadowing): command presence, delegation, decision-making under uncertainty, severity calls, handover, exec management. Requires two shadowed incidents and one simulated SEV1 before certification.
- **Communications Lead** (one day): status-page writing, customer tiering, legal boundaries, regulator triggers.
- Certification is valid 12 months and renewed via a simulation; the register of certified people is an audit artefact.
- Add on-call onboarding per team: a new joiner shadows two shifts before holding primary, and never holds primary in their first 90 days.
22. Simulation programme: game days, drills and wheel of misfortune (depends on: 21, 12, 20)
The process must be rehearsed before it meets a real SEV1. Simulations also build the commander pool and expose runbook gaps cheaply.
Run a standing calendar rather than one-off exercises.
- Monthly 60-minute tabletop ("wheel of misfortune") per engineering group, using a real past incident from the 31.
- Quarterly full-scale game day in production or a production-like environment: regional failover, ledger replica promotion, dependency failure, with the whole role structure activated and timed.
- Twice-yearly unannounced paging drill to measure real acknowledgement times at night.
- One security-incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Every exercise produces a lightweight postmortem and action items in the same system as real incidents.
- Measure and publish drill metrics: time to IC, time to first status update, time to correct mitigation decision.
23. Pilot with wave 0 teams (depends on: 22, 9, 11)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Run a six-week pilot with tight measurement and a public verdict.
- Select 5–6 teams: core payments, ledger/database, platform/Kubernetes, API gateway, plus two of the 12 teams already on-call.
- Activate the full stack for them: new severity scale, Duty IC rotation, single paging tool, alert budget, status-page policy, mandatory postmortems, paid on-call.
- Hold a weekly pilot retro; expect and document 20–40 process defects, and fix them in the standard before rollout.
- Validate the hard questions: does the 5-minute IC rule hold at 3 a.m.? Do cross-team pulls get answered? Is the severity tree unambiguous?
- Exit criteria: MTTD under 10 minutes for pilot services, IC assigned within 5 minutes in 95% of incidents, page volume down 50%, all postmortems on time, positive on-call sentiment.
- Publish a one-page pilot result to the whole company — this is the main adoption argument for the remaining teams.
24. Metrics, dashboards and the review cadence (depends on: 12, 18, 23)
Instrument the process itself, so improvement is visible and the audit has evidence of monitoring and review.
Define a small set of metrics with owners and a fixed meeting rhythm.
- Response metrics: MTTD, time to declare, time to IC assigned, MTTA, MTTM, MTTR, incidents per month by severity, % detected by customers first.
- Quality metrics: page volume per person per week, alert actionability rate, page budget breaches, postmortem on-time rate, action closure rate and ageing.
- Business metrics: SLA credits paid, availability against 99.95% per capability, error-budget consumption, repeat-incident rate.
- People metrics: on-call load distribution across teams, out-of-hours pages per person, on-call sentiment and attrition among on-call staff.
- Cadence: weekly Incident Review Board (postmortems and actions), monthly Reliability Review (metrics per team, alert hygiene, on-call load), quarterly Executive/Board review (credits, trends, investment asks), annual policy review.
- Every metric gets a target and a named owner; dashboards are self-serve and public inside the company.
25. Wave rollout across all 28 teams with readiness gates (depends on: 23, 24)
Roll out in four waves of six to eight teams, every three weeks, ordered by criticality. Each wave passes an explicit gate rather than a deadline.
Gates keep quality high and make the standard credible.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with 6+ certified responders, one IC candidate nominated, one drill passed.
- Assign each wave a named coach from the program team for three weeks of hands-on support.
- Gate criteria are checked and signed by the director; teams that fail are re-scheduled, not waived.
- Freeze legacy tooling per wave: after onboarding, the old alerting paths are disabled, not left as a fallback.
- Publish a live adoption scoreboard by team so progress is social, not administrative.
26. SOC 2 control mapping, evidence automation and internal dry-run audit (depends on: 18, 19, 25)
Design the process so audit evidence is a by-product of doing the work, then test it before the auditors do.
Engage the auditor early to confirm the interpretation of controls.
- Map the process to the Trust Services Criteria: CC7.3 and CC7.4 (incident identification, response, recovery), CC7.2 (monitoring), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), plus availability criteria A1.2.
- Produce and approve formal policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard — versioned, signed, annually reviewed.
- Automate evidence: incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, postmortem library, action-item closure reports, training and certification register, drill records.
- Confirm the observation window with the auditor and ensure the process is operating for a minimum of three months before fieldwork.
- Run an internal dry-run audit at month six: sample 15 incidents and walk the full evidence chain; fix gaps with 8 weeks to spare.
- Keep a remediation log for any incident where the process was not followed, with the corrective action — auditors respond better to documented exceptions than to a claim of perfection.
27. Change management, incentives and communications campaign (depends on: 3, 9, 23)
Run this in parallel from day one. The process will be judged by engineers on fairness, and by executives on visible results.
Communicate the deal explicitly and repeatedly.
- The deal in one sentence: **you are paid for on-call, you are paged only for what you own, a trained commander runs the incident, and your postmortem actions get real sprint capacity**.
- Launch communications: CTO all-hands, per-team roadshows, a one-page process card for laptops, an internal wiki hub and a Slack support channel with a 4-hour answer SLA.
- Recognition: incident-response contribution in promotion criteria and performance frameworks, quarterly awards for best postmortem and biggest alert-noise reduction, public thanks after every SEV1.
- Manager accountability: adoption, alert hygiene, action closure and on-call load in each engineering manager's quarterly objectives.
- Handle the exceptions: a written path for engineers who cannot do nights (caring responsibilities, health), covered by stipended volunteers elsewhere.
- Track sentiment quarterly and publish the results, including bad news, to keep credibility.
28. Program risk register and contingency planning (depends on: 1)
Name the ways this program fails and pre-commit the response. Review it monthly in the steering group.
The main risks are predictable.
- **Volunteer shortfall for the IC pool**: contingency is to make command a rostered duty for engineering managers and senior engineers until the pool reaches 25.
- **Compensation not approved in time**: fall back to time-off-in-lieu plus a phased stipend, but do not launch mandatory night on-call without some compensation.
- **Tool migration slipping**: keep the single-queue requirement and cut scope on the incident-management layer, not on paging consolidation.
- **Alert pruning causing a missed incident**: prune from paging to ticket first, observe for 30 days, then delete; keep a recovery path.
- **Burnout or attrition among the 12 experienced on-call teams**: monitor load weekly and cap individual page counts.
- **A major SEV1 mid-rollout**: pre-agree that the program lead becomes a full-time responder and the wave schedule slips by one wave, with the steering group informed the same day.
29. Continuous improvement, maturity roadmap and post-audit sustainability (depends on: 24, 25, 26)
Protect against the classic failure: the process decays once the audit is passed. Build the second-year plan before the first year ends.
Set a maturity model and a forward roadmap with owners.
- Quarterly process retrospective with the IC pool: what in the process itself slowed us down, what needs simplifying, is the severity scale calibrated?
- Re-baseline targets every six months; a process that hits all targets is under-ambitious.
- Year-two roadmap candidates: follow-the-sun coverage cell, automated mitigation and self-healing for the top three recurring causes, error-budget policy that gates releases, per-customer real-time impact reporting, and blast-radius reduction for the shared ledger cluster (the largest single structural risk).
- Move from lagging metrics (MTTR) to leading ones (error-budget burn, near-miss rate, drill performance).
- Make the annual policy review, certification renewal and drill calendar permanent calendar items owned by the Head of Reliability, independent of the audit cycle.
- Report to the board quarterly on availability, credits and incident trends so the process keeps executive attention after SOC 2 is signed.
--- PROPOSAL 2 (agent gpt5.6-sol_initial_2, openai/gpt-5.6-sol) ---
Estimated complexity: high
Success metrics: - Within 7 days, every suspected SEV0–SEV2 has one incident record, one channel, and a named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named incident commander for more than 10 minutes.
- By day 30, 100% of Tier 0 and Tier 1 services have a named owner, primary escalation, secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and subject-matter coverage.
- By day 120, 100% of production services have sustainable ownership and tested escalation paths.
- At least 95% of SEV0 and SEV1 pages are acknowledged within 5 minutes by month 3.
- At least 95% of SEV2 pages are acknowledged within 10 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to 10 minutes by day 90 and 5 minutes by month 6.
- The proportion of incidents first detected by customers falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to below 90 minutes by day 120 and below 60 minutes by month 6.
- At least 95% of qualifying incidents meet their initial customer-communication deadline by month 3.
- At least 95% of published incidents meet their required update cadence by month 3.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without loss of Tier 0 or Tier 1 detection coverage.
- Monthly pages fall from 3,400 to no more than 1,500 by day 90, with alert actionability and missed-detection reviews used as countermeasures against unsafe suppression.
- 100% of new paging alerts satisfy the owner, runbook, dashboard, action, severity, and escalation quality rules by day 60.
- 100% of required SEV0 and SEV1 postmortems are drafted within 3 business days and reviewed within 5 business days by month 2.
- At least 90% of postmortem actions are completed by their approved due dates by month 6.
- All 53 currently open historical actions are triaged within 30 days; all unaccepted high-risk items are completed within 90 days.
- Repeat incidents with the same unaddressed contributing factor decline by at least 50% within 6 months.
- Every primary rotation has at least six trained responders or a documented, time-limited executive exception by day 120.
- No responder is routinely scheduled more frequently than one primary week in six by day 120.
- Two end-to-end cross-company exercises, including regional and ledger scenarios, are completed before the audit, with all critical findings assigned and tracked.
- Monthly availability meets or exceeds the 99.95% contractual target by month 6, with exceptions reviewed at the executive reliability meeting.
- SLA credits decline by at least 50% on an annualized trailing basis by month 8.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (19):
1. Establish ownership, authority, and funding
Launch the program within 48 hours under an executive sponsor. Give one program owner authority to standardize incident management across all 28 teams.
- Name the CTO or equivalent as executive sponsor and a Head of Incident Management or Reliability as directly accountable owner.
- Form a working group with Engineering, SRE or Platform, Product, Support, Customer Success, Communications, Security, Legal, Compliance, Risk, HR, Finance, and Internal Audit.
- Approve authority for an incident commander to stop deployments, roll back releases, disable features, shift traffic, invoke continuity plans, and pause payment processing when integrity is at risk.
- Preserve financial controls. The incident commander may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access requirements.
- Fund paging tools, compensation, training, observability work, exercises, and dedicated reliability capacity.
- Reserve engineering capacity for incident remediation. Start with 10% of capacity and adjust through quarterly risk reviews.
- Record the current baselines: 31 customer-impacting incidents, 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M in credits, 3,400 monthly alerts, 85% noise, and 11 of 64 actions closed.
- Maintain a risk register for staffing gaps, shared-ledger concentration, regional failover, alert coverage, third parties, and audit readiness.
2. Install immediate minimum controls (depends on: 1)
Put an interim process in place during the first seven days. Do not wait for tool consolidation, policy perfection, or the SOC 2 audit.
- Publish a one-page interim severity guide and incident declaration procedure.
- Establish one continuously monitored incident declaration path through chat, telephone, and the paging system.
- Create a standard incident channel, conference bridge, incident document, and event naming convention.
- Staff an interim primary and backup incident commander at all times. Compensate this duty retroactively under the final compensation policy.
- Give trained duty personnel access to the status page, paging system, dashboards, support queue, service catalog, and emergency contacts.
- Require an incident commander to be named within 10 minutes for every suspected major incident.
- Direct Support to escalate credible customer reports immediately rather than waiting for engineering confirmation.
- Triage all 53 open historical postmortem actions. Complete, re-plan, or formally risk-accept the items affecting ledger integrity, payment duplication, regional resilience, security, and detection first.
- Hold a daily 15-minute operational review until permanent controls are working.
3. Create the service and dependency catalog (depends on: 1)
Build a reliable ownership map for all production services and customer journeys. This is the basis for paging, escalation, impact assessment, and audit evidence.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, data stores, queues, external processors, banking partners, and customer-facing endpoints.
- Assign each component a single accountable team, primary responder group, secondary escalation group, engineering manager, and product owner.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 according to financial integrity, customer impact, dependency centrality, and contractual obligations.
- Treat the ledger, payment orchestration, authentication, settlement, reconciliation, and critical shared infrastructure as Tier 0 or Tier 1.
- Map every important customer journey to its service, database, cloud-region, and third-party dependencies.
- Record SLOs, RTOs, RPOs, data classification, dashboards, runbooks, deployment controls, feature flags, and failover methods.
- Assign separate but coordinated responders for the ledger application and the shared PostgreSQL platform.
- Document whether each service is active-active, active-passive, or region-bound. Identify dependencies that make nominal regional redundancy ineffective.
- Make missing ownership or missing runbooks a release-blocking risk for Tier 0 and Tier 1 services.
4. Adopt severity and incident lifecycle standards (depends on: 1)
Approve one impact-based severity model for operational, security, data, and third-party incidents. When evidence is incomplete, start at the higher credible severity and downgrade later.
- SEV0, crisis: Use for actual or credible unauthorized, lost, duplicated, or corrupted movement of money; ledger integrity loss; material security compromise; material data exposure; both-region failure; or an event likely to require crisis or regulatory management. Page all roles immediately, engage executives, Security, Legal, Compliance, and Risk, and consider pausing payment activity.
- SEV1, critical: Use for widespread inability to initiate, process, settle, or reconcile payments; a core journey failing without a viable workaround; material regional impact; fast SLA-budget exhaustion; or an imminent integrity risk. Staff all incident roles, notify the executive duty officer, and publish customer communications.
- SEV2, major: Use for a material customer subset, one or more critical customers, significant degradation with a workaround, partial transaction failure, or a likely contractual impact. Assign an incident commander and subject-matter responders; add communications and scribe roles whenever customers are affected.
- SEV3, minor: Use for localized, low-impact degradation with no financial-integrity, security, regulatory, or material contractual risk. The owning team leads the response and keeps an internal record; external communication is not normally required.
- Base severity on actual or credible impact, not the seniority of the reporter, number of alerts, or presumed complexity of the fix.
- Permit any employee to declare an incident. Only the incident commander may lower severity after recording the evidence and rationale.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of customer impact. Define resolution only after stability, backlog processing, transaction recovery, and required ledger reconciliation are complete.
- Start measurement from the earliest reliable indication of impact, including telemetry, customer reports, and partner notifications.
5. Define roles and sustainable 24x7 staffing (depends on: 3, 4)
Separate command from technical remediation. This allows trained commanders to coordinate any incident without asking engineers to debug code they do not own.
- Incident commander: Owns severity, priorities, role assignment, escalation, decision cadence, mitigation strategy, handoffs, and final closure. One person has command at a time.
- Communications lead: Owns internal notices, status-page updates, account-manager briefs, approved customer language, and coordination with Legal or regulators.
- Scribe: Maintains a timestamped timeline of observations, decisions, commands, owners, and status changes. Automation may assist but does not replace human validation for SEV0 and SEV1.
- Subject-matter responders: Diagnose and mitigate only services or domains for which they have accepted ownership, training, access, and runbooks.
- Executive duty officer: Removes organizational obstacles and approves exceptional business decisions. This role does not take command unless a formal transfer occurs.
- Security, Legal, Compliance, Support, Vendor Management, and Business Continuity join according to predefined triggers.
- Create a company-wide incident-command rotation with at least eight certified primary commanders and eight qualified backups. Use weekly rotations with explicit handoffs.
- Create similarly sustainable communications and scribe pools using Engineering Operations, Support, Customer Operations, and Communications personnel.
- Group service responders into approximately 8–12 coherent product or platform domains rather than creating 28 fragile rotations. Each domain rotation should normally contain at least six trained responders.
- Do not place an engineer into another team's responder pool without training, access, runbooks, shadow shifts, and explicit acceptance by both teams.
- Maintain dedicated database-platform and ledger-application escalation coverage for the shared PostgreSQL environment.
- Require distinct people for commander, communications, and primary technical lead during SEV0 and SEV1 incidents.
- Require a verbal and written handoff when any incident role changes. Record the exact time and the new role owner.
6. Implement compensation and fatigue safeguards (depends on: 5)
End unpaid on-call before expanding coverage. Treat availability, interrupted personal time, and overnight recovery as compensable work.
- Pay a fixed stipend for each primary on-call week and a secondary stipend equal to a defined percentage of the primary amount.
- Pay a higher holiday stipend. Apply overtime and call-out rules to non-exempt employees as required by law.
- Give exempt employees a minimum call-out credit or equivalent paid recovery time for material after-hours work.
- Provide a paid recovery day after prolonged overnight work, a SEV0, or a qualifying SEV1. Managers must arrange daytime coverage rather than expecting normal output.
- Have HR, Finance, and employment counsel publish dollar amounts, tax treatment, eligibility, and payroll procedures within 14 days. Apply the policy consistently across teams and locations.
- Target rotations no more frequent than one week in six. Exceptions require a time-limited staffing plan and executive risk acceptance.
- Avoid consecutive primary and secondary weeks. A person must not be primary for two simultaneous domain rotations.
- Track after-hours pages, sleep interruptions, swaps, missed acknowledgements, and reported burnout by rotation.
- Trigger a staffing or alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
- Permit responders to declare themselves temporarily unfit after overnight work without performance penalty.
7. Consolidate incident and paging tooling (depends on: 4)
Create one operational system of record while migrating safely from the six current alerting tools. Consolidation must reduce ambiguity without creating a monitoring gap.
- Select one enterprise paging and escalation platform and one integrated incident record.
- Initially ingest events from all six tools. Deduplicate, correlate, and route them through the new platform before retiring sources.
- Integrate paging with chat, conference bridges, ticket tracking, the service catalog, observability tools, and the customer status page.
- Automatically capture declaration time, acknowledgements, role assignments, severity changes, messages, decisions, mitigated time, and resolved time.
- Use role-based access, multifactor authentication, break-glass controls, immutable audit logs, and periodic access reviews.
- Provide mobile and telephone fallback paths if chat, identity, or the primary paging tool is unavailable.
- Test paging, escalation, status publication, and conference access every week.
- Retire a legacy alert path only after its signals have named owners, successful end-to-end tests, and at least two weeks of verified operation in the new platform.
8. Improve detection and enforce alert quality (depends on: 3, 7)
Shift detection toward customer journeys, payment outcomes, and ledger integrity. Infrastructure metrics alone will not solve the current customer-first detection problem.
- Instrument payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, and reporting with SLOs and business-level success metrics.
- Run external synthetic transactions and API checks from outside the production boundary and from both AWS regions.
- Monitor transaction failure rates, processing latency, queue age, unprocessed volume, reconciliation breaks, unexpected ledger balances, duplicate identifiers, regional asymmetry, and third-party response quality.
- Correlate application telemetry with Kubernetes, AWS, PostgreSQL, network, deployment, and feature-flag events.
- Route high-priority support cases and credible partner notifications into the same incident declaration path within five minutes.
- Define noise as a page that is duplicate, informational, unactionable, non-production, or requires no timely human action.
- Require every paging alert to name an owner, affected service, urgency, customer or SLO risk, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Send non-urgent conditions to a ticket queue rather than a pager.
- Run new alerts in shadow mode for at least seven days unless an emergency risk exception is approved. Test both firing and recovery behavior.
- Review any alert with less than 50% actionability or more than three firings in seven days within two business days.
- Never silently disable a noisy alert. Verify compensating detection, record the decision, and assign a correction owner first.
- Review alert actionability, false positives, missed detection, and page load with every responder group each month.
9. Codify acknowledgement and escalation paths (depends on: 3, 4, 5, 7, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts when a qualifying signal or customer report enters the system.
- For SEV0 and SEV1, page the owning primary immediately; page the secondary after five unacknowledged minutes; page the domain manager and company incident commander at 10 minutes; and engage the executive duty officer by 15 minutes.
- For SEV2, require primary acknowledgement within 10 minutes and incident-command assignment within 15 minutes. Escalate to the secondary and manager when either target is missed.
- For SEV3, require acknowledgement within 30 minutes when immediate production action is needed. Otherwise create a prioritized work item.
- Automatically page the company incident commander for any credible integrity or security concern, cross-team event, customer-visible Tier 0 failure, regional event, or unresolved ownership question.
- If impact remains unknown after 15 minutes, raise severity rather than waiting for certainty.
- Let the incident commander summon dependency owners, cloud support, database support, payment processors, banking partners, and vendors through maintained escalation contacts.
- Test vendor contacts and premium-support entitlements quarterly.
- Route alerts with no valid owner to the central command rotation, then treat the missing ownership record as a control defect.
- Require human acknowledgement. Delivery to a device or chat channel does not count.
- Record every missed acknowledgement, failed escalation, and manual contact workaround for review.
10. Standardize live incident execution (depends on: 4, 5, 7, 9)
Give responders one concise operating procedure for the first minutes through resolution. Prioritize limiting customer and financial harm before proving a root cause.
- Open a dedicated channel, bridge, incident record, and timeline immediately for SEV0 through SEV2.
- Have the incident commander state severity, known impact, current hypothesis, immediate objective, assigned roles, and next update time.
- Freeze unrelated production changes during SEV0 and SEV1 incidents. Record exceptions approved by the incident commander.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Use pre-approved runbooks for region failover, Kubernetes recovery, PostgreSQL failover, credential rotation, queue recovery, and payment suspension.
- Guard against split-brain, replay, duplication, and out-of-order processing during regional or database recovery.
- Require reconciliation and controlled backlog processing before declaring payment or ledger incidents resolved.
- Keep diagnosis and mitigation workstreams separate when enough responders are available.
- State decisions and owners aloud and in the incident record. Avoid unrecorded direct-message command paths.
- Require stability for a severity-specific observation period before closure. Reopen the incident if impact recurs during that period.
- Conduct an explicit operational handback to the owning team, Support, and Customer Success.
11. Standardize internal, customer, and regulatory communications (depends on: 4, 5, 7, 10)
Communicate known impact early without waiting for a root cause. Use approved facts, acknowledge uncertainty, and give the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish an initial customer status within 15 minutes when disclosure is operationally and legally appropriate, then update every 15 minutes.
- For SEV1, notify internal stakeholders within 15 minutes, publish an initial customer status within 15 minutes, and update at least every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send an initial notice within 30 minutes. Update at least every 60 minutes.
- Do not normally publish SEV3 events. Notify specifically affected customers if contracts or material impact require it.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. State affected capabilities, customer symptoms, workarounds, regions, and next update time.
- Do not speculate about root cause, blame, security scope, recovery time, or data integrity.
- Give account managers a single approved briefing and an affected-customer list. Prohibit contradictory or improvised incident explanations.
- Maintain templates for outages, delays, data-integrity investigation, third-party failure, regional failure, security events, and resolution.
- Issue a resolution notice only after operational recovery and required reconciliation. Provide a customer-facing incident summary within five business days for qualifying events.
- Have Legal and Compliance maintain a jurisdiction, regulator, sponsor-bank, network, cyber-insurer, partner, and contract notification matrix.
- Where applicable, explicitly track the current New York cybersecurity-event notification clock, including the 72-hour requirement, without assuming every incident is reportable.
- Have Legal record the reportability decision, decision time, evidence, approver, deadline, and submission confirmation.
- Allow Security or Legal to limit public detail during an active threat, but require the reason and an alternative stakeholder plan to be recorded.
- Coordinate service-credit calculations and contractual notices with Finance and Customer Success from the same incident record.
12. Make postmortems mandatory and actionable (depends on: 4, 7, 10)
Use postmortems to improve systems and controls, not to assign personal blame. Keep performance or misconduct processes separate from the learning review.
- Require a postmortem for every SEV0 and SEV1.
- Require one for a SEV2 that affected customers, incurred credits, breached an SLO or contract, involved financial or data integrity, repeated a prior failure, exposed a control gap, or lasted more than two hours.
- Permit incident command, Security, Compliance, or the service owner to require a review for a near miss.
- Produce a factual draft within three business days and hold the cross-functional review within five business days.
- Use one template covering executive summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, what worked, what failed, and lessons.
- Include why monitoring did or did not detect the event before customers.
- Avoid a single-root-cause assumption. Examine technical, organizational, process, dependency, testing, and incentive factors.
- Give every action one owner, due date, priority, expected risk reduction, verification method, and linked engineering item.
- Classify actions as containment due within 7 days, corrective work due within 30 days, or strategic work normally due within 90 days.
- Require director approval and documented residual-risk acceptance for overdue high-risk actions.
- Verify effectiveness after implementation. Closing a ticket without evidence does not close the action.
- Publish broadly useful reviews internally. Maintain access-restricted versions for security, privacy, personnel, or legally privileged details.
13. Measure performance and review it routinely (depends on: 7, 8, 11, 12)
Use outcome, process, quality, and human-sustainability measures together. Do not reward teams for suppressing declarations or hiding incidents.
- Measure detection time from first impact to first internal signal, declaration time, acknowledgement time, role-staffing time, mitigation time, resolution time, and recurrence.
- Report both median and 90th percentile. Break results down by severity, service tier, customer journey, region, detection source, and owning domain.
- Track customer-first detection, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, and role conflicts.
- Track availability, error-budget consumption, failed-payment volume, delayed value, reconciliation breaks, impacted customers, contractual breaches, and service credits.
- Track alert volume, actionability, duplicates, after-hours pages, missed pages, pages per responder, and tool-source distribution.
- Track required postmortems completed on time, actions completed by due date, action age, verified effectiveness, and repeat contributing factors.
- Track rotation size, on-call frequency, swaps, recovery days, attrition signals, and quarterly responder sentiment.
- Hold a weekly operational review for recent incidents, overdue actions, alert problems, and upcoming risk.
- Hold a monthly executive reliability review covering trends, investment decisions, accepted risks, and SLA exposure.
- Hold a quarterly resilience and control review with Security, Compliance, Risk, Internal Audit, and Product leadership.
- Use team scorecards to direct investment and assistance, not individual performance penalties.
- Reconcile dashboard data against a monthly sample of incident records and customer cases to detect metric gaming or missing incidents.
14. Train and certify participants (depends on: 4, 5, 9, 10, 11, 12)
Train people before assigning full independent duty. Use paid working time for training, shadowing, exercises, and certification.
- Train all employees to recognize impact, declare an incident, and find the incident channel and status page.
- Train engineers and Support on severity, escalation, customer-report handling, evidence preservation, and financial-integrity precautions.
- Certify incident commanders through instruction, tabletop exercises, shadow incidents, and observed command performance.
- Train communications leads in status writing, contractual communications, regulator escalation, and avoiding unsupported claims.
- Train scribes in timestamping, decision capture, evidence hygiene, and separating fact from hypothesis.
- Require subject-matter responders to demonstrate dashboard, runbook, rollback, failover, and access competence for their assigned domain.
- Add training to new-hire onboarding and repeat role-specific certification annually.
- Appoint an incident-management champion in each of the 28 teams to collect feedback and support local adoption.
- Conduct listening sessions focused on pager fairness, cross-team boundaries, tooling friction, and psychological safety.
- Publish that command duty is process coordination, not responsibility for understanding or repairing another team's code.
15. Pilot and expand the on-call model (depends on: 3, 5, 6, 8, 9, 14)
Pilot the model on the highest-risk customer journeys before expanding it. Correct staffing, alert, access, and compensation defects at each stage gate.
- Start with the ledger, payment orchestration, Kubernetes platform, PostgreSQL platform, authentication, settlement, and Support intake.
- Run the command, communications, and domain rotations in parallel with existing paths for two weeks.
- Verify primary and secondary coverage, handoffs, access, runbooks, paging, conference access, status publication, and compensation processing.
- Require at least two shadow shifts before independent primary duty.
- Review every pilot page within one business day for routing accuracy, actionability, responder load, and missing context.
- Expand by customer journey and dependency domain, not by arbitrary team order.
- Provide central first-line triage if useful, but keep technical remediation with the accepted service owner.
- Do not use contractors or a managed service as the sole incident commander or sole owner of payment and ledger remediation.
- Permit temporary shared domain rotations only after service owners document training, access, runbooks, and escalation boundaries.
- Set an executive-reviewed deadline and remediation plan for any production service that cannot provide sustainable 24x7 ownership.
16. Exercise regional, ledger, and communication failures (depends on: 10, 11, 14, 15)
Validate the process under realistic conditions before relying on it. Begin in tabletop and staging environments, then use controlled production tests where risk permits.
- Run a company-wide incident-command tabletop within 30 days of policy approval.
- Exercise loss of one AWS region, Kubernetes control-plane degradation, shared PostgreSQL failure, payment-processor failure, queue backlog, credential compromise, and suspected duplicate payments.
- Exercise a simultaneous operational and security event to test command boundaries and disclosure control.
- Exercise status-page failure and loss of the primary chat or paging provider.
- Exercise overnight staffing, role handoff, executive escalation, account-manager messaging, and a potential regulator-notification decision.
- Validate backups, restore procedures, RPO, RTO, failover prerequisites, and post-recovery reconciliation.
- Do not inject uncontrolled changes into the production ledger. Use replicas, staging, simulations, or tightly governed production tests.
- Record exercise observations as tracked actions under the same ownership and due-date rules as incident actions.
- Run at least one domain exercise per quarter and two cross-company exercises before the SOC 2 audit.
17. Execute a time-boxed enterprise rollout (depends on: 2, 6, 7, 11, 12, 13, 15)
Use fixed implementation waves so the audit deadline does not become the start date. Report progress weekly and escalate missed stage gates as business risks.
- Days 0–7: Establish governance, interim command coverage, one declaration path, provisional severity, and daily operational reviews.
- By day 14: Approve the core policy, role definitions, communications timings, compensation design, and historical-action triage.
- By day 30: Complete Tier 0 ownership, certify the first command roster, begin the on-call pilot, enable standard incident records, and run the first tabletop.
- By day 60: Provide 24x7 coverage for all Tier 0 and Tier 1 customer journeys, integrate the six alert sources, and enforce postmortem tracking.
- By day 90: Migrate critical paging, implement customer-journey detection, complete status and regulatory playbooks, and materially reduce alert noise.
- By day 120: Assign sustainable ownership and escalation for every production service and complete the first controlled regional or continuity exercise.
- By day 180: Complete tool retirement decisions, verify action closure, rerun weak scenarios, and demonstrate improving detection and mitigation trends.
- In month 7: Conduct a mock audit and executive readiness review, leaving at least one month to correct evidence or operating defects.
- Use exception records with owners, expiry dates, compensating controls, and executive approval. Do not allow indefinite verbal exceptions.
18. Build SOC 2 evidence as the process operates (depends on: 1)
Design evidence collection at the start rather than reconstructing it before the audit. Demonstrate both control design and sustained operation.
- Map the incident process to applicable SOC 2 criteria with Compliance and the auditor, including detection, response, communication, change management, access, availability, and corrective action.
- Maintain approved, version-controlled policies, procedures, severity definitions, role descriptions, and exception records.
- Preserve rotation schedules, compensation activation, training attendance, certification, paging tests, access reviews, and exercise results.
- Preserve incident declarations, timestamps, role assignments, communications, decisions, status updates, postmortems, and corrective-action evidence.
- Record regulatory and contractual notification assessments, including decisions that no notification was required.
- Define retention, confidentiality, legal-hold, and access requirements for operational and security records.
- Sample evidence monthly and trace incidents from initial signal through action verification.
- Have Internal Audit or an independent control owner test the process in months 4 and 6.
- Correct control failures through tracked actions rather than editing historical records.
- Conduct the formal mock audit in month 7 using the same evidence populations expected for the external audit.
19. Sustain accountability and continuous improvement (depends on: 13, 17, 18)
Make incident management an operating discipline rather than an audit project. Keep policy, staffing, tools, and investment aligned with changing customer and system risk.
- Assign permanent owners for the incident policy, paging platform, status page, service catalog, training program, and metrics.
- Review severity thresholds, communication timings, compensation, and staffing at least annually and after material incidents.
- Use incident trends to prioritize architectural work on the shared ledger, regional independence, deployment safety, dependency isolation, and graceful degradation.
- Review repeat incidents and repeat contributing factors quarterly. Require executive action when remediation repeatedly loses priority.
- Survey responders quarterly and publish actions addressing fatigue, fairness, psychological safety, and tool friction.
- Recognize effective incident leadership, early declaration, useful postmortems, and preventive work.
- Prohibit retaliation for good-faith incident declaration or escalation.
- Provide the board or risk committee a quarterly summary of severe incidents, SLA exposure, regulatory events, overdue high-risk actions, and resilience investment.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
--- PROPOSAL 3 (agent qwen3.8-max_initial_3, alibaba/qwen3.8-max) ---
Estimated complexity: high
Success metrics: - **MTTD reduced from 22 minutes to <5 minutes** within 6 months of full rollout.
- **Customer-first detection rate reduced from 40% to <10%** within 6 months.
- **MTTM for SEV1 incidents reduced from 3 h 10 min to <60 minutes** within 9 months.
- **Monthly SLA credit payouts reduced from ~$325K to <$100K per quarter** within 12 months.
- **Alert volume reduced from 3,400/month to <600 actionable alerts/month** within 6 months; signal-to-noise ratio >80%.
- **Postmortem completion rate: 100% of SEV1/SEV2 incidents** have a blameless postmortem within 5 business days.
- **Postmortem action-item completion rate >90% within 30 days** of the postmortem (up from ~17%).
- **Zero incidents with >15 minutes of unowned command** (down from 2 incidents with >1 hour).
- **100% on-call coverage**: all 28 teams staffed with primary + secondary on-call 24×7 within 14 weeks.
- **On-call compensation adopted**: 100% of on-call engineers receiving stipends and page pay; on-call satisfaction score ≥4/5 in quarterly survey.
- **Status-page first-update within 15 minutes for SEV1 and 30 minutes for SEV2**, 100% compliance.
- **SOC 2 Type II audit passed** at month 8 with zero incident-response findings.
- **All 260 engineers trained**; 56+ certified ICs and 28+ certified CLs active within 14 weeks.
- **Six alerting tools consolidated to one** within 6 months; legacy tools decommissioned.
- **Regulator notification process tested**: at least one tabletop exercise includes a NY DFS / FinCEN notification drill, and the legal/compliance playbook is documented and approved.
- **Quarterly IMC reviews held consistently** with published KPI dashboards and action-item tracking.
- **On-call participation resistance resolved**: <10% of engineers report 'unwilling to participate' in the 6-month pulse survey (baseline to be measured in S1).
Steps (13):
1. Assess Current State and Baseline Metrics
Build an evidence-based picture of the current incident management reality before designing anything new.
Collect and catalog the last 12 months of incident data: all 31 customer-impacting incidents, 3,400 monthly alerts, on-call coverage gaps across the 28 teams, and the 11 of 64 closed postmortem action items. Interview one lead from each of the 28 teams to surface pain points, political concerns (the 'carrying a pager for other teams' pushback), and tool sprawl.
Deliverables to produce:
- **Alert inventory**: which of the six alert tools feed which teams, alert volume per team, noise rate per tool, and overlap between tools.
- **Incident timeline analysis**: median detection-to-notification-to-mitigation-to-resolution times, who detected first (internal vs. customer), who mitigated, and where handoff gaps occurred.
- **On-call coverage map**: which 12 teams have on-call, which 16 do not, rotation length, compensation status, and escalation paths (or lack thereof).
- **Postmortem audit**: format variance, action-item tracking gaps, and the two incidents with no clear owner for over an hour.
- **Tooling and integration audit**: Kubernetes observability stack, six alerting tools, status-page provider, communication channels (Slack, email, phone), and any existing runbooks.
- **Compliance gap analysis**: SOC 2 Type II CC7.3/CC7.4 requirements vs. current practice, with a risk register for the eight-month window.
- **Peer benchmarking**: incident management practices at 3–4 comparable B2B fintech platforms (e.g., Plaid, Stripe, Adyen) for severity scales, on-call comp, and MTTR targets.
2. Secure Executive Sponsorship and Form the IM Governance Body (depends on: 1)
Anchor the program with visible, top-down authority so that 28 teams adopt changes they did not individually request.
The CEO email about 'outages we hear about from clients' is a ready-made mandate. Convert it into a formal sponsorship structure.
- Appoint an **executive sponsor** (CTO or VP Engineering) who owns the program end-to-end and reports to the CEO monthly.
- Create an **Incident Management Office (IMO)**: one dedicated senior incident management lead, one tooling/platform engineer, and one part-time data analyst.
- Establish an **Incident Management Council (IMC)**: one engineering manager from each of the 28 teams, plus the VP of Customer Success, a compliance lead, and a security lead. The IMC meets bi-weekly during rollout, monthly thereafter.
- Draft and circulate an **executive mandate memo** that states: incident response is a shared operational obligation, not a per-team favor; participation in on-call rotations is a condition of employment for production-facing roles; and the program is not optional pending the SOC 2 audit.
- Allocate a dedicated budget line for on-call compensation, tooling consolidation, status-page licensing, training, and external facilitation.
3. Define Severity Levels and Automatic Triggers (depends on: 1)
Replace the current ad-hoc triage with a five-level severity taxonomy that every engineer, support agent, and account manager can apply in under 60 seconds.
- **SEV1 – Critical**: Ledger data corruption or loss, complete payment processing halt, confirmed data breach affecting customer PII or funds, or regulatory reporting breach. Triggers: automatic all-hands page to on-call, CTO + CEO paged within 10 minutes, dedicated bridge call within 5 minutes, status-page update within 15 minutes, regulator notification assessment within 1 hour, customer comms within 30 minutes.
- **SEV2 – High**: Payment processing degraded >30% throughput or >5% error rate, single-region failover failure, ledger read-only mode, or any condition likely to breach the 99.95% SLA within the current window. Triggers: primary + secondary on-call paged, incident commander assigned within 10 minutes, bridge call within 15 minutes, status-page update within 30 minutes, VP Engineering notified within 20 minutes.
- **SEV3 – Medium**: Non-critical service degradation affecting <30% of customers, single non-ledger microservice outage with working fallback, or elevated latency above SLA threshold but below halt. Triggers: primary on-call paged, IM notified within 30 minutes, status-page update within 1 hour if customer-visible, daily standup update.
- **SEV4 – Low**: Degraded internal tooling, minor UX bug with workaround, non-customer-facing alert. Triggers: next-business-day response, ticket created, no page unless on-call agrees.
- **SEV5 – Informational / Noise**: Cosmetic issues, planned-maintenance notifications, alert misfires logged for tuning. Triggers: no page, logged for weekly alert-quality review.
Define **escalation rules**: any SEV3 unresolved after 4 hours auto-escalates to SEV2; any SEV2 unresolved after 2 hours auto-escalates to SEV1. Severity can be **downgraded** only by the incident commander with IMC notification.
Publish the taxonomy as a one-page decision tree, a Slack slash-command (`/sev`), and an integration into the alerting tool so that every alert carries a suggested severity.
4. Define Incident Roles and Staffing Model (depends on: 3)
Codify four mandatory roles for every SEV1/SEV2 incident and optional roles for SEV3, then solve the 24×7 staffing problem across 28 teams.
**Roles**
- **Incident Commander (IC)**: owns the incident end-to-end, declares severity, assigns tasks, authorizes mitigations, decides when to escalate or stand down. Never writes code during the incident.
- **Communications Lead (CL)**: owns status-page updates, internal Slack channels, account-manager briefings, and regulator notifications. Separate from the IC so the IC can focus on mitigation.
- **Scribe / Timeline Keeper**: logs every decision, action, and timestamp in the incident channel and the incident-management tool. Produces the raw timeline for the postmortem.
- **Subject-Matter Responders (SMRs)**: 1–3 engineers from the owning team(s) who diagnose and fix. For the shared PostgreSQL ledger, a dedicated DBA responder is always required.
**24×7 Staffing via a Three-Tier Follow-the-Sun Model**
- **Tier 1 – Front-line on-call**: Primary + secondary responder per team, paged first. Covers the team's own services.
- **Tier 2 – Platform / SRE on-call**: A dedicated 6-person SRE rotation covering cross-cutting infrastructure: Kubernetes, the shared PostgreSQL ledger, networking, and the two AWS regions. This tier directly addresses the 'carrying a pager for other teams' concern by absorbing infrastructure incidents.
- **Tier 3 – IMC escalation**: Engineering managers and the IMO on-call for multi-team or SEV1 incidents. Provides the IC and CL when no team-level IC is available.
**Follow-the-Sun**: If any engineering hub exists in a second timezone, use it for overnight Tier 1 coverage. If not, partner with a managed on-call service for overnight first-response triage (severity declaration + paging the correct team), reducing 3 a.m. pages for NY-based engineers.
**IC and CL pools**: Nominate at least 2 ICs and 1 CL per team (56 ICs, 28 CLs minimum). ICs are trained and certified before they rotate. For SEV1 incidents, the IC must be a certified IC from the IMC pool, not just 'whoever is around'.
**Ledger-specific rule**: Because the PostgreSQL ledger is shared, a **Ledger Duty Officer** from the SRE Tier 2 is always on the bridge for any incident touching ledger services, regardless of which team owns the failing microservice.
5. Design On-Call Rotations, Compensation, and Alert-Quality Rules (depends on: 4)
Make on-call sustainable, fairly compensated, and free of alert noise so engineers stop resisting participation.
**Rotation Design**
- 7-day rotations, one primary + one secondary per team per week. No engineer is on-call more than one week in four.
- Minimum 48-hour rest between rotations. No on-call during approved PTO.
- All 28 teams participate. Teams without current on-call get a 90-day ramp with a shadow rotation before going live.
- Tier 2 SRE rotation: 6 engineers, one week on / five weeks off, with a dedicated backup.
**Compensation Package**
- **Base on-call stipend**: $500 per week of primary on-call, $250 for secondary, paid regardless of whether pages fire.
- **Page pay**: $75 per acknowledged page outside business hours; $150 if the page leads to active incident work.
- **Time-off-in-lieu (TOIL)**: Any engineer who works >4 hours overnight (00:00–06:00 local) gets a full TOIL day. >8 hours in a single incident gets 1.5 TOIL days.
- **SEV1 bonus**: $300 flat bonus for every engineer who actively works a SEV1 incident, paid within the next pay cycle.
- **Annual on-call cap**: No engineer exceeds 13 weeks of on-call per year. Exceeding the cap triggers a mandatory team-staffing review.
- Budget estimate: ~$420K/year for stipends and page pay across 28 teams; present this to the CFO as a fraction of the $1.3M annual SLA credit cost.
**Alert-Quality Rules (the '85% noise' problem)**
- Every alert must carry: owning team, suggested severity, runbook link, and a 30-day noise score.
- **Alert budget**: each team gets a maximum of 100 actionable alerts per month. Exceeding the budget triggers a mandatory alert-tuning session with the IMO.
- **Noise threshold**: any alert that fires >10 times in 7 days with no human action is auto-flagged for suppression or tuning within 14 days.
- **Alert review cadence**: weekly 30-minute alert-quality review per team; monthly cross-team alert review in the IMC.
- **Sunset rule**: alerts with no runbook are demoted to SEV5 after 30 days and suppressed after 60 days unless a runbook is written.
- Target: reduce monthly alert volume from 3,400 to <600 actionable alerts within 6 months.
6. Build Detection, Escalation, and Communication Paths (depends on: 3, 4)
Eliminate the 22-minute median detection gap and the 40% customer-first-detection rate with layered monitoring and a single escalation spine.
**Detection Layers**
- **Synthetic transactions**: run a payment end-to-end through the full stack (API → service → ledger → confirmation) every 60 seconds from both AWS regions. Alert if latency >2× baseline or any step fails. This catches what per-service metrics miss.
- **Customer-traffic anomaly detection**: monitor API error rates, payment success rates, and latency percentiles per customer cohort. Alert on >2σ deviation.
- **SLO-based alerting**: define SLIs for the 99.95% SLA (availability, latency p99, ledger consistency). Alert when error budget burn rate exceeds threshold, before the SLA actually breaches.
- **Infrastructure health**: Kubernetes node/pod health, PostgreSQL replication lag, disk I/O, and cross-region latency.
- **Support-ticket spike detection**: if >5 customers open tickets about the same symptom within 10 minutes, auto-create a SEV3 candidate.
**Escalation Path**
- Alert fires → PagerDuty routes to Tier 1 primary → 5-min no-ack → Tier 1 secondary → 10-min no-ack → Tier 2 SRE → 15-min no-ack → IMC on-call manager → 20-min no-ack → VP Engineering auto-page.
- Any SEV1 declaration auto-pages the CTO, opens a dedicated Slack channel + Zoom bridge, and notifies the CL.
- **No incident goes unowned for >15 minutes.** If no IC is assigned by minute 15, the IMC on-call manager assumes IC role by default.
**Internal Communications**
- Dedicated Slack channels: `#inc-sev1`, `#inc-sev2`, `#inc-sev3` (auto-created per incident), plus `#inc-updates` for broadcast.
- IC posts a structured update every 15 minutes (SEV1), 30 minutes (SEV2), 1 hour (SEV3) into the incident channel.
- CL posts a summary to `#inc-updates` and notifies relevant engineering managers.
**Customer Communications**
- **Status page**: auto-updated via API. SEV1: first update within 15 minutes, then every 30 minutes until resolved. SEV2: first update within 30 minutes, then every hour. SEV3: within 1 hour if customer-visible.
- **Account managers**: CL briefs AMs via a dedicated Slack channel within 30 minutes (SEV1) or 1 hour (SEV2). AMs contact their top-20 revenue accounts directly.
- **Customer email/SMS**: for SEV1 and SEV2, automated notification to all 2,100 customers via the status-page subscription system within 30 minutes.
- **Regulator notification**: Legal/Compliance assesses within 1 hour whether NY DFS, FinCEN, or card-network notification is required. If yes, file within the regulatory deadline (typically 72 hours for NY DFS cybersecurity events). Log the decision and filing in the incident record.
**Post-resolution**: CL publishes a 'resolved' update within 30 minutes of mitigation. For SEV1/SEV2, a preliminary customer-facing RCA summary is published within 5 business days.
7. Standardize Postmortems with Tracking and Accountability (depends on: 3, 6)
Fix the '11 of 64 action items closed' problem with a mandatory, uniform, blameless postmortem process backed by engineering-manager accountability.
**When Mandatory**
- All SEV1 and SEV2 incidents: postmortem required within 5 business days.
- SEV3 incidents: postmortem required if >50 customers affected, if the incident lasted >4 hours, or if it was customer-detected.
- SEV4/SEV5: optional, but any recurring SEV4 (≥3 times in 30 days) triggers a mandatory review.
**Format (single template, enforced by tooling)**
- Incident summary (severity, duration, customers affected, revenue impact, SLA credit exposure).
- Timeline (auto-generated from scribe notes + alert timestamps).
- Detection analysis: how was it detected, why did it take X minutes, could it have been faster.
- Root cause analysis using 5-Whys or fault-tree, not blame.
- Contributing factors (process, tooling, staffing, knowledge gaps).
- Impact quantification: customers affected, transactions failed, SLA credits triggered.
- Action items: each with a **named owner**, **due date**, **priority**, and **ticket in Jira**.
- Lessons learned and what went well.
**Blameless Review Meeting**
- Held within 5 business days, facilitated by the IMO or a trained facilitator (never the IC of that incident).
- All responders, the CL, relevant engineering managers, and an IMC representative attend.
- Ground rules: focus on system and process failures, not individual mistakes. The facilitator enforces this.
- Meeting recorded; notes published to the engineering-wide wiki within 48 hours.
**Action-Item Tracking and Accountability**
- Every action item is created as a Jira ticket with a due date and a named owner.
- **Engineering managers are accountable**: action-item completion is a standing agenda item in the bi-weekly IMC meeting. Any action item >7 days overdue is escalated to the VP Engineering.
- **Completion gate**: no team may close its postmortem until 100% of its action items have Jira tickets. Postmortem is 'closed' only when all tickets are resolved.
- **Quarterly audit**: the IMO audits action-item completion rates and reports to the IMC and the executive sponsor. Target: >90% completion within 30 days of the postmortem.
- Link postmortem quality and action-item completion to team health scores and engineering-manager performance reviews.
8. Define KPIs, Dashboards, and Governance Reviews (depends on: 7)
Create a measurable feedback loop so leadership can see whether the program is working and where to intervene.
**Primary KPIs (tracked weekly, reported monthly)**
- **MTTD** (Median Time to Detect): target <5 minutes (from 22).
- **MTTA** (Median Time to Acknowledge): target <5 minutes.
- **MTTM** (Median Time to Mitigate): target <60 minutes for SEV1 (from 3 h 10 min), <4 hours for SEV2.
- **Customer-first detection rate**: target <10% (from 40%).
- **SLA compliance**: maintain 99.95%; track monthly SLA credit payouts, target <$100K/quarter (from ~$325K/quarter).
- **Alert signal-to-noise ratio**: target >80% actionable (from ~15%).
- **Alert volume**: target <600/month (from 3,400).
- **Postmortem completion rate**: 100% for SEV1/SEV2 within 5 business days.
- **Action-item completion rate**: >90% within 30 days (from ~17%).
- **On-call health**: pages per engineer per week (target <5), TOIL usage, on-call satisfaction survey score.
- **Unowned incident duration**: target 0 incidents with >15 minutes without an IC.
**Dashboards**
- Real-time operational dashboard (Grafana): current incidents, active alerts, on-call roster, SLA error-budget burn.
- Weekly leadership dashboard (auto-generated): KPI trends, open action items, alert-noise report, on-call load distribution.
- Quarterly IMC scorecard per team.
**Review Cadence**
- **Weekly**: IMO publishes KPI snapshot to `#inc-updates`.
- **Bi-weekly IMC**: review open incidents, overdue action items, alert-quality exceptions, and on-call load.
- **Monthly executive review**: CTO presents KPI trends, SLA credit cost, and risk register to the CEO.
- **Quarterly incident-management review**: deep-dive into trends, training gaps, tooling needs, and process improvements. Output fed into the next quarter's roadmap.
9. Consolidate Tooling and Build the Incident Management Platform (depends on: 2, 3)
Replace six alerting tools and ad-hoc status-page updates with a single, integrated incident management stack.
**Target Tool Architecture**
- **Single alerting and on-call platform** (e.g., PagerDuty or Opsgenie): ingest all alerts, apply severity routing, manage on-call schedules, handle escalations, and send pages. Retire the other five tools within 6 months.
- **Observability consolidation**: standardize on one APM/metrics stack (e.g., Datadog or Grafana Cloud) for all 180 Kubernetes services across both AWS regions. Ensure the shared PostgreSQL ledger has dedicated dashboards.
- **Status page**: a dedicated, branded status page (e.g., Statuspage.io or Instatus) with API integration for auto-updates. Subscribe all 2,100 customers.
- **Incident coordination tool**: integrate incident-management workflows into Slack (auto-create channels, invite responders, post templates) and a dedicated incident record system (e.g., Jira Service Management, incident.io, or Rootly) for timelines, postmortems, and action-item tracking.
- **Runbook repository**: a central wiki (Confluence or Notion) with mandatory runbooks for every alert. No alert goes live without a linked runbook.
**Implementation Tasks**
- Migrate all 28 teams' alert rules into the single platform in three waves (highest-volume teams first).
- Build the severity-based routing rules and escalation policies per S3 and S6.
- Automate status-page updates triggered by severity declaration.
- Build the synthetic-transaction monitor and SLO-based alerting per S6.
- Integrate Jira for automatic action-item ticket creation from postmortems.
- Decommission legacy tools only after all teams have completed training on the new stack.
- Budget: allocate $150K–$250K/year for licensing, plus engineering time for migration.
10. Prepare for the SOC 2 Type II Audit (depends on: 7, 8, 9)
Ensure the incident management process produces the evidence the auditor will need, well before the audit window opens in eight months.
**SOC 2 Requirements to Address (CC7.3, CC7.4, CC7.5)**
- Documented incident response procedures (the severity taxonomy, role definitions, communication templates).
- Evidence of incident detection, response, and recovery for every SEV1/SEV2 incident during the audit period.
- Postmortem records with action-item tracking.
- On-call schedules, training records, and escalation evidence.
- Status-page update logs and customer notification records.
- Regulator notification logs (if any).
**Preparation Tasks**
- The IMO maintains a **SOC 2 evidence folder**: every incident record, postmortem, action-item ticket, status-page update, and training completion certificate is stored and indexed.
- Conduct a **mock SOC 2 audit** at month 5: an internal or external auditor reviews the incident management process end-to-end and identifies gaps.
- Remediate mock-audit findings before month 7.
- Ensure the incident management tool retains all records for at least 12 months (the SOC 2 Type II observation window).
- Document the **chain of custody** for incident records: who accessed, modified, or closed each record.
- Prepare a **narrative document** describing the incident management process, roles, and controls for the auditor.
- Coordinate with the compliance lead to align incident management evidence with the broader SOC 2 scope (access controls, change management, etc.).
11. Design and Deliver Training, Runbooks, and Change Management (depends on: 4, 5, 9)
Equip all 260 engineers, 28 team leads, account managers, and support staff with the knowledge and muscle memory to execute the new process.
**Training Tracks**
- **All 260 engineers** (2-hour session): severity taxonomy, how to acknowledge a page, how to join an incident bridge, how to hand off to an IC, and how to write a postmortem contribution. Delivered in team-level sessions over 4 weeks.
- **IC pool (56+ engineers)** (8-hour certification): incident command techniques, severity declaration, escalation decision-making, bridge facilitation, and blameless postmortem facilitation. Includes two tabletop exercises. Certification valid for 12 months, renewed annually.
- **CL pool (28+ staff)** (4-hour session): status-page writing, customer communication templates, regulator notification triggers, and AM briefing protocol.
- **Account managers and support staff** (1-hour session): how to read the status page, how to escalate a customer report into an incident, and what information to collect.
- **SRE Tier 2** (16-hour onboarding): Kubernetes and PostgreSQL ledger deep-dive, cross-region failover runbooks, and escalation authority.
**Runbooks**
- Every alert must have a runbook before it is routed to on-call. The IMO provides a runbook template and audits compliance weekly.
- Priority runbooks to write first: shared PostgreSQL ledger failover, Kubernetes cluster degradation, payment-processing pipeline failure, cross-region failover, and ledger data-integrity check.
- Runbooks are peer-reviewed and version-controlled.
**Change Management for Adoption**
- Address the 'carrying a pager for other teams' concern directly: publish an FAQ explaining the three-tier model, the SRE Tier 2 absorbing cross-team infrastructure, the compensation package, and the TOIL policy.
- Run **office hours** weekly for the first 8 weeks where any engineer can ask questions or raise concerns.
- Identify **team champions**: one engineer per team who volunteers as an early adopter and peer mentor.
- Publish a **weekly 'incident management newsletter'** during rollout: what changed, what improved, KPI trends, and success stories.
- Make on-call participation a documented expectation in job descriptions and performance reviews for production-facing roles.
12. Execute Phased Rollout, Tabletop Exercises, and Continuous Improvement (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11)
Introduce the process in three waves so teams are not overwhelmed, then validate with exercises and iterate continuously.
**Phase 1 – Weeks 1–6: Foundation**
- Publish the severity taxonomy, role definitions, and communication protocols (S3, S4, S6).
- Launch the single alerting platform for the 12 teams already on-call; begin migration for the other 16.
- Activate the SRE Tier 2 rotation for the shared PostgreSQL ledger and cross-cutting infrastructure.
- Deploy the status page and test the API integration.
- Begin IC and CL training (first cohort of 20 ICs, 10 CLs).
- Publish the on-call compensation package; HR integrates stipends into payroll.
- Write the top 10 priority runbooks.
**Phase 2 – Weeks 7–14: Expansion**
- All 28 teams live on the single alerting platform; legacy tools in read-only mode.
- All 28 teams on the on-call rotation schedule (the 16 new teams in shadow mode for the first 4 weeks).
- Second and third IC/CL training cohorts completed.
- First **tabletop exercise**: simulate a SEV1 ledger corruption scenario with all roles, test the escalation path, status-page updates, and AM briefings. Debrief and fix gaps.
- Postmortem template and Jira integration live; all new incidents use the standard process.
- Alert-tuning sprint: each team reduces its alert volume by 50%.
**Phase 3 – Weeks 15–24: Optimization**
- All teams fully live; legacy alerting tools decommissioned.
- Second **tabletop exercise**: simulate a SEV1 cross-region failure with regulator notification.
- First quarterly IMC review with full KPI dashboard.
- Mock SOC 2 audit (month 5) and remediation.
- Retrospective on the rollout: survey all 260 engineers for feedback, adjust compensation or rotation rules if needed.
- Establish the **continuous improvement cadence**: quarterly process review, annual severity-taxonomy review, and annual on-call compensation benchmarking.
**Ongoing Governance**
- The IMC owns the process document and approves changes.
- The IMO tracks all KPIs and reports to the CTO monthly.
- Any process change requires IMC approval and a 2-week notice period before enforcement.
- Annual external benchmarking against peer B2B payments platforms.
13. Establish Ongoing Governance, Annual Review, and Audit Readiness Cycle (depends on: 12)
Embed incident management as a permanent organizational capability, not a one-time project.
- **Annual process review**: the IMC reviews the severity taxonomy, role definitions, on-call structure, and compensation against industry benchmarks and internal KPIs. Update as needed.
- **Bi-annual tabletop exercises**: one SEV1 infrastructure scenario, one SEV1 data-breach/regulator scenario. Rotate the IC and CL assignments so everyone gets practice.
- **Quarterly alert-quality audit**: the IMO reviews alert volumes, noise rates, and runbook coverage across all 28 teams.
- **On-call health survey**: quarterly anonymous survey measuring burnout, fairness, and compensation satisfaction. Results reviewed by the IMC.
- **SOC 2 readiness cycle**: begin evidence collection immediately after each audit ends. The IMO maintains a rolling evidence folder. Mock audit at month 5 of every 12-month audit cycle.
- **Postmortem maturity tracking**: track the action-item completion rate monthly. If it drops below 80%, the VP Engineering intervenes.
- **Incident management maturity model**: adopt a 5-level maturity model (ad-hoc → defined → managed → optimized → predictive). Assess annually. Target: Level 3 within 12 months, Level 4 within 24 months.
- **Budget review**: annually review on-call compensation, tooling costs, and training budget against the reduction in SLA credits and incident frequency.
--- PROPOSAL 4 (agent grok4.6_initial_4, xai/grok-4.6) ---
Estimated complexity: high
Success metrics: - Median time to detect customer-impacting incidents ≤ 5 minutes within 6 months of go-live.
- Share of SEV-1/SEV-2 incidents first detected by customers ≤ 5% (from 40%).
- Median time to mitigate SEV-1/SEV-2 ≤ 45 minutes (from 3 h 10 min).
- Named Incident Commander assigned within 5 minutes for ≥ 95% of SEV-1/SEV-2.
- First status-page update within policy time for ≥ 95% of SEV-1/SEV-2.
- SLA credits down ≥ 80% versus the trailing $1.3M within 12 months.
- Paging volume ≤ 500 per month and noise ≤ 15% (from 3,400 and 85%).
- 100% of production services have a named owning team and a paging policy.
- Postmortems filed within 5 business days for 100% of SEV-1/SEV-2; action-item close rate ≥ 80% within 30 days.
- 24×7 IC and critical-path coverage with zero unfilled shifts per quarter.
- Paid on-call live for every rotation before that rotation pages humans.
- SOC 2 Type II incident-response controls evidenced for ≥ 5 months before the auditor's report.
- On-call pulse: ≥ 70% of engineers agree rotations are fair and limited to their services.
Steps (30):
1. Secure executive mandate and budget
Get a written CEO/CTO mandate that incident command is a company process, not a team hobby.
The mandate must state that **paid on-call** is required for production ownership. "No pager for other teams' code" is solved by named ownership, not by refusing coverage.
- Approve budget for tooling, stipends, training, and a dedicated program lead for six months.
- Name an executive sponsor (CTO or VP Engineering) who will chair the weekly incident review.
- Tie the clock to SOC 2 Type II: the process must be live in about 10 weeks so ~6 months of evidence remain.
- Commit that the CEO will hear about outages from this process, not from customers.
2. Form the working group and decision rights (depends on: 1)
Stand up a small group that can decide. Do not form a 28-team committee.
**Core seats:** SRE/platform lead, payments/ledger engineering manager, support lead, legal/compliance, HR, one rotating team EM, and a program manager.
- Meet twice a week for 10 weeks, then weekly.
- RACI: the group proposes; the sponsor decides in 48 hours; teams implement.
- Publish one Slack channel and one source-of-truth doc on day one.
- Time-box design to four weeks. Ship v1 rather than wait for consensus.
3. Inventory services, owners, and on-call gaps (depends on: 1)
Build a living catalog of all ~180 services: owning team, criticality, current on-call, alert sources, and runbook link.
Walk the last 31 customer-impacting incidents and the two events where **nobody was in charge**. Record who detected, who led, time to mitigate, and which alerts fired.
- Tag each service as critical-path, customer-visible, or internal.
- List the 16 teams with no on-call and every orphan service with no owner.
- Map all six alerting tools and the 3,400 monthly alerts onto services.
- Flag the shared PostgreSQL ledger and two-region failover as named special cases.
4. Map regulatory and contractual notification duties (depends on: 1)
Legal and compliance list every duty an incident can trigger. Do not invent clocks that violate a contract.
Cover SOC 2 CC7, customer MSA/SLA credit terms, money-transmitter rules, NYDFS 23 NYCRR 500 if applicable, PCI if in scope, and breach clocks.
- Extract **notification timings** from the largest customer contracts (status page, named AM, written notice).
- Define when legal, regulators, insurers, or the board must be told.
- Feed these clocks into severity triggers and the communications playbook.
5. Approve paid on-call and incident pay (depends on: 1, 4)
Unpaid on-call is why 16 teams refuse the pager and why nights are uncovered. Fix the money before asking for coverage.
HR, legal, and finance design a New York–compliant package: weekly stipend for primary and secondary, extra stipend for company IC and comms, and **after-hours incident pay** or comp time.
- Treat exempt vs non-exempt staff explicitly under NY wage-hour rules.
- Put stipend in the next pay cycle after policy publish, not "later."
- Cap consecutive night weeks. Fund hiring if a team cannot rotate fairly (minimum six people for 24x7 primary plus secondary).
- Publish the package before any new rotation starts. This is the main answer to pager pushback.
6. Ratify four severity levels and their triggers (depends on: 3, 4)
Adopt a business-impact scale. Engineers do not invent severity in the moment.
**SEV-1:** material payments failure; ledger down or inconsistent; security or customer-data incident; both regions impaired; or many customers already in SLA-credit territory.
**SEV-2:** degraded payments or a contracted feature down for multiple customers; SLA at risk.
**SEV-3:** narrow or single-customer impact with a workaround; no fleet-wide SLA risk.
**SEV-4:** no customer impact; ticket only.
- SEV-1 pages IC, comms, scribe, owning SMEs, and an exec; war room in 5 minutes; status page in 10; AM outreach in 20.
- SEV-2 pages IC and owning SMEs; comms may be the IC; status page in 15 minutes; updates every 30 minutes.
- SEV-3 pages the owning team only; customer notice only if that customer is affected.
- Anyone may declare. Only the IC may downgrade. When unsure, start high.
7. Define incident roles and their authority (depends on: 6)
Four roles. Separate coordination from debugging so "who is in charge" cannot stall for an hour again.
- **Incident Commander:** owns severity, the room, the clock, and the next action. Does not write code. May page anyone, freeze deploys, and invoke failover. Staffed from a company-wide trained pool, not from the failing team.
- **Communications Lead:** status page, customers, AMs, execs, regulators. Speaks only from IC-approved facts.
- **Scribe:** timeline in the incident tool. Required for SEV-1 and SEV-2.
- **SME responders:** the owning team's on-call. They mitigate. They do not run the room.
Publish a one-page authority card. The IC stays in charge even if a VP joins.
8. Design 24x7 coverage without 28 night rotations (depends on: 3, 5, 7)
Do not put 28 teams on 24x7. That is what engineers are rejecting.
Use **three layers** so people page for their own code, plus a trained commander.
- Layer A — company IC and SEV-1 comms: 24x7; about 24 trained people; week-long primary and secondary.
- Layer B — critical-path team on-call (ledger, payments processing, auth, API edge, platform/Kubernetes, data stores): 24x7 primary plus secondary.
- Layer C — all other teams: business-hours on-call; after hours the IC pages the team EM, who has a written escalation list.
Platform on-call is the safety net for unknown-owner pages, never the permanent owner. Every service must have a named team within 60 days or be scheduled to shut off.
9. Set rotation, handoff, and load rules (depends on: 8)
Write mechanical rules so rotations are fair and load is visible.
Primary week, then secondary week, then at least two weeks off. No one holds primary on two rotations at once.
- Handoff is a 30-minute overlap covering open incidents, silenced alerts, and upcoming changes.
- Page-load SLO: p50 ≤ 4 pages per 12-hour night shift; p95 ≤ 10. A breach opens an alert-quality action.
- Require a shadow week before a first IC shift or a first critical-path rotation.
- Swaps live in the paging tool. Managers own coverage gaps, not the last person on the roster.
10. Write detection and escalation paths (depends on: 6, 9)
Customers currently detect 40% of incidents and median time to detect is 22 minutes. That is the first failure mode.
Detection path: synthetic full-payment probes in both regions, SLO burn-rate alerts, support-to-incident intake, and one customer callback path that can create a SEV.
- A page must be acked in **5 minutes** or it auto-escalates to secondary, then IC, then the EM, then the VP.
- Support may declare SEV-2 or higher without engineering permission.
- If ownership is unclear for 10 minutes, the IC keeps the incident and assigns a temporary owner. Never wait.
- An exec bridge auto-opens for every SEV-1 at T+15 minutes.
11. Write internal, customer, and regulator communications (depends on: 4, 6, 7)
Stop "whoever is around" from writing the status page. Comms follow the clock, not convenience.
Timings from declaration:
- Internal war room: immediate. Exec summary for SEV-1/2 at 15 minutes, then every 30 minutes.
- **Public status page:** SEV-1 in 10 minutes, SEV-2 in 15. Updates at least every 30 minutes until resolve. Templates only. No speculation.
- Account managers get an affected-customer list and a script at T+20 minutes for SEV-1/2.
- Resolve notice and credit assessment within one business day.
- Comms pages legal on SEV-1 security, ledger integrity, or any outage that will breach contractual notice. Legal owns outbound regulatory letters; the IC owns facts.
12. Standardize blameless postmortems and action tracking (depends on: 6)
A written postmortem is mandatory for every SEV-1 and SEV-2 within 5 business days. SEV-3 if the IC or EM requests it.
Use one template: timeline, customer impact (volume, duration, credits), detection gap, what went well, what did not, process-focused five whys, and numbered actions with owner and due date.
- Review is **blameless** and scheduled. The IC attends. The exec sponsor reads every SEV-1.
- Actions live in one tracker, not in the doc. No action without an owner and a date. Default due date 14 days; 30 days max unless architecture work with a milestone.
- Close rate is a published metric. The old 11-of-64 pattern is a process failure.
13. Set alert quality rules that make paging acceptable (depends on: 3, 6)
3,400 alerts a month and 85% noise is why on-call feels like punishment. Pages are a product with a quality bar.
A page (not a ticket) must map to a customer-facing SLO or a hard dependency of one. It must have an owner team, a runbook link, and a default severity. It must be actionable at 3 a.m. by the person who is paged.
- Ban parallel paging from six tools. One paging policy: symptom-based; burn-rate preferred over raw thresholds.
- Every team gets a monthly noise budget. Exceeding it is a sprint task, not heroics.
- A human may silence a flapping alert only with a linked ticket.
14. Publish Incident Management Policy v1 (depends on: 5, 6, 7, 8, 9, 10, 11, 12, 13)
Collapse the design into a short policy people will open during an outage.
Ten pages or fewer, plus one-page cards for severity, roles, and comms timings. Host it where the incident tool can link it.
- Include the compensation summary and the rule: you are **not on-call for other teams' services**.
- Version it. v1 is mandatory from the pilot start date.
- Legal, HR, and the exec sponsor sign. Announce in all-hands, not only in Slack.
15. Implement a single incident command tool (depends on: 7, 10, 11)
Put one tool in the path that creates the room, pages roles from severity, records the timeline, and prompts status-page updates.
Requirements: Slack (or equivalent) incident bot, severity in one click, role assignment, stakeholder groups, and timeline export for postmortems and auditors.
- Integrate with the pager so IC, comms, and SME pages are automatic.
- Retain artifacts at least one year for SOC 2.
- Ad-hoc Zoom/Slack threads are no longer the system of record.
16. Consolidate six alerting tools onto one pager (depends on: 9, 13)
Pick one paging product. Connect existing monitors to it. Migrate **pages** first, tickets second.
- Inventory every page-producing rule. Delete or downgrade the noisy majority in S23.
- Route by service label → owning team schedule → the escalation policy from S10.
- IC and comms schedules live in the same product.
- Set a hard date after which pages outside the chosen tool are not valid on-call obligations.
17. Operationalize the status page and AM path (depends on: 11, 15)
Put the status page behind the Comms Lead role. Use templates for investigating, identified, mitigating, and resolved.
Subscribe AMs and customers whose contracts require it. Generate the affected-customer list from the incident (tenant, region, payment method).
- Dry-run a SEV-2 update before the pilot goes live.
- Record every public update in the incident timeline for the audit.
- Define partial vs full-outage wording so impact cannot be understated.
18. Stand up one postmortem repo and action board (depends on: 12)
Create the template, the filing location, and one Jira/Linear board with states: open, in progress, blocked, done, and won't-do (with exec reason).
Wire the incident tool so a SEV-1/2 automatically opens a draft postmortem and action tickets.
- Each week the program manager reports actions past due to the exec sponsor.
- If it is not in the board, it does not exist.
19. Assign owners and write critical-path runbooks (depends on: 3, 9)
Close the ownership gaps that cause "pager for other people's code."
Every production service gets a team in the catalog. Unowned services get an owner in 30 days or a decommission date.
- Write runbooks for ledger Postgres, regional failover, payments API, auth, and the Kubernetes control plane: symptoms, dashboards, mitigate vs escalate, customer impact.
- Link runbooks from alerts. If there is no runbook, the alert cannot page at night unless the EM accepts the gap in writing.
20. Detect payments failures before customers do (depends on: 3, 13)
Build or finish end-to-end synthetics: create payment, ledger write, webhook, both regions, both critical payment methods.
Alert on SLO burn, not on a single 500. Page SEV-2 or SEV-1 from these probes. This is the fastest lever on the 22-minute MTTD and the 40% customer-detected rate.
- Add ledger lag, replication, disk, and failover-readiness as first-class pages to the ledger team.
- Every postmortem asks first: **why did a customer see this first?**
21. Train the first cadre of ICs, comms leads, and scribes (depends on: 7, 14, 15)
Train about 24 ICs and 12 comms leads before the pilot. Classroom plus a recorded shadow of a simulated SEV-1.
Curriculum: severity, authority card, tool, comms timings, when to call legal, how to run a room of 20, how to hand off at 2 a.m.
- Certification: pass a tabletop. No certificate, no rotation.
- Recertify yearly and after any SEV-1 where process failed.
- Managers of ICs protect calendar time. This is part of the job.
22. Pilot on the payments critical path for six weeks (depends on: 5, 14, 15, 16, 17, 18, 19, 21)
Go live with policy, tool, paid rotations, and IC coverage for ledger, payments, API, platform, and support intake.
Keep old paths as backup for one week, then cut over. Real incidents use the new process only.
- Staff the program lead in every SEV-2+ as coach, not as secret IC.
- Collect friction daily. Fix tooling and wording in 48 hours.
- Expansion gate: named IC in under 5 minutes, first status update on time, no unpaid pages, postmortem filed.
23. Cut alert noise with a forced burn-down (depends on: 13, 16)
Give every team a numbered list of their noisiest alerts. Move noise from 85% to **under 15%**, and monthly pages from 3,400 toward 500.
Each sprint, critical-path teams must delete, debounce, or convert to ticket a fixed quota. Platform provides burn-rate and grouping libraries.
- Publish a weekly noise leaderboard. Shame systems, not people.
- After eight weeks, any page without a runbook or with >30% false pages in 14 days is auto-downgraded until fixed.
24. Roll remaining teams onto the model by risk (depends on: 22)
After the pilot gate, add teams in waves of four to five every two weeks. Highest customer-impact first.
Layer C teams get business-hours schedules and the EM night list. Do not surprise anyone with a pager.
- Each wave: ownership confirmed, alerts routed, runbooks for paging alerts, paid rotation in HR, one tabletop.
- Finish all 28 teams at least five months before the SOC 2 report date so the observation window covers the company.
- Orphan services still unowned at wave end are escalated to the sponsor for shutdown or reassignment.
25. Run tabletops and multi-region game days (depends on: 21, 22)
Schedule a monthly tabletop: SEV-1 ledger, SEV-1 region loss, SEV-2 degraded payments, customer-detected incident, and a "who is in charge" chaos drill.
Quarterly game day: fail a region or a ledger replica in staging or a controlled production drill.
- Include support, AMs, legal, and an exec. Process fails if only engineers show up.
- Capture actions on the same board as real postmortems.
- Use results as SOC 2 evidence that IR is tested.
26. Resolve ownership fights and pager culture (depends on: 5, 8, 22)
Treat pushback as design input, not defiance. Repeat the contract in office hours: you carry a pager for **your** services; IC is coordination; nights are paid; noise is a defect.
- EMs who cannot staff a fair rotation get headcount or have services reassigned. Do not run two-person 24x7.
- Publicly close the two historical "nobody in charge" incidents with what would be different now.
- Pulse-survey on-call at 60 and 120 days. If load or fairness is red, stop expansion until fixed.
27. Launch metrics, weekly review, and error budgets (depends on: 15, 22)
The CEO email exists because there was no operating rhythm. Stand up a dashboard and use it.
Track incident count by SEV, MTTD, MTTA, MTTM, customer-detected %, time-to-IC, time-to-status, pages per shift, noise %, SLA credits, postmortem SLA, and action close rate.
- **Weekly 30-minute review** chaired by the exec sponsor. Last week's incidents, open actions, alert SLO breaches. Mandatory for EMs of teams that had SEV-2+.
- Error budgets on critical SLOs: burn too fast and the team stops features and pays down reliability.
- Monthly report to the CEO so customer-first news does not recur.
28. Collect SOC 2 Type II evidence as a side effect (depends on: 4, 14, 22, 27)
Do not create a parallel "audit process." The real process is the evidence.
From pilot start, retain policy versions, on-call schedules, incident records, status-page history, postmortems, action tickets, training attendance, tabletop notes, and access reviews of the incident and pager tools.
- Map each artifact to SOC 2 CC7.2–CC7.5. Compliance owns the matrix; engineering owns completeness.
- Internal audit or a consultant runs readiness reviews at week 12 and week 20, not the week before the auditor.
29. Inspect and adapt the process after 90 days (depends on: 22, 27)
After 90 days live, revise v1 using data, not opinions.
Look at MTTD/MTTM, uncovered shifts, IC burnout, missed status updates, action close rate, and survey results.
- Change severity definitions if teams inflate or deflate them.
- Adjust Layer B vs C membership from real pages.
- Cut steps nobody follows. Add only what the last 90 days proved missing.
- Freeze v2 as the SOC 2 described process unless a defect is found.
30. Rehearse the audit and freeze the observation window (depends on: 24, 28, 29)
Six to eight weeks before the auditor, run a mock interview with an IC, a random engineer, support, and compliance.
Show one SEV-1 and one SEV-2 end-to-end: detect, page, command, comms, postmortem, closed actions.
- Fix evidence gaps immediately.
- Freeze process wording for the rest of the observation window; log exceptions.
- Brief the CEO and customer success with the metrics so the "outages we hear about from clients" story is retired.
--- PROPOSAL 5 (agent deepseek-v4-pro_initial_5, deepseek/deepseek-v4-pro) ---
Estimated complexity: high
Success metrics: - Customer-detected incidents decrease from 40% to less than 15% within six months.
- Median time to detect (MTTD) is under 5 minutes for SEV1 and SEV2 incidents.
- Median time to mitigate (MTTM) is under 60 minutes for SEV1 and under 2 hours for SEV2.
- Alert noise decreases from 85% to below 10% within three months.
- 100% of SEV1 and SEV2 incidents have a completed blameless postmortem within 5 business days.
- 100% of postmortem action items are tracked with an owner and due date; 90% are completed on time.
- 24x7 on-call coverage achieved across all 28 teams with no unpaid on-call.
- 99% of pages are acknowledged within 5 minutes.
- SLA credits paid reduce by at least 50% over the next 12 months.
- SOC 2 readiness: all incident response controls are documented, tested, and evidence is produced by month 7.
Steps (14):
1. Baseline current incident response and align stakeholders
Collect data from the last 12 months of incidents, all six alert tools, on-call practices, and team interviews. Identify gaps against the target incident management process and secure executive sponsorship.
- Gather incident timeline, detection source, mitigation time, customer impact, SLA credit, and postmortem status for all 31 incidents.
- Survey 28 teams on on-call burden, alert quality, and operational pain.
- Map current tools, escalation paths, and communication workflows.
- Create baseline metrics and a stakeholder map with executive sponsor and audit owner.
2. Define severity levels and response triggers (depends on: 1)
Define a four-level severity scale with objective business-impact criteria so any engineer can classify an incident consistently.
- SEV1: widespread transaction processing outage, data breach, security incident, or severe SLA breach; triggers full incident command, executive notification, and a 5-minute status page update.
- SEV2: major feature outage, significant degradation without workaround, or customer financial risk; triggers incident commander, full communications role, and status page updates.
- SEV3: partial impairment with workaround or limited customer impact; triggers on-call response, internal communication, and optional status page update.
- SEV4: minor or internal issue, no customer impact; handled during business hours through ticketing.
- Include an escalation matrix showing who can declare, downgrade, and invoke regulatory or legal involvement.
3. Define incident roles and decision authority (depends on: 1, 2)
Define incident roles, responsibilities, and decision authority using RACI to remove ambiguity about who is in charge.
- Incident Commander: owns the incident, declares severity, and coordinates resolution.
- Communications Lead: owns internal and external messaging, status page updates, and account manager notifications.
- Scribe: maintains timeline, incident log, and postmortem notes.
- Subject-matter responders: diagnose and fix the incident; may come from multiple teams.
- Executive sponsor: optional for SEV1; customer liaison: handles account managers.
- Define decision rights for severity declaration, escalation, rollback, customer communications, and incident closure.
4. Design 24x7 staffing model across 28 teams (depends on: 3)
Design 24x7 coverage across 28 teams without overloading engineers. Use service-based on-call plus a central incident command pool.
- Each service or domain team assigns primary and secondary on-call for its own services.
- Create central incident commander, communications, and scribe rotations staffed from a trained incident response guild across all teams; use follow-the-sun between the two AWS regions and time zones.
- Define escalation layers: service on-call to team lead or manager to service owner to executive.
- Define handoff times, shadow shifts, and load balancing; target at most one week of on-call per engineer per month.
- Bridge the current 12-team paid on-call to 28-team paid coverage; no team remains uncovered.
5. Define on-call rotations, compensation, and alert quality rules (depends on: 4)
Define sustainable rotations, pay, and rules that eliminate noisy pages.
- Rotations: weekly or biweekly, at least one primary and one secondary, with 12-hour shifts where possible or 24-hour for low-volume services.
- Compensation: monthly on-call stipend for all on-call engineers, additional incident response bonus for after-hours work, and time off in lieu; align with market rates.
- Alert quality rules: every page must be actionable, have a runbook link, specify a service owner, include severity, and be based on SLO burn or known failure signals; no dashboard-only alerts.
- Noise budget: reject or downgrade non-actionable alerts; all pages must go to on-call only after suppression and deduplication.
- Weekly alert review removes the top noisy alerts.
6. Design detection, escalation, and alert routing (depends on: 2, 4, 5)
Define how incidents are detected, routed, and escalated so nothing waits on a human to notice.
- Consolidate the six alert tools into one alerting and paging platform with routing by service, severity, and tags.
- Detection sources: infrastructure metrics, application synthetic transactions, log-based anomalies, business transaction SLI monitoring, and customer-reported issues through support or account managers.
- Routing: alert is paged to service on-call within 30 seconds; primary must acknowledge within 5 minutes; if no ack, page secondary then on-call manager.
- Escalation timeouts: unresolved SEV1 escalates to service owner at 15 minutes and to leadership at 30 minutes; any engineer can escalate to the incident commander.
- Define customer-reported incident intake and classification in the same tool.
7. Define internal and external communication protocols (depends on: 2, 3)
Define communication channels, templates, and timing for internal, customer, and regulator audiences.
- Internal: dedicated incident Slack channel, internal status page mirror, and war room bridge for SEV1; incident commander and communications lead own these channels.
- Status page: SEV1 post within 5 minutes, updates every 30 minutes or on material change, resolution within 60 minutes of mitigation; SEV2 post within 15 minutes, updates hourly; SEV3 optional.
- Account managers: SEV1 and SEV2 notify account managers within 15 minutes with an approved customer-facing description and expected impact.
- Regulators: legal or compliance determines notification for data breaches, security incidents, funds availability issues, or regulatory reportable events; criteria and timing follow legal and regulatory requirements; communications lead coordinates.
- Use pre-approved message templates and an approval chain; no ad-hoc wording.
8. Define postmortem policy and action tracking (depends on: 3)
Define mandatory blameless postmortems and action tracking.
- Mandatory for all SEV1 and SEV2 incidents, and any SEV3 that breaches SLA or is customer-detected.
- Format: impact, timeline, root causes, contributing factors, detection and response gaps, what worked well, and action items.
- Blameless: focus on system and process causes, not individual blame; use trained facilitators.
- Ownership: each action has an owner, due date, and tracking ID in a single backlog.
- Review postmortems at the weekly incident review; track action closure; expect 100% completion.
- Complete postmortems within 5 business days for SEV1 and SEV2 incidents.
9. Define metrics, dashboards, and review cadence (depends on: 2, 3, 8)
Define metrics and review cadence to measure process health.
- Metrics: MTTD, MTTM, customer detected percentage, alert noise percentage, on-call response time, on-call load, SLA credits paid, and postmortem action completion.
- Dashboards: real-time operational dashboard for on-call engineers and management.
- Weekly incident review: review all SEV1 and SEV2 incidents, action items, and noisy alerts.
- Monthly trends with leadership; quarterly review against SLOs and audit controls.
- Success thresholds: MTTD under 5 minutes, MTTM under 60 minutes for SEV1, customer detected under 15%, and alert noise under 10%.
10. Configure incident tooling and integrations (depends on: 5, 6, 7, 8, 9)
Implement and integrate the tools that automate the defined process.
- Aggregate alerts from the existing six tools into PagerDuty, Opsgenie, or a similar platform.
- Configure on-call schedules, escalation policies, and paging targeted at service owners.
- Integrate status page API for automated or one-click updates.
- Add Slack commands to declare incidents, start war rooms, assign roles, and post status updates.
- Integrate runbook and service catalog access; create postmortem templates in Jira or Notion with action item tracking.
- Ensure audit trails and role assignments are logged for SOC 2.
11. Pilot with 2-3 volunteer teams and iterate (depends on: 10)
Run a controlled pilot before full rollout to validate and refine the process.
- Select 2-3 volunteer teams with representative services and on-call patterns.
- Run the new severity, roles, on-call, alerting, and communication process for 2 weeks.
- Track metrics and gather feedback from on-call engineers, incident commanders, and communications leads.
- Iterate severity thresholds, alert rules, templates, and runbooks based on findings.
- Exit criteria: no SEV1 without a declared incident commander, alert noise below target, and positive on-call survey results.
12. Train and certify all 28 teams (depends on: 11)
Prepare all 28 teams so the process is not a paper exercise.
- Create role-specific training: all engineers, incident commanders, communications leads, scribes, subject-matter responders, and managers.
- Deliver interactive workshops and online modules.
- Run at least two game-day drills per quarter simulating SEV1 and SEV2 incidents, testing escalation, communication, and postmortem.
- Certify incident commanders and communications leads; only certified staff can hold those roles in rotations.
- Include scenario-specific training for account managers and legal or regulatory handling.
13. Full rollout and enforce process (depends on: 12)
Move every team onto the new process and retire old habits.
- Migrate all 28 teams into consolidated on-call schedules and alert routing.
- Retire old alerting tools or methods that bypass the process.
- Activate status page templates and postmortem action tracking globally.
- Monitor the first 30 days with a daily operational review and quick fixes.
- Communicate rollout status to leadership and all engineers; provide office hours and support.
14. SOC 2 readiness and continuous improvement (depends on: 13)
Map the process to SOC 2 and keep improving after go-live.
- Map each process component, including severity, roles, detection, communications, postmortems, and metrics, to SOC 2 control requirements.
- Produce runbooks, evidence of training, incident logs, postmortems, and action item tracking.
- Conduct an internal readiness assessment at month 6 and remediate gaps before the audit at month 8.
- Maintain weekly and monthly metrics reviews to drive continuous improvement.
- Refresh on-call compensation and alert quality rules based on operational data.
Your answer has these parts:
- "round_summary": one or two sentences on the round as a whole.
- "shared": a short list (four items at most) of what most proposals have in common: approaches, steps, priorities.
- "differences": a short list (four items at most) of what separates them, naming the proposals and the steps concerned.
- "proposals": one entry per proposal, each with "proposal" (its number), "summary" (a very concise summary of what the agent proposes: three or four sentences at most) and "approach" (the angle it takes, in a few words).
[ROUND 1]
[SYSTEM]
You are an expert reviewer of multi-agent planning processes.
Several LLM agents drafted plans for a task, refined them over a number of rounds while seeing each other's proposals, and finally voted for the best one.
Be exhaustive but precise: name concrete steps, ideas and metrics, never generalities. Judge plans by their fitness for the task as stated, their realism, their completeness, the soundness of their order and dependencies, how measurable their success is and how they handle things going wrong.
You are an impartial evaluator, not a chronicler: assess the proposals and the process on their merits, never rationalise what happened or assume that the outcome was right.
After your analysis, answer in the requested structure.
Every text field you write will be read by a busy person who skims. Make it easy to skim: short sentences and short paragraphs; when you name several things, prefer a list to a paragraph, with sub-items when an item has parts, but keep a single fact as a sentence; lead with the point and then the evidence; name proposals and steps by number (P2, step 4); no preamble, no repetition of the question, no closing summary; bold at most one key phrase per item or paragraph. Text fields accept Markdown: a blank line between paragraphs, "- " for lists, **bold**.
[HUMAN]
Task given to the agents: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
This is round 1, a refinement round: every agent received ALL the proposals of round 0 and wrote a new plan, improving on them or taking a different approach. By convention, the previous version of proposal N is proposal N of round 0, written by the same model.
PROPOSALS OF ROUND 0 (the previous versions):
--- PROPOSAL 1 (agent opus5_initial_1, anthropic/claude-opus-5) ---
Estimated complexity: high
Success metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 9 months.
- Customer-first detection falls from 40% of incidents to under 10% within 6 months and under 5% within 12.
- Median time to mitigate falls from 3 h 10 min to under 60 minutes within 12 months.
- An Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes.
- Status page updated within 15 minutes of SEV1 declaration and 30 minutes of SEV2 in 95% of cases.
- Monthly alert volume falls from 3,400 to under 500 pages, with actionability above 75%; out-of-hours pages under 2 per person per week.
- All six legacy alerting tools consolidated into one paging platform, legacy paging paths disabled, by week 16.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by their due date.
- SLA credits fall from $1.3M to under $400k in the first 12 months.
- Customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar.
- All 28 teams onboarded by week 24, with 24x7 rotations of 6+ certified responders for every Tier 0/1 team.
- 30+ certified Incident Commanders and 20+ certified Communications Leads, giving 24x7 primary and secondary command cover.
- Paid on-call policy approved by HR, Legal and Finance and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among engineers on rotation.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked action items.
Steps (29):
1. Program charter, executive mandate and funding
Convert the CEO's frustration into a named program with one accountable owner, a budget and a deadline that is earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, the 28 teams will opt out.
- Appoint a single Incident Management Program Lead (Head of Reliability/SRE) with direct exec sponsorship from CTO and CEO.
- Form a steering group: CTO, VP Eng, Head of Support/CS, CISO/Compliance, Legal, Finance (SLA credits), HR (on-call pay).
- Set the non-negotiables: one severity scale, one paging tool, one postmortem format, mandatory action tracking, paid on-call.
- Fix the timeline: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready by week 28 (four weeks of buffer before the audit).
- Approve budget lines: tooling (~$150–250k/yr), on-call compensation (~$600k–1M/yr), 2–3 dedicated program FTEs. Anchor it against $1.3M of credits plus incident cost.
2. Forensic baseline of the 31 incidents and the alert estate (depends on: 1)
Before designing anything, rebuild the facts. Re-open all 31 incidents and profile the 3,400 monthly alerts so every later design decision is evidence-based.
This also creates the "before" picture the exec and the auditor will compare against.
- Re-code each incident: trigger, service, detection source (customer vs monitor), timestamps for detect/acknowledge/declare/mitigate/resolve, who led, credits paid, root cause family.
- Quantify the 40% customer-first detections: which signal was missing in each case.
- Classify the two "nobody in charge" incidents minute by minute; use them as the burning-platform story.
- Audit the six alerting tools: volume per tool, per team, per alert rule; identify the top 50 rules that produce most of the 85% noise; find rules with no owner and no runbook.
- Baseline the numbers formally: MTTD 22 min, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed. Freeze them as the reference line.
3. Stakeholder listening tour and resistance map (depends on: 1)
Engineer pushback against "carrying a pager for other teams' code" is the main delivery risk. Treat it as a design input, not an attitude problem.
Run structured interviews across all 28 teams, plus Support, CS and Sales, in two weeks.
- Test the real objection: is it unpaid work, night sleep, unfamiliar code, poor runbooks, or fear of blame? Each has a different fix.
- Collect current informal practices — the 12 teams already on-call are the pilot candidates and the source of veterans.
- Document the promise that answers the objection: **you are only paged for services your team owns**, plus a trained commander who runs the incident and pulls in others.
- Map influencers and blockers by team; recruit 10–15 credible engineers as a design working group so the process is co-authored, not imposed.
- Survey baseline sentiment (trust in alerts, willingness to be on-call, burnout) to re-measure at 6 and 12 months.
4. Service ownership catalog and criticality tiering (depends on: 2)
You cannot page the right person across 180 services until each service has a named owning team. This is the foundation of both on-call fairness and severity mapping.
Build a machine-readable catalog (Backstage or equivalent) that is the single source of truth for routing.
- One owning team per service, a named engineering manager, a Slack channel, a paging escalation policy, a dependency list.
- Tier services by business impact: Tier 0 (money movement, ledger, auth, shared PostgreSQL cluster), Tier 1 (customer-facing but degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map each Tier 0/1 service to the customer-visible capability it supports (payment initiation, settlement, reporting, onboarding).
- Flag orphan services and cross-team shared components; force an ownership decision for each within 30 days, or schedule decommissioning.
- Publish coverage gaps to the steering group: any Tier 0 service without an owner is an executive escalation.
5. Severity scale and declaration criteria (depends on: 2, 4)
Define a five-level scale with objective, payments-specific triggers so declaration is a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade.
Each level triggers a fixed bundle of response, comms and postmortem obligations.
- **SEV1**: money movement stopped or incorrect, ledger integrity in doubt, data breach, full region loss, >10% of customers impacted. Triggers: immediate 24x7 page of IC + comms + exec, bridge within 5 min, status page within 15 min, mandatory postmortem, regulator assessment.
- **SEV2**: severe degradation, settlement at risk of missing a window, single large/strategic customer fully down, SLA breach likely. Triggers: IC paged, status page within 30 min, mandatory postmortem.
- **SEV3**: partial or workaround-available degradation, no credit exposure. Team-led, business-hours comms, postmortem optional but encouraged.
- **SEV4/5**: minor or internal-only; ticket-tracked, no paging.
- Add auto-escalation rules: any SEV3 open >2 h, or any incident touching the shared ledger cluster, becomes SEV2 automatically. Include a severity decision tree and 12 worked examples drawn from the 31 real incidents.
6. Incident roles, decision authority and handover rules (depends on: 5)
Solve the "nobody in charge for an hour" failure by making command explicit, transferable and logged.
Define five roles with written responsibilities, entry criteria and explicit authority.
- **Incident Commander**: owns the incident, not the fix. Authority to declare severity, pull any engineer, approve customer-impacting mitigations, invoke failover and authorise spend. The IC never types in the terminal.
- **Communications Lead**: owns status page, internal updates, account-manager briefings and the exec summary. Single voice to customers.
- **Scribe**: maintains the timeline, decisions and open questions; feeds the postmortem and the audit evidence trail.
- **Subject-Matter Responders**: engineers from owning teams; they investigate and remediate, and report to the IC.
- **Executive Liaison** (SEV1 only): shields the IC from exec questions and owns regulator/board escalation.
- Rules: the IC role is assumed within 5 minutes of declaration, stated explicitly in the channel ("I am IC"), and any handover is announced and logged. Roles may be combined below SEV2; never at SEV1.
7. 24x7 incident command coverage model (depends on: 6, 4)
Command is staffed by a small, trained, cross-team pool — not by 28 teams individually. This is what makes 24x7 realistic in one New York time zone.
Design a central rotation that scales with a growing certified pool.
- Create a **Duty Incident Commander** rotation of 25–35 certified volunteers (target ~1 per team, plus managers and senior engineers), giving each person roughly one week per 6–8 months.
- Pair with a Duty Comms Lead rotation (Support/CS leads plus engineering managers, ~15–20 people) and a Scribe pool (rotating, lowest barrier, used as the training entry point).
- Coverage: primary + secondary IC at all times; hard 5-minute acknowledgement SLA with automatic failover to secondary, then to the on-call engineering director.
- Night coverage options to evaluate in writing: US-only rotation with paid night stipend now; a Lisbon/Dublin or APAC follow-the-sun cell as a 12-month option; a 24x7 NOC-style triage desk for first-line detection.
- Eligibility: certification required (S21); commanders are volunteers with manager approval and can step out with 30 days' notice.
8. Team on-call structure, rotations and routing rules (depends on: 4, 6)
Rebuild team on-call around the principle that answers the pushback: **you are only paged for code your team owns**.
Apply a tiered obligation so 28 teams are not treated identically.
- Tier 0/1 owning teams (expected 14–18 teams): 24x7 primary + secondary, minimum 6 people per rotation, one-week shifts, handover Wednesday mornings.
- Tier 2/3 teams: business-hours on-call with a best-effort out-of-hours escalation path, no night paging.
- Platform/Infrastructure and Database teams: 24x7, since they own the shared PostgreSQL ledger cluster and the Kubernetes/regional layer.
- Rotations under 6 people are merged across teams or backfilled by hiring; no rotation of fewer than 4 is approved.
- Routing: every page resolves through the service catalog to the owning team's escalation policy; cross-team pages are made by the IC, never by an alert.
- Guardrails: maximum one week in four, no on-call in the week after a SEV1 you led, protected recovery time after any night page, and a per-person page budget (see S10).
9. On-call compensation, labour compliance and fairness policy (depends on: 8, 3)
Unpaid on-call is both a retention risk and a legal exposure in New York. Paying for it is the fastest way to convert resistance into participation.
Design the scheme with HR, Legal, Finance and Payroll, and publish it before asking anyone to sign up.
- Base stipend per week on rotation, differentiated by tier: e.g. $800–1,200 for 24x7 Tier 0/1, $300–500 for business-hours rotations, with premiums for holidays and weekends.
- Per-incident payment for out-of-hours activation (e.g. $150 per night page plus hourly beyond one hour) and guaranteed time-off-in-lieu after night work.
- Separate Duty IC stipend, since command is a distinct and heavier burden.
- Verify FLSA exempt/non-exempt treatment, NY State wage rules and overtime exposure for non-exempt staff; document the legal review.
- Budget and model the annual cost; get board/CFO approval as a line item, benchmarked against $1.3M of credits.
- Add non-cash elements: on-call time counted as delivery load (teams reduce sprint commitment by ~15%), incident leadership recognised in promotion criteria, and a public quarterly report of on-call load per team.
10. Alert quality standard and page budget (depends on: 2, 4)
3,400 alerts a month at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a contractual condition of paging someone.
Publish a standard, then enforce it mechanically.
- Every paging alert must have: a named owning team, a documented customer impact, a runbook link, a tested threshold, and a severity mapping. Alerts failing this are demoted to ticket or deleted.
- Page only on symptoms that affect customers (SLO burn rate, error budget, queue depth against settlement deadlines); cause-based CPU/memory alerts become dashboards or tickets.
- Set a **page budget**: maximum 2 out-of-hours pages per person per week. Breach triggers a mandatory alert-tuning sprint for the owning team and blocks new alert creation.
- Auto-quarantine: any alert that fires more than 5 times a month without action, or has >70% no-action acknowledgements, is silenced automatically and returned to its owner.
- Monthly alert review per team: kill, tune, or keep, with the numbers on screen.
- Target: 3,400 → under 500 pages a month, with actionability above 75% within six months.
11. Detection uplift: SLOs, synthetic journeys and ledger assurance (depends on: 10, 4)
The goal is to stop customers telling you first. Detection must be driven by customer-visible outcomes, not host metrics.
Instrument the money path end to end and alert on it.
- Define SLOs for each Tier 0/1 customer capability: payment initiation success rate, authorisation latency, settlement file timeliness, API availability and reporting freshness. Tie them to the 99.95% contractual SLA with a stricter internal target.
- Deploy synthetic transactions from outside the platform, in both regions, every 60 seconds, covering the full payment lifecycle including a small real-value canary flow where feasible.
- Add ledger assurance checks: continuous double-entry balance reconciliation, replication lag and failover-readiness alarms on the shared PostgreSQL cluster, and settlement-window countdown alerts.
- Build per-customer anomaly detection for the top 100 accounts (volume drop, error spike) so a single-tenant outage is detected before the account manager calls.
- Create an "inbound signal" bridge: any support ticket or account-manager report matching impact keywords auto-creates a triage incident within 5 minutes.
- Track every incident's detection source; make "customer detected first" a reviewed defect with its own follow-up action.
12. Tool consolidation and incident platform implementation (depends on: 5, 7, 8, 10)
Collapse six alerting tools into one paging and incident platform so there is a single queue, a single timeline and a single audit record.
Run a short, time-boxed selection and migrate within the pilot window.
- Select an integrated stack: paging/on-call scheduling plus an incident management layer (e.g. PagerDuty + incident.io/FireHydrant, or a single vendor) and a hosted status page.
- Implement one-command declaration in Slack (`/incident declare`) that creates the channel and bridge, pages the Duty IC, sets severity, opens the timeline and starts the clock.
- Migrate all monitoring sources to route into the one platform; decommission direct paging from the legacy six and block new integrations that bypass it.
- Automate the evidence trail: timestamps, role assignments, severity changes, comms sent, and postmortem linkage exported for SOC 2.
- Integrate with the service catalog for routing, Jira for actions, Salesforce/CS tooling for affected-customer lists, and Zoom/Slack Huddle for the bridge.
- Hard requirement: the platform must work when AWS in one region is down — verify out-of-band paging (SMS/phone) and a printed/offline fallback runbook.
13. Detection-to-escalation path and the five-minute command rule (depends on: 6, 7, 12)
Write the single path from "something looks wrong" to "someone is in charge", and make it impossible to skip.
The design target is detection to commander in under five minutes, any hour.
- Entry points: automated alert, engineer observation, support ticket, account manager, partner bank, customer-facing SEV hotline. All converge on the same declaration command.
- Anyone in the company may declare up to SEV2; nobody is punished for over-declaring. Publish that rule in writing and repeat it.
- Auto-page ladder: Duty IC (5 min) → secondary IC (5 min) → on-call Director (10 min) → CTO. Same ladder for the owning team's responder.
- Cross-team pull: the IC can page any team's on-call directly, with a 10-minute acknowledgement obligation. This is the reciprocal commitment that makes single-team ownership viable.
- Explicit takeover protocol: if no one claims IC within 5 minutes, the platform assigns it and announces it; the assignee cannot decline, only hand over.
- Define standing severity triggers for immediate regional failover, ledger read-only mode and partner-bank notification, with pre-authorised decision rights so the IC does not wait for an executive.
14. Internal communications protocol (depends on: 6, 12)
Standardise the internal channel so responders, executives and support see the same picture without interrupting the IC.
Separate the working channel from the audience channel.
- One incident channel per incident (auto-created), one bridge, and a read-only broadcast channel for executives, Support and Sales.
- Update cadence by severity: SEV1 every 30 minutes even if nothing has changed; SEV2 every 60 minutes; SEV3 at state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, ETA or next update time, current IC and Comms Lead.
- Exec briefing rule: executives ask questions only to the Executive Liaison; the IC is not interrupted. Publish this as a behavioural expectation signed by the exec team.
- Support/CS enablement: a live affected-customer list and a holding statement within 15 minutes of SEV1/SEV2 so the front line is never guessing.
- Handover protocol for incidents beyond 4 hours: formal IC handover checklist, fatigue rule, and staffing of a second shift.
15. Customer communications and status page policy (depends on: 5, 14)
Customers currently learn of outages from their own monitoring and hear from whoever happens to be around. Replace that with a timed, owned, pre-approved process.
The Comms Lead is the single author; templates remove the need to write under pressure.
- Timing commitments: status page posted within 15 minutes of SEV1 declaration and 30 minutes for SEV2; updates every 30/60 minutes; resolution notice within 30 minutes of mitigation; customer-facing summary within 5 business days for SEV1.
- Pre-approve 12–15 templates with Legal and Comms (degradation, delay in settlement, API errors, security event, third-party failure) so nothing needs legal review mid-incident.
- Subscription-based status page with per-component granularity mapped to the customer capabilities from S11, plus an email/webhook/RSS feed.
- Tiered outreach: top 100 accounts get a direct named call or email from their account manager within 30 minutes of SEV1, with a briefing pack from the Comms Lead; long tail gets the status page and a proactive email.
- Rules of language: state impact and next update time, never speculate on cause, never assign blame to a vendor before facts are confirmed.
- Run a quarterly customer-perception check with the top accounts on whether comms were timely and useful.
16. Regulatory, partner and legal notification playbook (depends on: 5, 15)
In payments, some incidents are reportable and the clock starts at detection. Build the assessment into the process so it is never an afterthought.
Work with Legal, Compliance and the CISO to produce a decision tree and contact matrix.
- Map obligations: NYDFS Part 500 (72-hour cybersecurity event notification), state breach laws, GLBA/FTC Safeguards, PCI DSS if card data is in scope, sponsor-bank and card-network contractual notice windows, and any FinCEN/OFAC implications.
- Add a mandatory regulatory-assessment checkpoint to every SEV1 and every security-related SEV2, owned by the Executive Liaison, completed within 2 hours of declaration and recorded even when the answer is "not reportable".
- Build the contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel, with 24x7 numbers and named backups.
- Pre-draft notification letters and hold them under legal privilege review.
- Check customer contracts for bespoke notification SLAs (often 1–4 hours for enterprise accounts) and encode them in the customer tiering.
- Test the playbook once per quarter as part of the simulation programme.
17. SLA credit and financial impact workflow (depends on: 5, 15)
Link incidents to money so severity, credits and prioritisation stay consistent — and so Finance stops being surprised.
Make credit calculation an automated output of the incident record, not a negotiation.
- Define the availability measurement method per contract, per component, and agree it with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and the per-capability telemetry; generate a proposed credit schedule within 5 business days of resolution.
- Decide the posture: proactive credits for the top tier (reputation upside) versus claims-based for the rest; document the approval chain.
- Track credits per incident and per root-cause family; feed a quarterly report showing which reliability investments would have prevented which credits.
- Set a target: reduce credits from $1.3M to under $400k in year one, and use that delta as the ongoing business case.
18. Postmortem standard and blameless review forum (depends on: 5, 6)
Replace "some incidents, various formats" with a mandatory, single-format, blameless process with fixed deadlines.
The discipline is in the deadlines and the forum, not in the template.
- Mandatory for: every SEV1 and SEV2, every incident where a customer detected it first, every incident over 2 hours, every repeat of a known cause, and every near-miss involving the ledger. Optional but templated for SEV3.
- Fixed timeline: draft within 3 business days, peer review within 5, published company-wide within 10. The IC owns delivery; the owning team's manager is accountable.
- One template: timeline, customer and financial impact, detection analysis (why not sooner), response analysis (why mitigation took as long as it did), contributing factors, what went well, action items with owner and due date.
- Blameless rules in writing: describe systems and decisions in the context available at the time; no individual named as a cause; HR and management commit that postmortems are never used in performance reviews.
- Weekly 60-minute Incident Review Board: reviews all postmortems from the prior week, challenges quality, ratifies severity, and approves or rejects action items. Attendance by engineering directors is mandatory.
- Publish a searchable postmortem library and a quarterly "top five recurring causes" analysis.
19. Action item ownership and tracking system (depends on: 18, 12)
11 of 64 actions closed is the single clearest symptom of a process nobody enforces. Give actions the same status as customer commitments.
Track them where engineering work already lives, with visible escalation.
- Every action gets: a named individual owner (not a team), a priority class, a due date and a Jira ticket auto-created from the postmortem.
- Priority classes with hard SLAs: P0 prevents recurrence of a SEV1, due in 30 days; P1 in 60 days; P2 in 90 days. P0s are committed into the next sprint before any roadmap work.
- Capacity rule: teams reserve a standing 15–20% of sprint capacity for reliability and incident actions. Without reserved capacity, the actions will not land.
- Escalation ladder for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Monthly reporting of closure rate by team in the engineering leadership review; include it in manager performance objectives.
- Target: 90% of P0/P1 actions closed on time within two quarters.
20. Runbooks, major-incident playbooks and the on-call readiness bar (depends on: 4, 8)
Nobody can respond well to unfamiliar systems at 3 a.m. without runbooks — and poor runbooks are a real part of the pager resistance.
Define a minimum readiness bar that a service must meet before it is allowed to page anyone.
- Readiness checklist per Tier 0/1 service: current architecture diagram, dependency map, dashboard link, alert-to-runbook mapping, rollback procedure, feature-flag kill switches, escalation contacts, and a data-loss/latency impact statement.
- Write major-incident playbooks for the top failure modes derived from S2: shared PostgreSQL ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank/third-party outage, settlement-window breach, and suspected security compromise.
- Prioritise the shared ledger: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and a clearly stated data-loss tolerance (RPO/RTO) signed off by the exec.
- Runbooks must be tested at least twice a year in a drill; untested runbooks are marked stale in the catalog.
- Enforcement: a service without readiness sign-off cannot create paging alerts, and the gap is reported to its director.
21. Training, certification and the commander academy (depends on: 6, 13, 14, 18)
Command is a skill, not a title. Build a certification path so 24x7 coverage is staffed by people who have practised.
Use a tiered curriculum with real assessment.
- **Scribe** (2 hours): timeline discipline and tooling. The entry point for everyone.
- **Responder** (half a day): severity scale, declaration, escalation, runbook use, comms hygiene. Mandatory for every engineer joining an on-call rotation.
- **Incident Commander** (two days plus shadowing): command presence, delegation, decision-making under uncertainty, severity calls, handover, exec management. Requires two shadowed incidents and one simulated SEV1 before certification.
- **Communications Lead** (one day): status-page writing, customer tiering, legal boundaries, regulator triggers.
- Certification is valid 12 months and renewed via a simulation; the register of certified people is an audit artefact.
- Add on-call onboarding per team: a new joiner shadows two shifts before holding primary, and never holds primary in their first 90 days.
22. Simulation programme: game days, drills and wheel of misfortune (depends on: 21, 12, 20)
The process must be rehearsed before it meets a real SEV1. Simulations also build the commander pool and expose runbook gaps cheaply.
Run a standing calendar rather than one-off exercises.
- Monthly 60-minute tabletop ("wheel of misfortune") per engineering group, using a real past incident from the 31.
- Quarterly full-scale game day in production or a production-like environment: regional failover, ledger replica promotion, dependency failure, with the whole role structure activated and timed.
- Twice-yearly unannounced paging drill to measure real acknowledgement times at night.
- One security-incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Every exercise produces a lightweight postmortem and action items in the same system as real incidents.
- Measure and publish drill metrics: time to IC, time to first status update, time to correct mitigation decision.
23. Pilot with wave 0 teams (depends on: 22, 9, 11)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Run a six-week pilot with tight measurement and a public verdict.
- Select 5–6 teams: core payments, ledger/database, platform/Kubernetes, API gateway, plus two of the 12 teams already on-call.
- Activate the full stack for them: new severity scale, Duty IC rotation, single paging tool, alert budget, status-page policy, mandatory postmortems, paid on-call.
- Hold a weekly pilot retro; expect and document 20–40 process defects, and fix them in the standard before rollout.
- Validate the hard questions: does the 5-minute IC rule hold at 3 a.m.? Do cross-team pulls get answered? Is the severity tree unambiguous?
- Exit criteria: MTTD under 10 minutes for pilot services, IC assigned within 5 minutes in 95% of incidents, page volume down 50%, all postmortems on time, positive on-call sentiment.
- Publish a one-page pilot result to the whole company — this is the main adoption argument for the remaining teams.
24. Metrics, dashboards and the review cadence (depends on: 12, 18, 23)
Instrument the process itself, so improvement is visible and the audit has evidence of monitoring and review.
Define a small set of metrics with owners and a fixed meeting rhythm.
- Response metrics: MTTD, time to declare, time to IC assigned, MTTA, MTTM, MTTR, incidents per month by severity, % detected by customers first.
- Quality metrics: page volume per person per week, alert actionability rate, page budget breaches, postmortem on-time rate, action closure rate and ageing.
- Business metrics: SLA credits paid, availability against 99.95% per capability, error-budget consumption, repeat-incident rate.
- People metrics: on-call load distribution across teams, out-of-hours pages per person, on-call sentiment and attrition among on-call staff.
- Cadence: weekly Incident Review Board (postmortems and actions), monthly Reliability Review (metrics per team, alert hygiene, on-call load), quarterly Executive/Board review (credits, trends, investment asks), annual policy review.
- Every metric gets a target and a named owner; dashboards are self-serve and public inside the company.
25. Wave rollout across all 28 teams with readiness gates (depends on: 23, 24)
Roll out in four waves of six to eight teams, every three weeks, ordered by criticality. Each wave passes an explicit gate rather than a deadline.
Gates keep quality high and make the standard credible.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with 6+ certified responders, one IC candidate nominated, one drill passed.
- Assign each wave a named coach from the program team for three weeks of hands-on support.
- Gate criteria are checked and signed by the director; teams that fail are re-scheduled, not waived.
- Freeze legacy tooling per wave: after onboarding, the old alerting paths are disabled, not left as a fallback.
- Publish a live adoption scoreboard by team so progress is social, not administrative.
26. SOC 2 control mapping, evidence automation and internal dry-run audit (depends on: 18, 19, 25)
Design the process so audit evidence is a by-product of doing the work, then test it before the auditors do.
Engage the auditor early to confirm the interpretation of controls.
- Map the process to the Trust Services Criteria: CC7.3 and CC7.4 (incident identification, response, recovery), CC7.2 (monitoring), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), plus availability criteria A1.2.
- Produce and approve formal policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard — versioned, signed, annually reviewed.
- Automate evidence: incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, postmortem library, action-item closure reports, training and certification register, drill records.
- Confirm the observation window with the auditor and ensure the process is operating for a minimum of three months before fieldwork.
- Run an internal dry-run audit at month six: sample 15 incidents and walk the full evidence chain; fix gaps with 8 weeks to spare.
- Keep a remediation log for any incident where the process was not followed, with the corrective action — auditors respond better to documented exceptions than to a claim of perfection.
27. Change management, incentives and communications campaign (depends on: 3, 9, 23)
Run this in parallel from day one. The process will be judged by engineers on fairness, and by executives on visible results.
Communicate the deal explicitly and repeatedly.
- The deal in one sentence: **you are paid for on-call, you are paged only for what you own, a trained commander runs the incident, and your postmortem actions get real sprint capacity**.
- Launch communications: CTO all-hands, per-team roadshows, a one-page process card for laptops, an internal wiki hub and a Slack support channel with a 4-hour answer SLA.
- Recognition: incident-response contribution in promotion criteria and performance frameworks, quarterly awards for best postmortem and biggest alert-noise reduction, public thanks after every SEV1.
- Manager accountability: adoption, alert hygiene, action closure and on-call load in each engineering manager's quarterly objectives.
- Handle the exceptions: a written path for engineers who cannot do nights (caring responsibilities, health), covered by stipended volunteers elsewhere.
- Track sentiment quarterly and publish the results, including bad news, to keep credibility.
28. Program risk register and contingency planning (depends on: 1)
Name the ways this program fails and pre-commit the response. Review it monthly in the steering group.
The main risks are predictable.
- **Volunteer shortfall for the IC pool**: contingency is to make command a rostered duty for engineering managers and senior engineers until the pool reaches 25.
- **Compensation not approved in time**: fall back to time-off-in-lieu plus a phased stipend, but do not launch mandatory night on-call without some compensation.
- **Tool migration slipping**: keep the single-queue requirement and cut scope on the incident-management layer, not on paging consolidation.
- **Alert pruning causing a missed incident**: prune from paging to ticket first, observe for 30 days, then delete; keep a recovery path.
- **Burnout or attrition among the 12 experienced on-call teams**: monitor load weekly and cap individual page counts.
- **A major SEV1 mid-rollout**: pre-agree that the program lead becomes a full-time responder and the wave schedule slips by one wave, with the steering group informed the same day.
29. Continuous improvement, maturity roadmap and post-audit sustainability (depends on: 24, 25, 26)
Protect against the classic failure: the process decays once the audit is passed. Build the second-year plan before the first year ends.
Set a maturity model and a forward roadmap with owners.
- Quarterly process retrospective with the IC pool: what in the process itself slowed us down, what needs simplifying, is the severity scale calibrated?
- Re-baseline targets every six months; a process that hits all targets is under-ambitious.
- Year-two roadmap candidates: follow-the-sun coverage cell, automated mitigation and self-healing for the top three recurring causes, error-budget policy that gates releases, per-customer real-time impact reporting, and blast-radius reduction for the shared ledger cluster (the largest single structural risk).
- Move from lagging metrics (MTTR) to leading ones (error-budget burn, near-miss rate, drill performance).
- Make the annual policy review, certification renewal and drill calendar permanent calendar items owned by the Head of Reliability, independent of the audit cycle.
- Report to the board quarterly on availability, credits and incident trends so the process keeps executive attention after SOC 2 is signed.
--- PROPOSAL 2 (agent gpt5.6-sol_initial_2, openai/gpt-5.6-sol) ---
Estimated complexity: high
Success metrics: - Within 7 days, every suspected SEV0–SEV2 has one incident record, one channel, and a named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named incident commander for more than 10 minutes.
- By day 30, 100% of Tier 0 and Tier 1 services have a named owner, primary escalation, secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and subject-matter coverage.
- By day 120, 100% of production services have sustainable ownership and tested escalation paths.
- At least 95% of SEV0 and SEV1 pages are acknowledged within 5 minutes by month 3.
- At least 95% of SEV2 pages are acknowledged within 10 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to 10 minutes by day 90 and 5 minutes by month 6.
- The proportion of incidents first detected by customers falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to below 90 minutes by day 120 and below 60 minutes by month 6.
- At least 95% of qualifying incidents meet their initial customer-communication deadline by month 3.
- At least 95% of published incidents meet their required update cadence by month 3.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without loss of Tier 0 or Tier 1 detection coverage.
- Monthly pages fall from 3,400 to no more than 1,500 by day 90, with alert actionability and missed-detection reviews used as countermeasures against unsafe suppression.
- 100% of new paging alerts satisfy the owner, runbook, dashboard, action, severity, and escalation quality rules by day 60.
- 100% of required SEV0 and SEV1 postmortems are drafted within 3 business days and reviewed within 5 business days by month 2.
- At least 90% of postmortem actions are completed by their approved due dates by month 6.
- All 53 currently open historical actions are triaged within 30 days; all unaccepted high-risk items are completed within 90 days.
- Repeat incidents with the same unaddressed contributing factor decline by at least 50% within 6 months.
- Every primary rotation has at least six trained responders or a documented, time-limited executive exception by day 120.
- No responder is routinely scheduled more frequently than one primary week in six by day 120.
- Two end-to-end cross-company exercises, including regional and ledger scenarios, are completed before the audit, with all critical findings assigned and tracked.
- Monthly availability meets or exceeds the 99.95% contractual target by month 6, with exceptions reviewed at the executive reliability meeting.
- SLA credits decline by at least 50% on an annualized trailing basis by month 8.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (19):
1. Establish ownership, authority, and funding
Launch the program within 48 hours under an executive sponsor. Give one program owner authority to standardize incident management across all 28 teams.
- Name the CTO or equivalent as executive sponsor and a Head of Incident Management or Reliability as directly accountable owner.
- Form a working group with Engineering, SRE or Platform, Product, Support, Customer Success, Communications, Security, Legal, Compliance, Risk, HR, Finance, and Internal Audit.
- Approve authority for an incident commander to stop deployments, roll back releases, disable features, shift traffic, invoke continuity plans, and pause payment processing when integrity is at risk.
- Preserve financial controls. The incident commander may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access requirements.
- Fund paging tools, compensation, training, observability work, exercises, and dedicated reliability capacity.
- Reserve engineering capacity for incident remediation. Start with 10% of capacity and adjust through quarterly risk reviews.
- Record the current baselines: 31 customer-impacting incidents, 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M in credits, 3,400 monthly alerts, 85% noise, and 11 of 64 actions closed.
- Maintain a risk register for staffing gaps, shared-ledger concentration, regional failover, alert coverage, third parties, and audit readiness.
2. Install immediate minimum controls (depends on: 1)
Put an interim process in place during the first seven days. Do not wait for tool consolidation, policy perfection, or the SOC 2 audit.
- Publish a one-page interim severity guide and incident declaration procedure.
- Establish one continuously monitored incident declaration path through chat, telephone, and the paging system.
- Create a standard incident channel, conference bridge, incident document, and event naming convention.
- Staff an interim primary and backup incident commander at all times. Compensate this duty retroactively under the final compensation policy.
- Give trained duty personnel access to the status page, paging system, dashboards, support queue, service catalog, and emergency contacts.
- Require an incident commander to be named within 10 minutes for every suspected major incident.
- Direct Support to escalate credible customer reports immediately rather than waiting for engineering confirmation.
- Triage all 53 open historical postmortem actions. Complete, re-plan, or formally risk-accept the items affecting ledger integrity, payment duplication, regional resilience, security, and detection first.
- Hold a daily 15-minute operational review until permanent controls are working.
3. Create the service and dependency catalog (depends on: 1)
Build a reliable ownership map for all production services and customer journeys. This is the basis for paging, escalation, impact assessment, and audit evidence.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, data stores, queues, external processors, banking partners, and customer-facing endpoints.
- Assign each component a single accountable team, primary responder group, secondary escalation group, engineering manager, and product owner.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 according to financial integrity, customer impact, dependency centrality, and contractual obligations.
- Treat the ledger, payment orchestration, authentication, settlement, reconciliation, and critical shared infrastructure as Tier 0 or Tier 1.
- Map every important customer journey to its service, database, cloud-region, and third-party dependencies.
- Record SLOs, RTOs, RPOs, data classification, dashboards, runbooks, deployment controls, feature flags, and failover methods.
- Assign separate but coordinated responders for the ledger application and the shared PostgreSQL platform.
- Document whether each service is active-active, active-passive, or region-bound. Identify dependencies that make nominal regional redundancy ineffective.
- Make missing ownership or missing runbooks a release-blocking risk for Tier 0 and Tier 1 services.
4. Adopt severity and incident lifecycle standards (depends on: 1)
Approve one impact-based severity model for operational, security, data, and third-party incidents. When evidence is incomplete, start at the higher credible severity and downgrade later.
- SEV0, crisis: Use for actual or credible unauthorized, lost, duplicated, or corrupted movement of money; ledger integrity loss; material security compromise; material data exposure; both-region failure; or an event likely to require crisis or regulatory management. Page all roles immediately, engage executives, Security, Legal, Compliance, and Risk, and consider pausing payment activity.
- SEV1, critical: Use for widespread inability to initiate, process, settle, or reconcile payments; a core journey failing without a viable workaround; material regional impact; fast SLA-budget exhaustion; or an imminent integrity risk. Staff all incident roles, notify the executive duty officer, and publish customer communications.
- SEV2, major: Use for a material customer subset, one or more critical customers, significant degradation with a workaround, partial transaction failure, or a likely contractual impact. Assign an incident commander and subject-matter responders; add communications and scribe roles whenever customers are affected.
- SEV3, minor: Use for localized, low-impact degradation with no financial-integrity, security, regulatory, or material contractual risk. The owning team leads the response and keeps an internal record; external communication is not normally required.
- Base severity on actual or credible impact, not the seniority of the reporter, number of alerts, or presumed complexity of the fix.
- Permit any employee to declare an incident. Only the incident commander may lower severity after recording the evidence and rationale.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of customer impact. Define resolution only after stability, backlog processing, transaction recovery, and required ledger reconciliation are complete.
- Start measurement from the earliest reliable indication of impact, including telemetry, customer reports, and partner notifications.
5. Define roles and sustainable 24x7 staffing (depends on: 3, 4)
Separate command from technical remediation. This allows trained commanders to coordinate any incident without asking engineers to debug code they do not own.
- Incident commander: Owns severity, priorities, role assignment, escalation, decision cadence, mitigation strategy, handoffs, and final closure. One person has command at a time.
- Communications lead: Owns internal notices, status-page updates, account-manager briefs, approved customer language, and coordination with Legal or regulators.
- Scribe: Maintains a timestamped timeline of observations, decisions, commands, owners, and status changes. Automation may assist but does not replace human validation for SEV0 and SEV1.
- Subject-matter responders: Diagnose and mitigate only services or domains for which they have accepted ownership, training, access, and runbooks.
- Executive duty officer: Removes organizational obstacles and approves exceptional business decisions. This role does not take command unless a formal transfer occurs.
- Security, Legal, Compliance, Support, Vendor Management, and Business Continuity join according to predefined triggers.
- Create a company-wide incident-command rotation with at least eight certified primary commanders and eight qualified backups. Use weekly rotations with explicit handoffs.
- Create similarly sustainable communications and scribe pools using Engineering Operations, Support, Customer Operations, and Communications personnel.
- Group service responders into approximately 8–12 coherent product or platform domains rather than creating 28 fragile rotations. Each domain rotation should normally contain at least six trained responders.
- Do not place an engineer into another team's responder pool without training, access, runbooks, shadow shifts, and explicit acceptance by both teams.
- Maintain dedicated database-platform and ledger-application escalation coverage for the shared PostgreSQL environment.
- Require distinct people for commander, communications, and primary technical lead during SEV0 and SEV1 incidents.
- Require a verbal and written handoff when any incident role changes. Record the exact time and the new role owner.
6. Implement compensation and fatigue safeguards (depends on: 5)
End unpaid on-call before expanding coverage. Treat availability, interrupted personal time, and overnight recovery as compensable work.
- Pay a fixed stipend for each primary on-call week and a secondary stipend equal to a defined percentage of the primary amount.
- Pay a higher holiday stipend. Apply overtime and call-out rules to non-exempt employees as required by law.
- Give exempt employees a minimum call-out credit or equivalent paid recovery time for material after-hours work.
- Provide a paid recovery day after prolonged overnight work, a SEV0, or a qualifying SEV1. Managers must arrange daytime coverage rather than expecting normal output.
- Have HR, Finance, and employment counsel publish dollar amounts, tax treatment, eligibility, and payroll procedures within 14 days. Apply the policy consistently across teams and locations.
- Target rotations no more frequent than one week in six. Exceptions require a time-limited staffing plan and executive risk acceptance.
- Avoid consecutive primary and secondary weeks. A person must not be primary for two simultaneous domain rotations.
- Track after-hours pages, sleep interruptions, swaps, missed acknowledgements, and reported burnout by rotation.
- Trigger a staffing or alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
- Permit responders to declare themselves temporarily unfit after overnight work without performance penalty.
7. Consolidate incident and paging tooling (depends on: 4)
Create one operational system of record while migrating safely from the six current alerting tools. Consolidation must reduce ambiguity without creating a monitoring gap.
- Select one enterprise paging and escalation platform and one integrated incident record.
- Initially ingest events from all six tools. Deduplicate, correlate, and route them through the new platform before retiring sources.
- Integrate paging with chat, conference bridges, ticket tracking, the service catalog, observability tools, and the customer status page.
- Automatically capture declaration time, acknowledgements, role assignments, severity changes, messages, decisions, mitigated time, and resolved time.
- Use role-based access, multifactor authentication, break-glass controls, immutable audit logs, and periodic access reviews.
- Provide mobile and telephone fallback paths if chat, identity, or the primary paging tool is unavailable.
- Test paging, escalation, status publication, and conference access every week.
- Retire a legacy alert path only after its signals have named owners, successful end-to-end tests, and at least two weeks of verified operation in the new platform.
8. Improve detection and enforce alert quality (depends on: 3, 7)
Shift detection toward customer journeys, payment outcomes, and ledger integrity. Infrastructure metrics alone will not solve the current customer-first detection problem.
- Instrument payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, and reporting with SLOs and business-level success metrics.
- Run external synthetic transactions and API checks from outside the production boundary and from both AWS regions.
- Monitor transaction failure rates, processing latency, queue age, unprocessed volume, reconciliation breaks, unexpected ledger balances, duplicate identifiers, regional asymmetry, and third-party response quality.
- Correlate application telemetry with Kubernetes, AWS, PostgreSQL, network, deployment, and feature-flag events.
- Route high-priority support cases and credible partner notifications into the same incident declaration path within five minutes.
- Define noise as a page that is duplicate, informational, unactionable, non-production, or requires no timely human action.
- Require every paging alert to name an owner, affected service, urgency, customer or SLO risk, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Send non-urgent conditions to a ticket queue rather than a pager.
- Run new alerts in shadow mode for at least seven days unless an emergency risk exception is approved. Test both firing and recovery behavior.
- Review any alert with less than 50% actionability or more than three firings in seven days within two business days.
- Never silently disable a noisy alert. Verify compensating detection, record the decision, and assign a correction owner first.
- Review alert actionability, false positives, missed detection, and page load with every responder group each month.
9. Codify acknowledgement and escalation paths (depends on: 3, 4, 5, 7, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts when a qualifying signal or customer report enters the system.
- For SEV0 and SEV1, page the owning primary immediately; page the secondary after five unacknowledged minutes; page the domain manager and company incident commander at 10 minutes; and engage the executive duty officer by 15 minutes.
- For SEV2, require primary acknowledgement within 10 minutes and incident-command assignment within 15 minutes. Escalate to the secondary and manager when either target is missed.
- For SEV3, require acknowledgement within 30 minutes when immediate production action is needed. Otherwise create a prioritized work item.
- Automatically page the company incident commander for any credible integrity or security concern, cross-team event, customer-visible Tier 0 failure, regional event, or unresolved ownership question.
- If impact remains unknown after 15 minutes, raise severity rather than waiting for certainty.
- Let the incident commander summon dependency owners, cloud support, database support, payment processors, banking partners, and vendors through maintained escalation contacts.
- Test vendor contacts and premium-support entitlements quarterly.
- Route alerts with no valid owner to the central command rotation, then treat the missing ownership record as a control defect.
- Require human acknowledgement. Delivery to a device or chat channel does not count.
- Record every missed acknowledgement, failed escalation, and manual contact workaround for review.
10. Standardize live incident execution (depends on: 4, 5, 7, 9)
Give responders one concise operating procedure for the first minutes through resolution. Prioritize limiting customer and financial harm before proving a root cause.
- Open a dedicated channel, bridge, incident record, and timeline immediately for SEV0 through SEV2.
- Have the incident commander state severity, known impact, current hypothesis, immediate objective, assigned roles, and next update time.
- Freeze unrelated production changes during SEV0 and SEV1 incidents. Record exceptions approved by the incident commander.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Use pre-approved runbooks for region failover, Kubernetes recovery, PostgreSQL failover, credential rotation, queue recovery, and payment suspension.
- Guard against split-brain, replay, duplication, and out-of-order processing during regional or database recovery.
- Require reconciliation and controlled backlog processing before declaring payment or ledger incidents resolved.
- Keep diagnosis and mitigation workstreams separate when enough responders are available.
- State decisions and owners aloud and in the incident record. Avoid unrecorded direct-message command paths.
- Require stability for a severity-specific observation period before closure. Reopen the incident if impact recurs during that period.
- Conduct an explicit operational handback to the owning team, Support, and Customer Success.
11. Standardize internal, customer, and regulatory communications (depends on: 4, 5, 7, 10)
Communicate known impact early without waiting for a root cause. Use approved facts, acknowledge uncertainty, and give the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish an initial customer status within 15 minutes when disclosure is operationally and legally appropriate, then update every 15 minutes.
- For SEV1, notify internal stakeholders within 15 minutes, publish an initial customer status within 15 minutes, and update at least every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send an initial notice within 30 minutes. Update at least every 60 minutes.
- Do not normally publish SEV3 events. Notify specifically affected customers if contracts or material impact require it.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. State affected capabilities, customer symptoms, workarounds, regions, and next update time.
- Do not speculate about root cause, blame, security scope, recovery time, or data integrity.
- Give account managers a single approved briefing and an affected-customer list. Prohibit contradictory or improvised incident explanations.
- Maintain templates for outages, delays, data-integrity investigation, third-party failure, regional failure, security events, and resolution.
- Issue a resolution notice only after operational recovery and required reconciliation. Provide a customer-facing incident summary within five business days for qualifying events.
- Have Legal and Compliance maintain a jurisdiction, regulator, sponsor-bank, network, cyber-insurer, partner, and contract notification matrix.
- Where applicable, explicitly track the current New York cybersecurity-event notification clock, including the 72-hour requirement, without assuming every incident is reportable.
- Have Legal record the reportability decision, decision time, evidence, approver, deadline, and submission confirmation.
- Allow Security or Legal to limit public detail during an active threat, but require the reason and an alternative stakeholder plan to be recorded.
- Coordinate service-credit calculations and contractual notices with Finance and Customer Success from the same incident record.
12. Make postmortems mandatory and actionable (depends on: 4, 7, 10)
Use postmortems to improve systems and controls, not to assign personal blame. Keep performance or misconduct processes separate from the learning review.
- Require a postmortem for every SEV0 and SEV1.
- Require one for a SEV2 that affected customers, incurred credits, breached an SLO or contract, involved financial or data integrity, repeated a prior failure, exposed a control gap, or lasted more than two hours.
- Permit incident command, Security, Compliance, or the service owner to require a review for a near miss.
- Produce a factual draft within three business days and hold the cross-functional review within five business days.
- Use one template covering executive summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, what worked, what failed, and lessons.
- Include why monitoring did or did not detect the event before customers.
- Avoid a single-root-cause assumption. Examine technical, organizational, process, dependency, testing, and incentive factors.
- Give every action one owner, due date, priority, expected risk reduction, verification method, and linked engineering item.
- Classify actions as containment due within 7 days, corrective work due within 30 days, or strategic work normally due within 90 days.
- Require director approval and documented residual-risk acceptance for overdue high-risk actions.
- Verify effectiveness after implementation. Closing a ticket without evidence does not close the action.
- Publish broadly useful reviews internally. Maintain access-restricted versions for security, privacy, personnel, or legally privileged details.
13. Measure performance and review it routinely (depends on: 7, 8, 11, 12)
Use outcome, process, quality, and human-sustainability measures together. Do not reward teams for suppressing declarations or hiding incidents.
- Measure detection time from first impact to first internal signal, declaration time, acknowledgement time, role-staffing time, mitigation time, resolution time, and recurrence.
- Report both median and 90th percentile. Break results down by severity, service tier, customer journey, region, detection source, and owning domain.
- Track customer-first detection, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, and role conflicts.
- Track availability, error-budget consumption, failed-payment volume, delayed value, reconciliation breaks, impacted customers, contractual breaches, and service credits.
- Track alert volume, actionability, duplicates, after-hours pages, missed pages, pages per responder, and tool-source distribution.
- Track required postmortems completed on time, actions completed by due date, action age, verified effectiveness, and repeat contributing factors.
- Track rotation size, on-call frequency, swaps, recovery days, attrition signals, and quarterly responder sentiment.
- Hold a weekly operational review for recent incidents, overdue actions, alert problems, and upcoming risk.
- Hold a monthly executive reliability review covering trends, investment decisions, accepted risks, and SLA exposure.
- Hold a quarterly resilience and control review with Security, Compliance, Risk, Internal Audit, and Product leadership.
- Use team scorecards to direct investment and assistance, not individual performance penalties.
- Reconcile dashboard data against a monthly sample of incident records and customer cases to detect metric gaming or missing incidents.
14. Train and certify participants (depends on: 4, 5, 9, 10, 11, 12)
Train people before assigning full independent duty. Use paid working time for training, shadowing, exercises, and certification.
- Train all employees to recognize impact, declare an incident, and find the incident channel and status page.
- Train engineers and Support on severity, escalation, customer-report handling, evidence preservation, and financial-integrity precautions.
- Certify incident commanders through instruction, tabletop exercises, shadow incidents, and observed command performance.
- Train communications leads in status writing, contractual communications, regulator escalation, and avoiding unsupported claims.
- Train scribes in timestamping, decision capture, evidence hygiene, and separating fact from hypothesis.
- Require subject-matter responders to demonstrate dashboard, runbook, rollback, failover, and access competence for their assigned domain.
- Add training to new-hire onboarding and repeat role-specific certification annually.
- Appoint an incident-management champion in each of the 28 teams to collect feedback and support local adoption.
- Conduct listening sessions focused on pager fairness, cross-team boundaries, tooling friction, and psychological safety.
- Publish that command duty is process coordination, not responsibility for understanding or repairing another team's code.
15. Pilot and expand the on-call model (depends on: 3, 5, 6, 8, 9, 14)
Pilot the model on the highest-risk customer journeys before expanding it. Correct staffing, alert, access, and compensation defects at each stage gate.
- Start with the ledger, payment orchestration, Kubernetes platform, PostgreSQL platform, authentication, settlement, and Support intake.
- Run the command, communications, and domain rotations in parallel with existing paths for two weeks.
- Verify primary and secondary coverage, handoffs, access, runbooks, paging, conference access, status publication, and compensation processing.
- Require at least two shadow shifts before independent primary duty.
- Review every pilot page within one business day for routing accuracy, actionability, responder load, and missing context.
- Expand by customer journey and dependency domain, not by arbitrary team order.
- Provide central first-line triage if useful, but keep technical remediation with the accepted service owner.
- Do not use contractors or a managed service as the sole incident commander or sole owner of payment and ledger remediation.
- Permit temporary shared domain rotations only after service owners document training, access, runbooks, and escalation boundaries.
- Set an executive-reviewed deadline and remediation plan for any production service that cannot provide sustainable 24x7 ownership.
16. Exercise regional, ledger, and communication failures (depends on: 10, 11, 14, 15)
Validate the process under realistic conditions before relying on it. Begin in tabletop and staging environments, then use controlled production tests where risk permits.
- Run a company-wide incident-command tabletop within 30 days of policy approval.
- Exercise loss of one AWS region, Kubernetes control-plane degradation, shared PostgreSQL failure, payment-processor failure, queue backlog, credential compromise, and suspected duplicate payments.
- Exercise a simultaneous operational and security event to test command boundaries and disclosure control.
- Exercise status-page failure and loss of the primary chat or paging provider.
- Exercise overnight staffing, role handoff, executive escalation, account-manager messaging, and a potential regulator-notification decision.
- Validate backups, restore procedures, RPO, RTO, failover prerequisites, and post-recovery reconciliation.
- Do not inject uncontrolled changes into the production ledger. Use replicas, staging, simulations, or tightly governed production tests.
- Record exercise observations as tracked actions under the same ownership and due-date rules as incident actions.
- Run at least one domain exercise per quarter and two cross-company exercises before the SOC 2 audit.
17. Execute a time-boxed enterprise rollout (depends on: 2, 6, 7, 11, 12, 13, 15)
Use fixed implementation waves so the audit deadline does not become the start date. Report progress weekly and escalate missed stage gates as business risks.
- Days 0–7: Establish governance, interim command coverage, one declaration path, provisional severity, and daily operational reviews.
- By day 14: Approve the core policy, role definitions, communications timings, compensation design, and historical-action triage.
- By day 30: Complete Tier 0 ownership, certify the first command roster, begin the on-call pilot, enable standard incident records, and run the first tabletop.
- By day 60: Provide 24x7 coverage for all Tier 0 and Tier 1 customer journeys, integrate the six alert sources, and enforce postmortem tracking.
- By day 90: Migrate critical paging, implement customer-journey detection, complete status and regulatory playbooks, and materially reduce alert noise.
- By day 120: Assign sustainable ownership and escalation for every production service and complete the first controlled regional or continuity exercise.
- By day 180: Complete tool retirement decisions, verify action closure, rerun weak scenarios, and demonstrate improving detection and mitigation trends.
- In month 7: Conduct a mock audit and executive readiness review, leaving at least one month to correct evidence or operating defects.
- Use exception records with owners, expiry dates, compensating controls, and executive approval. Do not allow indefinite verbal exceptions.
18. Build SOC 2 evidence as the process operates (depends on: 1)
Design evidence collection at the start rather than reconstructing it before the audit. Demonstrate both control design and sustained operation.
- Map the incident process to applicable SOC 2 criteria with Compliance and the auditor, including detection, response, communication, change management, access, availability, and corrective action.
- Maintain approved, version-controlled policies, procedures, severity definitions, role descriptions, and exception records.
- Preserve rotation schedules, compensation activation, training attendance, certification, paging tests, access reviews, and exercise results.
- Preserve incident declarations, timestamps, role assignments, communications, decisions, status updates, postmortems, and corrective-action evidence.
- Record regulatory and contractual notification assessments, including decisions that no notification was required.
- Define retention, confidentiality, legal-hold, and access requirements for operational and security records.
- Sample evidence monthly and trace incidents from initial signal through action verification.
- Have Internal Audit or an independent control owner test the process in months 4 and 6.
- Correct control failures through tracked actions rather than editing historical records.
- Conduct the formal mock audit in month 7 using the same evidence populations expected for the external audit.
19. Sustain accountability and continuous improvement (depends on: 13, 17, 18)
Make incident management an operating discipline rather than an audit project. Keep policy, staffing, tools, and investment aligned with changing customer and system risk.
- Assign permanent owners for the incident policy, paging platform, status page, service catalog, training program, and metrics.
- Review severity thresholds, communication timings, compensation, and staffing at least annually and after material incidents.
- Use incident trends to prioritize architectural work on the shared ledger, regional independence, deployment safety, dependency isolation, and graceful degradation.
- Review repeat incidents and repeat contributing factors quarterly. Require executive action when remediation repeatedly loses priority.
- Survey responders quarterly and publish actions addressing fatigue, fairness, psychological safety, and tool friction.
- Recognize effective incident leadership, early declaration, useful postmortems, and preventive work.
- Prohibit retaliation for good-faith incident declaration or escalation.
- Provide the board or risk committee a quarterly summary of severe incidents, SLA exposure, regulatory events, overdue high-risk actions, and resilience investment.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
--- PROPOSAL 3 (agent qwen3.8-max_initial_3, alibaba/qwen3.8-max) ---
Estimated complexity: high
Success metrics: - **MTTD reduced from 22 minutes to <5 minutes** within 6 months of full rollout.
- **Customer-first detection rate reduced from 40% to <10%** within 6 months.
- **MTTM for SEV1 incidents reduced from 3 h 10 min to <60 minutes** within 9 months.
- **Monthly SLA credit payouts reduced from ~$325K to <$100K per quarter** within 12 months.
- **Alert volume reduced from 3,400/month to <600 actionable alerts/month** within 6 months; signal-to-noise ratio >80%.
- **Postmortem completion rate: 100% of SEV1/SEV2 incidents** have a blameless postmortem within 5 business days.
- **Postmortem action-item completion rate >90% within 30 days** of the postmortem (up from ~17%).
- **Zero incidents with >15 minutes of unowned command** (down from 2 incidents with >1 hour).
- **100% on-call coverage**: all 28 teams staffed with primary + secondary on-call 24×7 within 14 weeks.
- **On-call compensation adopted**: 100% of on-call engineers receiving stipends and page pay; on-call satisfaction score ≥4/5 in quarterly survey.
- **Status-page first-update within 15 minutes for SEV1 and 30 minutes for SEV2**, 100% compliance.
- **SOC 2 Type II audit passed** at month 8 with zero incident-response findings.
- **All 260 engineers trained**; 56+ certified ICs and 28+ certified CLs active within 14 weeks.
- **Six alerting tools consolidated to one** within 6 months; legacy tools decommissioned.
- **Regulator notification process tested**: at least one tabletop exercise includes a NY DFS / FinCEN notification drill, and the legal/compliance playbook is documented and approved.
- **Quarterly IMC reviews held consistently** with published KPI dashboards and action-item tracking.
- **On-call participation resistance resolved**: <10% of engineers report 'unwilling to participate' in the 6-month pulse survey (baseline to be measured in S1).
Steps (13):
1. Assess Current State and Baseline Metrics
Build an evidence-based picture of the current incident management reality before designing anything new.
Collect and catalog the last 12 months of incident data: all 31 customer-impacting incidents, 3,400 monthly alerts, on-call coverage gaps across the 28 teams, and the 11 of 64 closed postmortem action items. Interview one lead from each of the 28 teams to surface pain points, political concerns (the 'carrying a pager for other teams' pushback), and tool sprawl.
Deliverables to produce:
- **Alert inventory**: which of the six alert tools feed which teams, alert volume per team, noise rate per tool, and overlap between tools.
- **Incident timeline analysis**: median detection-to-notification-to-mitigation-to-resolution times, who detected first (internal vs. customer), who mitigated, and where handoff gaps occurred.
- **On-call coverage map**: which 12 teams have on-call, which 16 do not, rotation length, compensation status, and escalation paths (or lack thereof).
- **Postmortem audit**: format variance, action-item tracking gaps, and the two incidents with no clear owner for over an hour.
- **Tooling and integration audit**: Kubernetes observability stack, six alerting tools, status-page provider, communication channels (Slack, email, phone), and any existing runbooks.
- **Compliance gap analysis**: SOC 2 Type II CC7.3/CC7.4 requirements vs. current practice, with a risk register for the eight-month window.
- **Peer benchmarking**: incident management practices at 3–4 comparable B2B fintech platforms (e.g., Plaid, Stripe, Adyen) for severity scales, on-call comp, and MTTR targets.
2. Secure Executive Sponsorship and Form the IM Governance Body (depends on: 1)
Anchor the program with visible, top-down authority so that 28 teams adopt changes they did not individually request.
The CEO email about 'outages we hear about from clients' is a ready-made mandate. Convert it into a formal sponsorship structure.
- Appoint an **executive sponsor** (CTO or VP Engineering) who owns the program end-to-end and reports to the CEO monthly.
- Create an **Incident Management Office (IMO)**: one dedicated senior incident management lead, one tooling/platform engineer, and one part-time data analyst.
- Establish an **Incident Management Council (IMC)**: one engineering manager from each of the 28 teams, plus the VP of Customer Success, a compliance lead, and a security lead. The IMC meets bi-weekly during rollout, monthly thereafter.
- Draft and circulate an **executive mandate memo** that states: incident response is a shared operational obligation, not a per-team favor; participation in on-call rotations is a condition of employment for production-facing roles; and the program is not optional pending the SOC 2 audit.
- Allocate a dedicated budget line for on-call compensation, tooling consolidation, status-page licensing, training, and external facilitation.
3. Define Severity Levels and Automatic Triggers (depends on: 1)
Replace the current ad-hoc triage with a five-level severity taxonomy that every engineer, support agent, and account manager can apply in under 60 seconds.
- **SEV1 – Critical**: Ledger data corruption or loss, complete payment processing halt, confirmed data breach affecting customer PII or funds, or regulatory reporting breach. Triggers: automatic all-hands page to on-call, CTO + CEO paged within 10 minutes, dedicated bridge call within 5 minutes, status-page update within 15 minutes, regulator notification assessment within 1 hour, customer comms within 30 minutes.
- **SEV2 – High**: Payment processing degraded >30% throughput or >5% error rate, single-region failover failure, ledger read-only mode, or any condition likely to breach the 99.95% SLA within the current window. Triggers: primary + secondary on-call paged, incident commander assigned within 10 minutes, bridge call within 15 minutes, status-page update within 30 minutes, VP Engineering notified within 20 minutes.
- **SEV3 – Medium**: Non-critical service degradation affecting <30% of customers, single non-ledger microservice outage with working fallback, or elevated latency above SLA threshold but below halt. Triggers: primary on-call paged, IM notified within 30 minutes, status-page update within 1 hour if customer-visible, daily standup update.
- **SEV4 – Low**: Degraded internal tooling, minor UX bug with workaround, non-customer-facing alert. Triggers: next-business-day response, ticket created, no page unless on-call agrees.
- **SEV5 – Informational / Noise**: Cosmetic issues, planned-maintenance notifications, alert misfires logged for tuning. Triggers: no page, logged for weekly alert-quality review.
Define **escalation rules**: any SEV3 unresolved after 4 hours auto-escalates to SEV2; any SEV2 unresolved after 2 hours auto-escalates to SEV1. Severity can be **downgraded** only by the incident commander with IMC notification.
Publish the taxonomy as a one-page decision tree, a Slack slash-command (`/sev`), and an integration into the alerting tool so that every alert carries a suggested severity.
4. Define Incident Roles and Staffing Model (depends on: 3)
Codify four mandatory roles for every SEV1/SEV2 incident and optional roles for SEV3, then solve the 24×7 staffing problem across 28 teams.
**Roles**
- **Incident Commander (IC)**: owns the incident end-to-end, declares severity, assigns tasks, authorizes mitigations, decides when to escalate or stand down. Never writes code during the incident.
- **Communications Lead (CL)**: owns status-page updates, internal Slack channels, account-manager briefings, and regulator notifications. Separate from the IC so the IC can focus on mitigation.
- **Scribe / Timeline Keeper**: logs every decision, action, and timestamp in the incident channel and the incident-management tool. Produces the raw timeline for the postmortem.
- **Subject-Matter Responders (SMRs)**: 1–3 engineers from the owning team(s) who diagnose and fix. For the shared PostgreSQL ledger, a dedicated DBA responder is always required.
**24×7 Staffing via a Three-Tier Follow-the-Sun Model**
- **Tier 1 – Front-line on-call**: Primary + secondary responder per team, paged first. Covers the team's own services.
- **Tier 2 – Platform / SRE on-call**: A dedicated 6-person SRE rotation covering cross-cutting infrastructure: Kubernetes, the shared PostgreSQL ledger, networking, and the two AWS regions. This tier directly addresses the 'carrying a pager for other teams' concern by absorbing infrastructure incidents.
- **Tier 3 – IMC escalation**: Engineering managers and the IMO on-call for multi-team or SEV1 incidents. Provides the IC and CL when no team-level IC is available.
**Follow-the-Sun**: If any engineering hub exists in a second timezone, use it for overnight Tier 1 coverage. If not, partner with a managed on-call service for overnight first-response triage (severity declaration + paging the correct team), reducing 3 a.m. pages for NY-based engineers.
**IC and CL pools**: Nominate at least 2 ICs and 1 CL per team (56 ICs, 28 CLs minimum). ICs are trained and certified before they rotate. For SEV1 incidents, the IC must be a certified IC from the IMC pool, not just 'whoever is around'.
**Ledger-specific rule**: Because the PostgreSQL ledger is shared, a **Ledger Duty Officer** from the SRE Tier 2 is always on the bridge for any incident touching ledger services, regardless of which team owns the failing microservice.
5. Design On-Call Rotations, Compensation, and Alert-Quality Rules (depends on: 4)
Make on-call sustainable, fairly compensated, and free of alert noise so engineers stop resisting participation.
**Rotation Design**
- 7-day rotations, one primary + one secondary per team per week. No engineer is on-call more than one week in four.
- Minimum 48-hour rest between rotations. No on-call during approved PTO.
- All 28 teams participate. Teams without current on-call get a 90-day ramp with a shadow rotation before going live.
- Tier 2 SRE rotation: 6 engineers, one week on / five weeks off, with a dedicated backup.
**Compensation Package**
- **Base on-call stipend**: $500 per week of primary on-call, $250 for secondary, paid regardless of whether pages fire.
- **Page pay**: $75 per acknowledged page outside business hours; $150 if the page leads to active incident work.
- **Time-off-in-lieu (TOIL)**: Any engineer who works >4 hours overnight (00:00–06:00 local) gets a full TOIL day. >8 hours in a single incident gets 1.5 TOIL days.
- **SEV1 bonus**: $300 flat bonus for every engineer who actively works a SEV1 incident, paid within the next pay cycle.
- **Annual on-call cap**: No engineer exceeds 13 weeks of on-call per year. Exceeding the cap triggers a mandatory team-staffing review.
- Budget estimate: ~$420K/year for stipends and page pay across 28 teams; present this to the CFO as a fraction of the $1.3M annual SLA credit cost.
**Alert-Quality Rules (the '85% noise' problem)**
- Every alert must carry: owning team, suggested severity, runbook link, and a 30-day noise score.
- **Alert budget**: each team gets a maximum of 100 actionable alerts per month. Exceeding the budget triggers a mandatory alert-tuning session with the IMO.
- **Noise threshold**: any alert that fires >10 times in 7 days with no human action is auto-flagged for suppression or tuning within 14 days.
- **Alert review cadence**: weekly 30-minute alert-quality review per team; monthly cross-team alert review in the IMC.
- **Sunset rule**: alerts with no runbook are demoted to SEV5 after 30 days and suppressed after 60 days unless a runbook is written.
- Target: reduce monthly alert volume from 3,400 to <600 actionable alerts within 6 months.
6. Build Detection, Escalation, and Communication Paths (depends on: 3, 4)
Eliminate the 22-minute median detection gap and the 40% customer-first-detection rate with layered monitoring and a single escalation spine.
**Detection Layers**
- **Synthetic transactions**: run a payment end-to-end through the full stack (API → service → ledger → confirmation) every 60 seconds from both AWS regions. Alert if latency >2× baseline or any step fails. This catches what per-service metrics miss.
- **Customer-traffic anomaly detection**: monitor API error rates, payment success rates, and latency percentiles per customer cohort. Alert on >2σ deviation.
- **SLO-based alerting**: define SLIs for the 99.95% SLA (availability, latency p99, ledger consistency). Alert when error budget burn rate exceeds threshold, before the SLA actually breaches.
- **Infrastructure health**: Kubernetes node/pod health, PostgreSQL replication lag, disk I/O, and cross-region latency.
- **Support-ticket spike detection**: if >5 customers open tickets about the same symptom within 10 minutes, auto-create a SEV3 candidate.
**Escalation Path**
- Alert fires → PagerDuty routes to Tier 1 primary → 5-min no-ack → Tier 1 secondary → 10-min no-ack → Tier 2 SRE → 15-min no-ack → IMC on-call manager → 20-min no-ack → VP Engineering auto-page.
- Any SEV1 declaration auto-pages the CTO, opens a dedicated Slack channel + Zoom bridge, and notifies the CL.
- **No incident goes unowned for >15 minutes.** If no IC is assigned by minute 15, the IMC on-call manager assumes IC role by default.
**Internal Communications**
- Dedicated Slack channels: `#inc-sev1`, `#inc-sev2`, `#inc-sev3` (auto-created per incident), plus `#inc-updates` for broadcast.
- IC posts a structured update every 15 minutes (SEV1), 30 minutes (SEV2), 1 hour (SEV3) into the incident channel.
- CL posts a summary to `#inc-updates` and notifies relevant engineering managers.
**Customer Communications**
- **Status page**: auto-updated via API. SEV1: first update within 15 minutes, then every 30 minutes until resolved. SEV2: first update within 30 minutes, then every hour. SEV3: within 1 hour if customer-visible.
- **Account managers**: CL briefs AMs via a dedicated Slack channel within 30 minutes (SEV1) or 1 hour (SEV2). AMs contact their top-20 revenue accounts directly.
- **Customer email/SMS**: for SEV1 and SEV2, automated notification to all 2,100 customers via the status-page subscription system within 30 minutes.
- **Regulator notification**: Legal/Compliance assesses within 1 hour whether NY DFS, FinCEN, or card-network notification is required. If yes, file within the regulatory deadline (typically 72 hours for NY DFS cybersecurity events). Log the decision and filing in the incident record.
**Post-resolution**: CL publishes a 'resolved' update within 30 minutes of mitigation. For SEV1/SEV2, a preliminary customer-facing RCA summary is published within 5 business days.
7. Standardize Postmortems with Tracking and Accountability (depends on: 3, 6)
Fix the '11 of 64 action items closed' problem with a mandatory, uniform, blameless postmortem process backed by engineering-manager accountability.
**When Mandatory**
- All SEV1 and SEV2 incidents: postmortem required within 5 business days.
- SEV3 incidents: postmortem required if >50 customers affected, if the incident lasted >4 hours, or if it was customer-detected.
- SEV4/SEV5: optional, but any recurring SEV4 (≥3 times in 30 days) triggers a mandatory review.
**Format (single template, enforced by tooling)**
- Incident summary (severity, duration, customers affected, revenue impact, SLA credit exposure).
- Timeline (auto-generated from scribe notes + alert timestamps).
- Detection analysis: how was it detected, why did it take X minutes, could it have been faster.
- Root cause analysis using 5-Whys or fault-tree, not blame.
- Contributing factors (process, tooling, staffing, knowledge gaps).
- Impact quantification: customers affected, transactions failed, SLA credits triggered.
- Action items: each with a **named owner**, **due date**, **priority**, and **ticket in Jira**.
- Lessons learned and what went well.
**Blameless Review Meeting**
- Held within 5 business days, facilitated by the IMO or a trained facilitator (never the IC of that incident).
- All responders, the CL, relevant engineering managers, and an IMC representative attend.
- Ground rules: focus on system and process failures, not individual mistakes. The facilitator enforces this.
- Meeting recorded; notes published to the engineering-wide wiki within 48 hours.
**Action-Item Tracking and Accountability**
- Every action item is created as a Jira ticket with a due date and a named owner.
- **Engineering managers are accountable**: action-item completion is a standing agenda item in the bi-weekly IMC meeting. Any action item >7 days overdue is escalated to the VP Engineering.
- **Completion gate**: no team may close its postmortem until 100% of its action items have Jira tickets. Postmortem is 'closed' only when all tickets are resolved.
- **Quarterly audit**: the IMO audits action-item completion rates and reports to the IMC and the executive sponsor. Target: >90% completion within 30 days of the postmortem.
- Link postmortem quality and action-item completion to team health scores and engineering-manager performance reviews.
8. Define KPIs, Dashboards, and Governance Reviews (depends on: 7)
Create a measurable feedback loop so leadership can see whether the program is working and where to intervene.
**Primary KPIs (tracked weekly, reported monthly)**
- **MTTD** (Median Time to Detect): target <5 minutes (from 22).
- **MTTA** (Median Time to Acknowledge): target <5 minutes.
- **MTTM** (Median Time to Mitigate): target <60 minutes for SEV1 (from 3 h 10 min), <4 hours for SEV2.
- **Customer-first detection rate**: target <10% (from 40%).
- **SLA compliance**: maintain 99.95%; track monthly SLA credit payouts, target <$100K/quarter (from ~$325K/quarter).
- **Alert signal-to-noise ratio**: target >80% actionable (from ~15%).
- **Alert volume**: target <600/month (from 3,400).
- **Postmortem completion rate**: 100% for SEV1/SEV2 within 5 business days.
- **Action-item completion rate**: >90% within 30 days (from ~17%).
- **On-call health**: pages per engineer per week (target <5), TOIL usage, on-call satisfaction survey score.
- **Unowned incident duration**: target 0 incidents with >15 minutes without an IC.
**Dashboards**
- Real-time operational dashboard (Grafana): current incidents, active alerts, on-call roster, SLA error-budget burn.
- Weekly leadership dashboard (auto-generated): KPI trends, open action items, alert-noise report, on-call load distribution.
- Quarterly IMC scorecard per team.
**Review Cadence**
- **Weekly**: IMO publishes KPI snapshot to `#inc-updates`.
- **Bi-weekly IMC**: review open incidents, overdue action items, alert-quality exceptions, and on-call load.
- **Monthly executive review**: CTO presents KPI trends, SLA credit cost, and risk register to the CEO.
- **Quarterly incident-management review**: deep-dive into trends, training gaps, tooling needs, and process improvements. Output fed into the next quarter's roadmap.
9. Consolidate Tooling and Build the Incident Management Platform (depends on: 2, 3)
Replace six alerting tools and ad-hoc status-page updates with a single, integrated incident management stack.
**Target Tool Architecture**
- **Single alerting and on-call platform** (e.g., PagerDuty or Opsgenie): ingest all alerts, apply severity routing, manage on-call schedules, handle escalations, and send pages. Retire the other five tools within 6 months.
- **Observability consolidation**: standardize on one APM/metrics stack (e.g., Datadog or Grafana Cloud) for all 180 Kubernetes services across both AWS regions. Ensure the shared PostgreSQL ledger has dedicated dashboards.
- **Status page**: a dedicated, branded status page (e.g., Statuspage.io or Instatus) with API integration for auto-updates. Subscribe all 2,100 customers.
- **Incident coordination tool**: integrate incident-management workflows into Slack (auto-create channels, invite responders, post templates) and a dedicated incident record system (e.g., Jira Service Management, incident.io, or Rootly) for timelines, postmortems, and action-item tracking.
- **Runbook repository**: a central wiki (Confluence or Notion) with mandatory runbooks for every alert. No alert goes live without a linked runbook.
**Implementation Tasks**
- Migrate all 28 teams' alert rules into the single platform in three waves (highest-volume teams first).
- Build the severity-based routing rules and escalation policies per S3 and S6.
- Automate status-page updates triggered by severity declaration.
- Build the synthetic-transaction monitor and SLO-based alerting per S6.
- Integrate Jira for automatic action-item ticket creation from postmortems.
- Decommission legacy tools only after all teams have completed training on the new stack.
- Budget: allocate $150K–$250K/year for licensing, plus engineering time for migration.
10. Prepare for the SOC 2 Type II Audit (depends on: 7, 8, 9)
Ensure the incident management process produces the evidence the auditor will need, well before the audit window opens in eight months.
**SOC 2 Requirements to Address (CC7.3, CC7.4, CC7.5)**
- Documented incident response procedures (the severity taxonomy, role definitions, communication templates).
- Evidence of incident detection, response, and recovery for every SEV1/SEV2 incident during the audit period.
- Postmortem records with action-item tracking.
- On-call schedules, training records, and escalation evidence.
- Status-page update logs and customer notification records.
- Regulator notification logs (if any).
**Preparation Tasks**
- The IMO maintains a **SOC 2 evidence folder**: every incident record, postmortem, action-item ticket, status-page update, and training completion certificate is stored and indexed.
- Conduct a **mock SOC 2 audit** at month 5: an internal or external auditor reviews the incident management process end-to-end and identifies gaps.
- Remediate mock-audit findings before month 7.
- Ensure the incident management tool retains all records for at least 12 months (the SOC 2 Type II observation window).
- Document the **chain of custody** for incident records: who accessed, modified, or closed each record.
- Prepare a **narrative document** describing the incident management process, roles, and controls for the auditor.
- Coordinate with the compliance lead to align incident management evidence with the broader SOC 2 scope (access controls, change management, etc.).
11. Design and Deliver Training, Runbooks, and Change Management (depends on: 4, 5, 9)
Equip all 260 engineers, 28 team leads, account managers, and support staff with the knowledge and muscle memory to execute the new process.
**Training Tracks**
- **All 260 engineers** (2-hour session): severity taxonomy, how to acknowledge a page, how to join an incident bridge, how to hand off to an IC, and how to write a postmortem contribution. Delivered in team-level sessions over 4 weeks.
- **IC pool (56+ engineers)** (8-hour certification): incident command techniques, severity declaration, escalation decision-making, bridge facilitation, and blameless postmortem facilitation. Includes two tabletop exercises. Certification valid for 12 months, renewed annually.
- **CL pool (28+ staff)** (4-hour session): status-page writing, customer communication templates, regulator notification triggers, and AM briefing protocol.
- **Account managers and support staff** (1-hour session): how to read the status page, how to escalate a customer report into an incident, and what information to collect.
- **SRE Tier 2** (16-hour onboarding): Kubernetes and PostgreSQL ledger deep-dive, cross-region failover runbooks, and escalation authority.
**Runbooks**
- Every alert must have a runbook before it is routed to on-call. The IMO provides a runbook template and audits compliance weekly.
- Priority runbooks to write first: shared PostgreSQL ledger failover, Kubernetes cluster degradation, payment-processing pipeline failure, cross-region failover, and ledger data-integrity check.
- Runbooks are peer-reviewed and version-controlled.
**Change Management for Adoption**
- Address the 'carrying a pager for other teams' concern directly: publish an FAQ explaining the three-tier model, the SRE Tier 2 absorbing cross-team infrastructure, the compensation package, and the TOIL policy.
- Run **office hours** weekly for the first 8 weeks where any engineer can ask questions or raise concerns.
- Identify **team champions**: one engineer per team who volunteers as an early adopter and peer mentor.
- Publish a **weekly 'incident management newsletter'** during rollout: what changed, what improved, KPI trends, and success stories.
- Make on-call participation a documented expectation in job descriptions and performance reviews for production-facing roles.
12. Execute Phased Rollout, Tabletop Exercises, and Continuous Improvement (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11)
Introduce the process in three waves so teams are not overwhelmed, then validate with exercises and iterate continuously.
**Phase 1 – Weeks 1–6: Foundation**
- Publish the severity taxonomy, role definitions, and communication protocols (S3, S4, S6).
- Launch the single alerting platform for the 12 teams already on-call; begin migration for the other 16.
- Activate the SRE Tier 2 rotation for the shared PostgreSQL ledger and cross-cutting infrastructure.
- Deploy the status page and test the API integration.
- Begin IC and CL training (first cohort of 20 ICs, 10 CLs).
- Publish the on-call compensation package; HR integrates stipends into payroll.
- Write the top 10 priority runbooks.
**Phase 2 – Weeks 7–14: Expansion**
- All 28 teams live on the single alerting platform; legacy tools in read-only mode.
- All 28 teams on the on-call rotation schedule (the 16 new teams in shadow mode for the first 4 weeks).
- Second and third IC/CL training cohorts completed.
- First **tabletop exercise**: simulate a SEV1 ledger corruption scenario with all roles, test the escalation path, status-page updates, and AM briefings. Debrief and fix gaps.
- Postmortem template and Jira integration live; all new incidents use the standard process.
- Alert-tuning sprint: each team reduces its alert volume by 50%.
**Phase 3 – Weeks 15–24: Optimization**
- All teams fully live; legacy alerting tools decommissioned.
- Second **tabletop exercise**: simulate a SEV1 cross-region failure with regulator notification.
- First quarterly IMC review with full KPI dashboard.
- Mock SOC 2 audit (month 5) and remediation.
- Retrospective on the rollout: survey all 260 engineers for feedback, adjust compensation or rotation rules if needed.
- Establish the **continuous improvement cadence**: quarterly process review, annual severity-taxonomy review, and annual on-call compensation benchmarking.
**Ongoing Governance**
- The IMC owns the process document and approves changes.
- The IMO tracks all KPIs and reports to the CTO monthly.
- Any process change requires IMC approval and a 2-week notice period before enforcement.
- Annual external benchmarking against peer B2B payments platforms.
13. Establish Ongoing Governance, Annual Review, and Audit Readiness Cycle (depends on: 12)
Embed incident management as a permanent organizational capability, not a one-time project.
- **Annual process review**: the IMC reviews the severity taxonomy, role definitions, on-call structure, and compensation against industry benchmarks and internal KPIs. Update as needed.
- **Bi-annual tabletop exercises**: one SEV1 infrastructure scenario, one SEV1 data-breach/regulator scenario. Rotate the IC and CL assignments so everyone gets practice.
- **Quarterly alert-quality audit**: the IMO reviews alert volumes, noise rates, and runbook coverage across all 28 teams.
- **On-call health survey**: quarterly anonymous survey measuring burnout, fairness, and compensation satisfaction. Results reviewed by the IMC.
- **SOC 2 readiness cycle**: begin evidence collection immediately after each audit ends. The IMO maintains a rolling evidence folder. Mock audit at month 5 of every 12-month audit cycle.
- **Postmortem maturity tracking**: track the action-item completion rate monthly. If it drops below 80%, the VP Engineering intervenes.
- **Incident management maturity model**: adopt a 5-level maturity model (ad-hoc → defined → managed → optimized → predictive). Assess annually. Target: Level 3 within 12 months, Level 4 within 24 months.
- **Budget review**: annually review on-call compensation, tooling costs, and training budget against the reduction in SLA credits and incident frequency.
--- PROPOSAL 4 (agent grok4.6_initial_4, xai/grok-4.6) ---
Estimated complexity: high
Success metrics: - Median time to detect customer-impacting incidents ≤ 5 minutes within 6 months of go-live.
- Share of SEV-1/SEV-2 incidents first detected by customers ≤ 5% (from 40%).
- Median time to mitigate SEV-1/SEV-2 ≤ 45 minutes (from 3 h 10 min).
- Named Incident Commander assigned within 5 minutes for ≥ 95% of SEV-1/SEV-2.
- First status-page update within policy time for ≥ 95% of SEV-1/SEV-2.
- SLA credits down ≥ 80% versus the trailing $1.3M within 12 months.
- Paging volume ≤ 500 per month and noise ≤ 15% (from 3,400 and 85%).
- 100% of production services have a named owning team and a paging policy.
- Postmortems filed within 5 business days for 100% of SEV-1/SEV-2; action-item close rate ≥ 80% within 30 days.
- 24×7 IC and critical-path coverage with zero unfilled shifts per quarter.
- Paid on-call live for every rotation before that rotation pages humans.
- SOC 2 Type II incident-response controls evidenced for ≥ 5 months before the auditor's report.
- On-call pulse: ≥ 70% of engineers agree rotations are fair and limited to their services.
Steps (30):
1. Secure executive mandate and budget
Get a written CEO/CTO mandate that incident command is a company process, not a team hobby.
The mandate must state that **paid on-call** is required for production ownership. "No pager for other teams' code" is solved by named ownership, not by refusing coverage.
- Approve budget for tooling, stipends, training, and a dedicated program lead for six months.
- Name an executive sponsor (CTO or VP Engineering) who will chair the weekly incident review.
- Tie the clock to SOC 2 Type II: the process must be live in about 10 weeks so ~6 months of evidence remain.
- Commit that the CEO will hear about outages from this process, not from customers.
2. Form the working group and decision rights (depends on: 1)
Stand up a small group that can decide. Do not form a 28-team committee.
**Core seats:** SRE/platform lead, payments/ledger engineering manager, support lead, legal/compliance, HR, one rotating team EM, and a program manager.
- Meet twice a week for 10 weeks, then weekly.
- RACI: the group proposes; the sponsor decides in 48 hours; teams implement.
- Publish one Slack channel and one source-of-truth doc on day one.
- Time-box design to four weeks. Ship v1 rather than wait for consensus.
3. Inventory services, owners, and on-call gaps (depends on: 1)
Build a living catalog of all ~180 services: owning team, criticality, current on-call, alert sources, and runbook link.
Walk the last 31 customer-impacting incidents and the two events where **nobody was in charge**. Record who detected, who led, time to mitigate, and which alerts fired.
- Tag each service as critical-path, customer-visible, or internal.
- List the 16 teams with no on-call and every orphan service with no owner.
- Map all six alerting tools and the 3,400 monthly alerts onto services.
- Flag the shared PostgreSQL ledger and two-region failover as named special cases.
4. Map regulatory and contractual notification duties (depends on: 1)
Legal and compliance list every duty an incident can trigger. Do not invent clocks that violate a contract.
Cover SOC 2 CC7, customer MSA/SLA credit terms, money-transmitter rules, NYDFS 23 NYCRR 500 if applicable, PCI if in scope, and breach clocks.
- Extract **notification timings** from the largest customer contracts (status page, named AM, written notice).
- Define when legal, regulators, insurers, or the board must be told.
- Feed these clocks into severity triggers and the communications playbook.
5. Approve paid on-call and incident pay (depends on: 1, 4)
Unpaid on-call is why 16 teams refuse the pager and why nights are uncovered. Fix the money before asking for coverage.
HR, legal, and finance design a New York–compliant package: weekly stipend for primary and secondary, extra stipend for company IC and comms, and **after-hours incident pay** or comp time.
- Treat exempt vs non-exempt staff explicitly under NY wage-hour rules.
- Put stipend in the next pay cycle after policy publish, not "later."
- Cap consecutive night weeks. Fund hiring if a team cannot rotate fairly (minimum six people for 24x7 primary plus secondary).
- Publish the package before any new rotation starts. This is the main answer to pager pushback.
6. Ratify four severity levels and their triggers (depends on: 3, 4)
Adopt a business-impact scale. Engineers do not invent severity in the moment.
**SEV-1:** material payments failure; ledger down or inconsistent; security or customer-data incident; both regions impaired; or many customers already in SLA-credit territory.
**SEV-2:** degraded payments or a contracted feature down for multiple customers; SLA at risk.
**SEV-3:** narrow or single-customer impact with a workaround; no fleet-wide SLA risk.
**SEV-4:** no customer impact; ticket only.
- SEV-1 pages IC, comms, scribe, owning SMEs, and an exec; war room in 5 minutes; status page in 10; AM outreach in 20.
- SEV-2 pages IC and owning SMEs; comms may be the IC; status page in 15 minutes; updates every 30 minutes.
- SEV-3 pages the owning team only; customer notice only if that customer is affected.
- Anyone may declare. Only the IC may downgrade. When unsure, start high.
7. Define incident roles and their authority (depends on: 6)
Four roles. Separate coordination from debugging so "who is in charge" cannot stall for an hour again.
- **Incident Commander:** owns severity, the room, the clock, and the next action. Does not write code. May page anyone, freeze deploys, and invoke failover. Staffed from a company-wide trained pool, not from the failing team.
- **Communications Lead:** status page, customers, AMs, execs, regulators. Speaks only from IC-approved facts.
- **Scribe:** timeline in the incident tool. Required for SEV-1 and SEV-2.
- **SME responders:** the owning team's on-call. They mitigate. They do not run the room.
Publish a one-page authority card. The IC stays in charge even if a VP joins.
8. Design 24x7 coverage without 28 night rotations (depends on: 3, 5, 7)
Do not put 28 teams on 24x7. That is what engineers are rejecting.
Use **three layers** so people page for their own code, plus a trained commander.
- Layer A — company IC and SEV-1 comms: 24x7; about 24 trained people; week-long primary and secondary.
- Layer B — critical-path team on-call (ledger, payments processing, auth, API edge, platform/Kubernetes, data stores): 24x7 primary plus secondary.
- Layer C — all other teams: business-hours on-call; after hours the IC pages the team EM, who has a written escalation list.
Platform on-call is the safety net for unknown-owner pages, never the permanent owner. Every service must have a named team within 60 days or be scheduled to shut off.
9. Set rotation, handoff, and load rules (depends on: 8)
Write mechanical rules so rotations are fair and load is visible.
Primary week, then secondary week, then at least two weeks off. No one holds primary on two rotations at once.
- Handoff is a 30-minute overlap covering open incidents, silenced alerts, and upcoming changes.
- Page-load SLO: p50 ≤ 4 pages per 12-hour night shift; p95 ≤ 10. A breach opens an alert-quality action.
- Require a shadow week before a first IC shift or a first critical-path rotation.
- Swaps live in the paging tool. Managers own coverage gaps, not the last person on the roster.
10. Write detection and escalation paths (depends on: 6, 9)
Customers currently detect 40% of incidents and median time to detect is 22 minutes. That is the first failure mode.
Detection path: synthetic full-payment probes in both regions, SLO burn-rate alerts, support-to-incident intake, and one customer callback path that can create a SEV.
- A page must be acked in **5 minutes** or it auto-escalates to secondary, then IC, then the EM, then the VP.
- Support may declare SEV-2 or higher without engineering permission.
- If ownership is unclear for 10 minutes, the IC keeps the incident and assigns a temporary owner. Never wait.
- An exec bridge auto-opens for every SEV-1 at T+15 minutes.
11. Write internal, customer, and regulator communications (depends on: 4, 6, 7)
Stop "whoever is around" from writing the status page. Comms follow the clock, not convenience.
Timings from declaration:
- Internal war room: immediate. Exec summary for SEV-1/2 at 15 minutes, then every 30 minutes.
- **Public status page:** SEV-1 in 10 minutes, SEV-2 in 15. Updates at least every 30 minutes until resolve. Templates only. No speculation.
- Account managers get an affected-customer list and a script at T+20 minutes for SEV-1/2.
- Resolve notice and credit assessment within one business day.
- Comms pages legal on SEV-1 security, ledger integrity, or any outage that will breach contractual notice. Legal owns outbound regulatory letters; the IC owns facts.
12. Standardize blameless postmortems and action tracking (depends on: 6)
A written postmortem is mandatory for every SEV-1 and SEV-2 within 5 business days. SEV-3 if the IC or EM requests it.
Use one template: timeline, customer impact (volume, duration, credits), detection gap, what went well, what did not, process-focused five whys, and numbered actions with owner and due date.
- Review is **blameless** and scheduled. The IC attends. The exec sponsor reads every SEV-1.
- Actions live in one tracker, not in the doc. No action without an owner and a date. Default due date 14 days; 30 days max unless architecture work with a milestone.
- Close rate is a published metric. The old 11-of-64 pattern is a process failure.
13. Set alert quality rules that make paging acceptable (depends on: 3, 6)
3,400 alerts a month and 85% noise is why on-call feels like punishment. Pages are a product with a quality bar.
A page (not a ticket) must map to a customer-facing SLO or a hard dependency of one. It must have an owner team, a runbook link, and a default severity. It must be actionable at 3 a.m. by the person who is paged.
- Ban parallel paging from six tools. One paging policy: symptom-based; burn-rate preferred over raw thresholds.
- Every team gets a monthly noise budget. Exceeding it is a sprint task, not heroics.
- A human may silence a flapping alert only with a linked ticket.
14. Publish Incident Management Policy v1 (depends on: 5, 6, 7, 8, 9, 10, 11, 12, 13)
Collapse the design into a short policy people will open during an outage.
Ten pages or fewer, plus one-page cards for severity, roles, and comms timings. Host it where the incident tool can link it.
- Include the compensation summary and the rule: you are **not on-call for other teams' services**.
- Version it. v1 is mandatory from the pilot start date.
- Legal, HR, and the exec sponsor sign. Announce in all-hands, not only in Slack.
15. Implement a single incident command tool (depends on: 7, 10, 11)
Put one tool in the path that creates the room, pages roles from severity, records the timeline, and prompts status-page updates.
Requirements: Slack (or equivalent) incident bot, severity in one click, role assignment, stakeholder groups, and timeline export for postmortems and auditors.
- Integrate with the pager so IC, comms, and SME pages are automatic.
- Retain artifacts at least one year for SOC 2.
- Ad-hoc Zoom/Slack threads are no longer the system of record.
16. Consolidate six alerting tools onto one pager (depends on: 9, 13)
Pick one paging product. Connect existing monitors to it. Migrate **pages** first, tickets second.
- Inventory every page-producing rule. Delete or downgrade the noisy majority in S23.
- Route by service label → owning team schedule → the escalation policy from S10.
- IC and comms schedules live in the same product.
- Set a hard date after which pages outside the chosen tool are not valid on-call obligations.
17. Operationalize the status page and AM path (depends on: 11, 15)
Put the status page behind the Comms Lead role. Use templates for investigating, identified, mitigating, and resolved.
Subscribe AMs and customers whose contracts require it. Generate the affected-customer list from the incident (tenant, region, payment method).
- Dry-run a SEV-2 update before the pilot goes live.
- Record every public update in the incident timeline for the audit.
- Define partial vs full-outage wording so impact cannot be understated.
18. Stand up one postmortem repo and action board (depends on: 12)
Create the template, the filing location, and one Jira/Linear board with states: open, in progress, blocked, done, and won't-do (with exec reason).
Wire the incident tool so a SEV-1/2 automatically opens a draft postmortem and action tickets.
- Each week the program manager reports actions past due to the exec sponsor.
- If it is not in the board, it does not exist.
19. Assign owners and write critical-path runbooks (depends on: 3, 9)
Close the ownership gaps that cause "pager for other people's code."
Every production service gets a team in the catalog. Unowned services get an owner in 30 days or a decommission date.
- Write runbooks for ledger Postgres, regional failover, payments API, auth, and the Kubernetes control plane: symptoms, dashboards, mitigate vs escalate, customer impact.
- Link runbooks from alerts. If there is no runbook, the alert cannot page at night unless the EM accepts the gap in writing.
20. Detect payments failures before customers do (depends on: 3, 13)
Build or finish end-to-end synthetics: create payment, ledger write, webhook, both regions, both critical payment methods.
Alert on SLO burn, not on a single 500. Page SEV-2 or SEV-1 from these probes. This is the fastest lever on the 22-minute MTTD and the 40% customer-detected rate.
- Add ledger lag, replication, disk, and failover-readiness as first-class pages to the ledger team.
- Every postmortem asks first: **why did a customer see this first?**
21. Train the first cadre of ICs, comms leads, and scribes (depends on: 7, 14, 15)
Train about 24 ICs and 12 comms leads before the pilot. Classroom plus a recorded shadow of a simulated SEV-1.
Curriculum: severity, authority card, tool, comms timings, when to call legal, how to run a room of 20, how to hand off at 2 a.m.
- Certification: pass a tabletop. No certificate, no rotation.
- Recertify yearly and after any SEV-1 where process failed.
- Managers of ICs protect calendar time. This is part of the job.
22. Pilot on the payments critical path for six weeks (depends on: 5, 14, 15, 16, 17, 18, 19, 21)
Go live with policy, tool, paid rotations, and IC coverage for ledger, payments, API, platform, and support intake.
Keep old paths as backup for one week, then cut over. Real incidents use the new process only.
- Staff the program lead in every SEV-2+ as coach, not as secret IC.
- Collect friction daily. Fix tooling and wording in 48 hours.
- Expansion gate: named IC in under 5 minutes, first status update on time, no unpaid pages, postmortem filed.
23. Cut alert noise with a forced burn-down (depends on: 13, 16)
Give every team a numbered list of their noisiest alerts. Move noise from 85% to **under 15%**, and monthly pages from 3,400 toward 500.
Each sprint, critical-path teams must delete, debounce, or convert to ticket a fixed quota. Platform provides burn-rate and grouping libraries.
- Publish a weekly noise leaderboard. Shame systems, not people.
- After eight weeks, any page without a runbook or with >30% false pages in 14 days is auto-downgraded until fixed.
24. Roll remaining teams onto the model by risk (depends on: 22)
After the pilot gate, add teams in waves of four to five every two weeks. Highest customer-impact first.
Layer C teams get business-hours schedules and the EM night list. Do not surprise anyone with a pager.
- Each wave: ownership confirmed, alerts routed, runbooks for paging alerts, paid rotation in HR, one tabletop.
- Finish all 28 teams at least five months before the SOC 2 report date so the observation window covers the company.
- Orphan services still unowned at wave end are escalated to the sponsor for shutdown or reassignment.
25. Run tabletops and multi-region game days (depends on: 21, 22)
Schedule a monthly tabletop: SEV-1 ledger, SEV-1 region loss, SEV-2 degraded payments, customer-detected incident, and a "who is in charge" chaos drill.
Quarterly game day: fail a region or a ledger replica in staging or a controlled production drill.
- Include support, AMs, legal, and an exec. Process fails if only engineers show up.
- Capture actions on the same board as real postmortems.
- Use results as SOC 2 evidence that IR is tested.
26. Resolve ownership fights and pager culture (depends on: 5, 8, 22)
Treat pushback as design input, not defiance. Repeat the contract in office hours: you carry a pager for **your** services; IC is coordination; nights are paid; noise is a defect.
- EMs who cannot staff a fair rotation get headcount or have services reassigned. Do not run two-person 24x7.
- Publicly close the two historical "nobody in charge" incidents with what would be different now.
- Pulse-survey on-call at 60 and 120 days. If load or fairness is red, stop expansion until fixed.
27. Launch metrics, weekly review, and error budgets (depends on: 15, 22)
The CEO email exists because there was no operating rhythm. Stand up a dashboard and use it.
Track incident count by SEV, MTTD, MTTA, MTTM, customer-detected %, time-to-IC, time-to-status, pages per shift, noise %, SLA credits, postmortem SLA, and action close rate.
- **Weekly 30-minute review** chaired by the exec sponsor. Last week's incidents, open actions, alert SLO breaches. Mandatory for EMs of teams that had SEV-2+.
- Error budgets on critical SLOs: burn too fast and the team stops features and pays down reliability.
- Monthly report to the CEO so customer-first news does not recur.
28. Collect SOC 2 Type II evidence as a side effect (depends on: 4, 14, 22, 27)
Do not create a parallel "audit process." The real process is the evidence.
From pilot start, retain policy versions, on-call schedules, incident records, status-page history, postmortems, action tickets, training attendance, tabletop notes, and access reviews of the incident and pager tools.
- Map each artifact to SOC 2 CC7.2–CC7.5. Compliance owns the matrix; engineering owns completeness.
- Internal audit or a consultant runs readiness reviews at week 12 and week 20, not the week before the auditor.
29. Inspect and adapt the process after 90 days (depends on: 22, 27)
After 90 days live, revise v1 using data, not opinions.
Look at MTTD/MTTM, uncovered shifts, IC burnout, missed status updates, action close rate, and survey results.
- Change severity definitions if teams inflate or deflate them.
- Adjust Layer B vs C membership from real pages.
- Cut steps nobody follows. Add only what the last 90 days proved missing.
- Freeze v2 as the SOC 2 described process unless a defect is found.
30. Rehearse the audit and freeze the observation window (depends on: 24, 28, 29)
Six to eight weeks before the auditor, run a mock interview with an IC, a random engineer, support, and compliance.
Show one SEV-1 and one SEV-2 end-to-end: detect, page, command, comms, postmortem, closed actions.
- Fix evidence gaps immediately.
- Freeze process wording for the rest of the observation window; log exceptions.
- Brief the CEO and customer success with the metrics so the "outages we hear about from clients" story is retired.
--- PROPOSAL 5 (agent deepseek-v4-pro_initial_5, deepseek/deepseek-v4-pro) ---
Estimated complexity: high
Success metrics: - Customer-detected incidents decrease from 40% to less than 15% within six months.
- Median time to detect (MTTD) is under 5 minutes for SEV1 and SEV2 incidents.
- Median time to mitigate (MTTM) is under 60 minutes for SEV1 and under 2 hours for SEV2.
- Alert noise decreases from 85% to below 10% within three months.
- 100% of SEV1 and SEV2 incidents have a completed blameless postmortem within 5 business days.
- 100% of postmortem action items are tracked with an owner and due date; 90% are completed on time.
- 24x7 on-call coverage achieved across all 28 teams with no unpaid on-call.
- 99% of pages are acknowledged within 5 minutes.
- SLA credits paid reduce by at least 50% over the next 12 months.
- SOC 2 readiness: all incident response controls are documented, tested, and evidence is produced by month 7.
Steps (14):
1. Baseline current incident response and align stakeholders
Collect data from the last 12 months of incidents, all six alert tools, on-call practices, and team interviews. Identify gaps against the target incident management process and secure executive sponsorship.
- Gather incident timeline, detection source, mitigation time, customer impact, SLA credit, and postmortem status for all 31 incidents.
- Survey 28 teams on on-call burden, alert quality, and operational pain.
- Map current tools, escalation paths, and communication workflows.
- Create baseline metrics and a stakeholder map with executive sponsor and audit owner.
2. Define severity levels and response triggers (depends on: 1)
Define a four-level severity scale with objective business-impact criteria so any engineer can classify an incident consistently.
- SEV1: widespread transaction processing outage, data breach, security incident, or severe SLA breach; triggers full incident command, executive notification, and a 5-minute status page update.
- SEV2: major feature outage, significant degradation without workaround, or customer financial risk; triggers incident commander, full communications role, and status page updates.
- SEV3: partial impairment with workaround or limited customer impact; triggers on-call response, internal communication, and optional status page update.
- SEV4: minor or internal issue, no customer impact; handled during business hours through ticketing.
- Include an escalation matrix showing who can declare, downgrade, and invoke regulatory or legal involvement.
3. Define incident roles and decision authority (depends on: 1, 2)
Define incident roles, responsibilities, and decision authority using RACI to remove ambiguity about who is in charge.
- Incident Commander: owns the incident, declares severity, and coordinates resolution.
- Communications Lead: owns internal and external messaging, status page updates, and account manager notifications.
- Scribe: maintains timeline, incident log, and postmortem notes.
- Subject-matter responders: diagnose and fix the incident; may come from multiple teams.
- Executive sponsor: optional for SEV1; customer liaison: handles account managers.
- Define decision rights for severity declaration, escalation, rollback, customer communications, and incident closure.
4. Design 24x7 staffing model across 28 teams (depends on: 3)
Design 24x7 coverage across 28 teams without overloading engineers. Use service-based on-call plus a central incident command pool.
- Each service or domain team assigns primary and secondary on-call for its own services.
- Create central incident commander, communications, and scribe rotations staffed from a trained incident response guild across all teams; use follow-the-sun between the two AWS regions and time zones.
- Define escalation layers: service on-call to team lead or manager to service owner to executive.
- Define handoff times, shadow shifts, and load balancing; target at most one week of on-call per engineer per month.
- Bridge the current 12-team paid on-call to 28-team paid coverage; no team remains uncovered.
5. Define on-call rotations, compensation, and alert quality rules (depends on: 4)
Define sustainable rotations, pay, and rules that eliminate noisy pages.
- Rotations: weekly or biweekly, at least one primary and one secondary, with 12-hour shifts where possible or 24-hour for low-volume services.
- Compensation: monthly on-call stipend for all on-call engineers, additional incident response bonus for after-hours work, and time off in lieu; align with market rates.
- Alert quality rules: every page must be actionable, have a runbook link, specify a service owner, include severity, and be based on SLO burn or known failure signals; no dashboard-only alerts.
- Noise budget: reject or downgrade non-actionable alerts; all pages must go to on-call only after suppression and deduplication.
- Weekly alert review removes the top noisy alerts.
6. Design detection, escalation, and alert routing (depends on: 2, 4, 5)
Define how incidents are detected, routed, and escalated so nothing waits on a human to notice.
- Consolidate the six alert tools into one alerting and paging platform with routing by service, severity, and tags.
- Detection sources: infrastructure metrics, application synthetic transactions, log-based anomalies, business transaction SLI monitoring, and customer-reported issues through support or account managers.
- Routing: alert is paged to service on-call within 30 seconds; primary must acknowledge within 5 minutes; if no ack, page secondary then on-call manager.
- Escalation timeouts: unresolved SEV1 escalates to service owner at 15 minutes and to leadership at 30 minutes; any engineer can escalate to the incident commander.
- Define customer-reported incident intake and classification in the same tool.
7. Define internal and external communication protocols (depends on: 2, 3)
Define communication channels, templates, and timing for internal, customer, and regulator audiences.
- Internal: dedicated incident Slack channel, internal status page mirror, and war room bridge for SEV1; incident commander and communications lead own these channels.
- Status page: SEV1 post within 5 minutes, updates every 30 minutes or on material change, resolution within 60 minutes of mitigation; SEV2 post within 15 minutes, updates hourly; SEV3 optional.
- Account managers: SEV1 and SEV2 notify account managers within 15 minutes with an approved customer-facing description and expected impact.
- Regulators: legal or compliance determines notification for data breaches, security incidents, funds availability issues, or regulatory reportable events; criteria and timing follow legal and regulatory requirements; communications lead coordinates.
- Use pre-approved message templates and an approval chain; no ad-hoc wording.
8. Define postmortem policy and action tracking (depends on: 3)
Define mandatory blameless postmortems and action tracking.
- Mandatory for all SEV1 and SEV2 incidents, and any SEV3 that breaches SLA or is customer-detected.
- Format: impact, timeline, root causes, contributing factors, detection and response gaps, what worked well, and action items.
- Blameless: focus on system and process causes, not individual blame; use trained facilitators.
- Ownership: each action has an owner, due date, and tracking ID in a single backlog.
- Review postmortems at the weekly incident review; track action closure; expect 100% completion.
- Complete postmortems within 5 business days for SEV1 and SEV2 incidents.
9. Define metrics, dashboards, and review cadence (depends on: 2, 3, 8)
Define metrics and review cadence to measure process health.
- Metrics: MTTD, MTTM, customer detected percentage, alert noise percentage, on-call response time, on-call load, SLA credits paid, and postmortem action completion.
- Dashboards: real-time operational dashboard for on-call engineers and management.
- Weekly incident review: review all SEV1 and SEV2 incidents, action items, and noisy alerts.
- Monthly trends with leadership; quarterly review against SLOs and audit controls.
- Success thresholds: MTTD under 5 minutes, MTTM under 60 minutes for SEV1, customer detected under 15%, and alert noise under 10%.
10. Configure incident tooling and integrations (depends on: 5, 6, 7, 8, 9)
Implement and integrate the tools that automate the defined process.
- Aggregate alerts from the existing six tools into PagerDuty, Opsgenie, or a similar platform.
- Configure on-call schedules, escalation policies, and paging targeted at service owners.
- Integrate status page API for automated or one-click updates.
- Add Slack commands to declare incidents, start war rooms, assign roles, and post status updates.
- Integrate runbook and service catalog access; create postmortem templates in Jira or Notion with action item tracking.
- Ensure audit trails and role assignments are logged for SOC 2.
11. Pilot with 2-3 volunteer teams and iterate (depends on: 10)
Run a controlled pilot before full rollout to validate and refine the process.
- Select 2-3 volunteer teams with representative services and on-call patterns.
- Run the new severity, roles, on-call, alerting, and communication process for 2 weeks.
- Track metrics and gather feedback from on-call engineers, incident commanders, and communications leads.
- Iterate severity thresholds, alert rules, templates, and runbooks based on findings.
- Exit criteria: no SEV1 without a declared incident commander, alert noise below target, and positive on-call survey results.
12. Train and certify all 28 teams (depends on: 11)
Prepare all 28 teams so the process is not a paper exercise.
- Create role-specific training: all engineers, incident commanders, communications leads, scribes, subject-matter responders, and managers.
- Deliver interactive workshops and online modules.
- Run at least two game-day drills per quarter simulating SEV1 and SEV2 incidents, testing escalation, communication, and postmortem.
- Certify incident commanders and communications leads; only certified staff can hold those roles in rotations.
- Include scenario-specific training for account managers and legal or regulatory handling.
13. Full rollout and enforce process (depends on: 12)
Move every team onto the new process and retire old habits.
- Migrate all 28 teams into consolidated on-call schedules and alert routing.
- Retire old alerting tools or methods that bypass the process.
- Activate status page templates and postmortem action tracking globally.
- Monitor the first 30 days with a daily operational review and quick fixes.
- Communicate rollout status to leadership and all engineers; provide office hours and support.
14. SOC 2 readiness and continuous improvement (depends on: 13)
Map the process to SOC 2 and keep improving after go-live.
- Map each process component, including severity, roles, detection, communications, postmortems, and metrics, to SOC 2 control requirements.
- Produce runbooks, evidence of training, incident logs, postmortems, and action item tracking.
- Conduct an internal readiness assessment at month 6 and remediate gaps before the audit at month 8.
- Maintain weekly and monthly metrics reviews to drive continuous improvement.
- Refresh on-call compensation and alert quality rules based on operational data.
Step-level differences computed by the tool:
Proposal 1 vs the previous-round proposal it resembles most (opus5_initial_1): 18 steps kept, added ['Executive mandate, single owner, funding and non-negotiables', 'Seven-day interim command bridge', '24x7 coverage model: central command corps, local expertise', 'Detection uplift on the money path', 'Consolidate to one pager, one incident record, one status page', 'Live incident execution doctrine', 'Action ownership, capacity reservation and enforcement', 'Runbooks, readiness bar and ledger blast-radius reduction', 'Exercise programme: tabletops, game days and unannounced drills', 'Publish Incident Management Policy v1', 'Pilot on the payments critical path', 'Alert noise burn-down campaign', 'SOC 2 evidence by design, internal testing and mock audit', 'Ninety-day inspect-and-adapt, then year-two sustainability'], removed ['Program charter, executive mandate and funding', '24x7 incident command coverage model', 'Team on-call structure, rotations and routing rules', 'Detection uplift: SLOs, synthetic journeys and ledger assurance', 'Tool consolidation and incident platform implementation', 'Action item ownership and tracking system', 'Runbooks, major-incident playbooks and the on-call readiness bar', 'Simulation programme: game days, drills and wheel of misfortune', 'Pilot with wave 0 teams', 'SOC 2 control mapping, evidence automation and internal dry-run audit', 'Continuous improvement, maturity roadmap and post-audit sustainability']
Proposal 2 vs the previous-round proposal it resembles most (gpt5.6-sol_initial_2): 7 steps kept, added ['Install an interim process in seven days', 'Build the baseline, ownership catalog, and risk map', 'Design compliance and evidence controls from day one', 'Define roles, authority, and handoffs', 'Create sustainable 24×7 coverage across the service estate', 'Set the service readiness and runbook standard', 'Establish one paging and incident system of record', 'Enforce alert quality and burn down noise safely', 'Detect payment and ledger failures before customers', 'Operationalize regulatory, partner, contract, and credit decisions', 'Enforce action ownership and effectiveness tracking', 'Measure outcomes, controls, business impact, and human load', 'Train participants and address pager resistance', 'Pilot on the payment critical path', 'Roll out by customer journey and risk', 'Exercise command, regional resilience, and ledger recovery', 'Test SOC 2 operating effectiveness before the auditor'], removed ['Install immediate minimum controls', 'Create the service and dependency catalog', 'Define roles and sustainable 24x7 staffing', 'Consolidate incident and paging tooling', 'Improve detection and enforce alert quality', 'Codify acknowledgement and escalation paths', 'Measure performance and review it routinely', 'Train and certify participants', 'Pilot and expand the on-call model', 'Exercise regional, ledger, and communication failures', 'Execute a time-boxed enterprise rollout', 'Build SOC 2 evidence as the process operates']
Proposal 3 vs the previous-round proposal it resembles most (opus5_initial_1): 18 steps kept, added ['Executive mandate, program funding, and governance', 'Evidence baseline from incidents, alerts, and coverage gaps', 'Paid on-call, fatigue safeguards, and HR compliance', 'On-call staffing model and rotation rules', 'Detection uplift across payments, ledger, and customer signals', 'Single incident platform and alert-tool consolidation', 'Customer status page, account-manager outreach, and SLA credit workflow', 'Simulation program and game days', 'Critical-path pilot and gate review'], removed ['Program charter, executive mandate and funding', 'Forensic baseline of the 31 incidents and the alert estate', '24x7 incident command coverage model', 'Team on-call structure, rotations and routing rules', 'On-call compensation, labour compliance and fairness policy', 'Detection uplift: SLOs, synthetic journeys and ledger assurance', 'Tool consolidation and incident platform implementation', 'Customer communications and status page policy', 'SLA credit and financial impact workflow', 'Simulation programme: game days, drills and wheel of misfortune', 'Pilot with wave 0 teams']
Proposal 4 vs the previous-round proposal it resembles most (opus5_initial_1): 11 steps kept, added ['Immediate 7-day operating floor', 'Roles, authority, and ledger dual-control', 'Three-layer 24x7 coverage model', 'Customer-journey detection and ledger assurance', 'Single incident and paging platform', 'Detection-to-command escalation path', 'Live execution and major-incident playbooks', 'Internal, customer, and regulatory communications', 'Blameless postmortems and action tracking', 'Simulations and game days', 'Critical-path pilot', 'Metrics, reviews, and error budgets', 'SOC 2 evidence, mock audit, and sustainability'], removed ['Incident roles, decision authority and handover rules', '24x7 incident command coverage model', 'Team on-call structure, rotations and routing rules', 'Detection uplift: SLOs, synthetic journeys and ledger assurance', 'Tool consolidation and incident platform implementation', 'Detection-to-escalation path and the five-minute command rule', 'Internal communications protocol', 'Customer communications and status page policy', 'Regulatory, partner and legal notification playbook', 'Postmortem standard and blameless review forum', 'Action item ownership and tracking system', 'Runbooks, major-incident playbooks and the on-call readiness bar', 'Simulation programme: game days, drills and wheel of misfortune', 'Pilot with wave 0 teams', 'Metrics, dashboards and the review cadence', 'SOC 2 control mapping, evidence automation and internal dry-run audit', 'Program risk register and contingency planning', 'Continuous improvement, maturity roadmap and post-audit sustainability']
Proposal 5 vs the previous-round proposal it resembles most (opus5_initial_1): 15 steps kept, added ['Executive mandate, program governance, and interim incident command', 'Baseline data and alert estate analysis', 'Design paid on-call compensation and fatigue safeguards', 'Define acknowledgement and escalation paths', 'Standardize customer and status page communications', 'Standardize blameless postmortems', 'Track postmortem actions with owner and due date', 'Create runbooks and service readiness bar', 'Train, certify, and simulate incident response', 'Pilot on critical services and iterate', 'Wave rollout to all teams and decommission legacy paths'], removed ['Program charter, executive mandate and funding', 'Forensic baseline of the 31 incidents and the alert estate', 'On-call compensation, labour compliance and fairness policy', 'Detection-to-escalation path and the five-minute command rule', 'Customer communications and status page policy', 'Postmortem standard and blameless review forum', 'Action item ownership and tracking system', 'Runbooks, major-incident playbooks and the on-call readiness bar', 'Training, certification and the commander academy', 'Simulation programme: game days, drills and wheel of misfortune', 'Pilot with wave 0 teams', 'Wave rollout across all 28 teams with readiness gates', 'Change management, incentives and communications campaign', 'Program risk register and contingency planning']
Origin of the steps of the new proposals, matched by title by the tool (evidence for "taken"; ideas can also travel without a matching title):
Proposal 1: 18 of its 32 steps match its own previous version, 10 are new; step 14 resembles step 10 of proposal 2; steps 24, 25, 28 resemble steps 14, 22, 23 of proposal 4
Proposal 2: 7 of its 24 steps match its own previous version, 11 are new; steps 3, 6, 17 resemble steps 4, 6, 19 of proposal 1; steps 2, 12, 20 resemble steps 29, 20, 22 of proposal 4
Proposal 3: 1 of its 27 steps match its own previous version, 9 are new; steps 3, 4, 6, 9, 12, 13, 15, 16, 17, 18, 19, 22, 23, 24, 25, 26, 27 resemble steps 3, 4, 6, 10, 13, 14, 16, 18, 19, 20, 21, 25, 24, 27, 26, 28, 29 of proposal 1
Proposal 4: 4 of its 24 steps match its own previous version, 8 are new; steps 1, 3, 4, 5, 6, 9, 10, 16, 18, 22, 23 resemble steps 1, 2, 3, 4, 5, 9, 10, 17, 21, 25, 27 of proposal 1; step 12 resembles step 7 of proposal 2
Proposal 5: 4 of its 26 steps match its own previous version, 6 are new; steps 3, 4, 6, 10, 11, 16, 17, 25, 26 resemble steps 3, 4, 6, 10, 11, 16, 17, 26, 29 of proposal 1; steps 9, 12, 13, 14, 15 resemble steps 6, 7, 9, 11, 11 of proposal 2; step 7 resembles step 4 of proposal 3; step 18 resembles step 12 of proposal 4
PROPOSALS OF ROUND 1 (to assess):
--- PROPOSAL 1 (agent opus5_refine_1, anthropic/claude-opus-5) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (32):
1. Executive mandate, single owner, funding and non-negotiables
Convert the CEO email into a chartered program with one accountable owner and authority over all 28 teams. Incident response becomes a **company operating process**, not a per-team choice.
- Name an executive sponsor (CTO) and one accountable owner (Head of Reliability / Incident Management) with a small permanent office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group with Engineering, SRE/Platform, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours. It is not a 28-person committee.
- Fix the non-negotiables now: one severity scale, one paging tool, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.
- Set the clock deliberately earlier than the audit: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready week 28.
- Approve budget against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises.
2. Seven-day interim command bridge (depends on: 1)
Do not let design work leave the company unprotected for six weeks. Put a crude but real process in place within seven days and improve it later.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster from engineering managers and senior SREs, primary plus backup, 24x7. Pay it retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Tell Support to escalate credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items: complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This becomes both the design input and the "before" picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start / detect / declare / commander-assigned / mitigate / resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. This list drives the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome. Identify the top 50 rules producing most of the noise and every rule with no owner or runbook.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS and Sales in two weeks. Test what the objection actually is: unpaid work, lost sleep, unfamiliar code, missing runbooks, or fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey (fairness, trust in alerts, willingness, burnout) to re-measure at 60, 120 and 365 days.
5. Service catalog, ownership and money-path tiering (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that is the single source of truth for paging, impact and audit.
- One accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not by technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions and third parties, including sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Every orphan service gets an owner in 30 days or a decommission date approved by the sponsor. Tier 0 without an owner is an executive escalation.
6. Severity scale, declaration rules and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- **SEV3 (major, contained):** narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- **SEV4:** no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring; only the commander may downgrade, with the evidence recorded.
- Lifecycle: Detected → Declared → Triaged → **Mitigated** (customer impact ends) → Monitoring → **Resolved** (backlog processed and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Incident roles, authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, roles, cadence and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit spend. Retains command when a VP joins.
- **Communications Lead:** single voice for status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates for SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team; they mitigate, they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance and Vendor Management join on defined triggers.
- Rules: command claimed within 5 minutes, stated in channel ("I am IC"), distinct people for command, comms and technical lead at SEV1/SEV2, and every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation or privileged-access rules.
8. 24x7 coverage model: central command corps, local expertise (depends on: 5, 7)
Do not create 28 night rotations. Centralise coordination in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- **Layer A — Incident Command corps:** ~30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of ~18 from Support, CS and engineering management; a scribe pool used as the training entry point.
- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- **Layer C — everyone else:** business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing, with cost and timeline: a US-only paid night rotation now, a Lisbon or APAC follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
9. On-call compensation, labour compliance and fatigue safeguards (depends on: 8, 4)
Unpaid on-call in New York is both a retention problem and a legal exposure. Pay for it before asking anyone to sign up, and publish the numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 for business-hours rotations, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover rather than expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, tax treatment and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a ~15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
10. Alert quality standard, page budget and noise burn-down (depends on: 3, 5)
3,400 alerts at 85% noise is the reason detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human, and burn the backlog down deliberately rather than by mass silencing.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a **page budget** of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements; return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
- Target 3,400 → under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
11. Detection uplift on the money path (depends on: 10, 5)
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls; five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrating without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack or the identity provider is unavailable. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live incident execution doctrine (depends on: 7, 12, 13)
Give responders one short operating procedure for the first minutes through closure. Priority is limiting customer and financial harm, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; controlled backlog drain; **reconciliation completed before any payment or ledger incident is declared resolved**.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and Customer Success, and reopening if impact recurs.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook and RSS; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack; the long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity or blame; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded **even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review; Legal owns outbound regulatory text, the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
18. SLA credit and financial impact workflow (depends on: 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one, and use the delta as the standing business case for on-call pay and reliability capacity.
19. Blameless postmortem standard and Incident Review Board (depends on: 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first?** and **why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, capacity reservation and enforcement (depends on: 19, 12)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
21. Runbooks, readiness bar and ledger blast-radius reduction (depends on: 5, 8)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the single largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel workstream on blast-radius reduction — tenant or function partitioning, read replicas, and isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (depends on: 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days + shadowing):** command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence for their own domain before independent primary duty; two shadow shifts minimum, never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (depends on: 22, 12, 21)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 11, 12, 21, 22, 9)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect (from first impact), time to declare, time to commander, acknowledgement, mitigate, resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; team scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Wave rollout to all 28 teams with readiness gates (depends on: 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director; failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
28. Alert noise burn-down campaign (depends on: 10, 12, 25)
Run the noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
29. Change management, fairness and pager culture (depends on: 4, 9, 25)
Run this from day one in parallel. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (depends on: 24, 26, 27)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and contingencies (depends on: 1)
Name the ways this programme fails and pre-commit the response. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 26, 27, 30)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 2 (agent gpt5.6-sol_refine_2, openai/gpt-5.6-sol) ---
Estimated complexity: high
Success metrics: - By day 7, every suspected SEV0–SEV2 uses one incident record, one coordination channel, and one named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named commander for more than 10 minutes; at least 95% are assigned within 5 minutes.
- By day 30, 100% of Tier 0 services have a named owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, every Tier 0 and Tier 1 customer journey has compensated 24×7 command and subject-matter coverage.
- By day 120, all 180 production services have an owner, response tier, tested escalation path, and appropriate coverage model.
- No mandatory night rotation starts before its compensation, training, access, and staffing controls are active.
- Every critical primary rotation has at least six qualified responders or a documented, expiring executive exception by day 120.
- No responder is routinely scheduled for primary duty more often than one week in six by day 120.
- At least 95% of critical pages are acknowledged within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by day 120 and under 60 minutes by month 8.
- At least 95% of SEV0 and SEV1 customer notices are issued within 15 minutes of declaration by month 3.
- At least 95% of customer-visible SEV2 notices are issued within 30 minutes by month 3.
- At least 95% of incidents meet their required update cadence by month 3.
- Monthly paging volume falls from 3,400 to no more than 1,500 by day 90 and no more than 700 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6 without loss of critical detection coverage.
- All new paging alerts satisfy the owner, impact, action, dashboard, runbook, deduplication, and escalation standard by day 60.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 currently open historical actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date by month 6, with effectiveness evidence.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Two end-to-end cross-company exercises, including regional failure and ledger recovery, are completed before the audit.
- Monthly availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- SLA credits decline by at least 50% on an annualized trailing basis within 12 months.
- Quarterly on-call surveys show improving fairness and sustainability, with at least 75% favorable responses by month 6.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (24):
1. Create the mandate, ownership, and funding
Launch incident management as a company operating program within 48 hours. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable program owner.
- Form a small steering group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Fund two to three implementation staff, paging and incident tooling, observability work, training, exercises, and on-call compensation.
- Reserve 10%–15% of engineering capacity for alert remediation, runbooks, and incident actions.
- Authorize incident commanders to freeze deployments, order rollback, disable features, shift traffic, and invoke continuity plans.
- Preserve financial controls. Incident commanders may coordinate ledger recovery but may not bypass dual approval, privileged-access controls, or reconciliation.
- Set the delivery target: critical controls operational within 60 days, enterprise rollout within 120 days, and a mock audit in month 7.
2. Install an interim process in seven days (depends on: 1)
Do not wait for new tools or the final policy. Put a minimum viable incident process into operation immediately and start collecting evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communications schedule.
- Establish one monitored declaration path through chat, telephone, and the existing paging tools.
- Create a standard incident channel, bridge, incident document, and naming convention.
- Staff temporary primary and backup incident commanders 24×7 from the existing on-call teams and engineering leadership.
- Compensate interim duties retroactively under the final compensation policy.
- Require a named incident commander within 10 minutes for every suspected major incident.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Hold a daily 15-minute control review until the permanent process is live.
3. Build the baseline, ownership catalog, and risk map (depends on: 1)
Establish the facts behind the current failures and assign every production component an owner. Use the resulting catalog as the source for routing, escalation, and audit evidence.
- Reconstruct all 31 customer-impacting incidents, including first impact, detection source, declaration, command assignment, mitigation, resolution, customer communications, and credits.
- Analyze the two incidents with no clear leader and every case detected first by customers.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, databases, queues, regional dependencies, payment processors, and banking partners.
- Assign one accountable team, engineering manager, product owner, primary escalation, secondary escalation, dashboard, and runbook to each service.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 based on financial integrity, customer impact, contractual exposure, and dependency centrality.
- Map critical customer journeys to their application, PostgreSQL, Kubernetes, regional, and third-party dependencies.
- Inventory the six alert sources and all 3,400 monthly alerts by owner, volume, actionability, and duplication.
- Interview representatives from all 28 teams and baseline on-call sentiment, fatigue, and objections.
- Give orphan services an owner or decommissioning decision within 30 days.
4. Design compliance and evidence controls from day one (depends on: 1)
Map the operating process to audit, legal, contractual, and record-retention requirements before finalizing it. Confirm the expected SOC 2 observation period with the auditor immediately.
- Map controls to applicable SOC 2 criteria for monitoring, incident identification, response, recovery, communications, corrective action, access, and availability.
- Define evidence required for declarations, pages, acknowledgements, role assignments, decisions, status updates, postmortems, actions, training, drills, and exceptions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve the Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, and Postmortem Standard.
- Record control exceptions with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing evidence before the audit.
5. Adopt one severity and lifecycle standard (depends on: 2, 3, 4)
Use four impact-based severity levels across operational, security, data, and third-party incidents. Start at the highest credible severity when facts are uncertain, then downgrade with recorded evidence.
- **SEV0 — financial or security crisis:** suspected ledger corruption, unauthorized or duplicated funds movement, material data compromise, both-region loss, or a decision to suspend payment processing. Page all roles immediately; engage executives, Security, Legal, Compliance, and Risk; assess regulatory duties within one hour; use distinct role holders; require a postmortem.
- **SEV1 — critical availability event:** a core payment journey is unavailable, payment failures exceed an initial 10% guardrail for five minutes, regional loss has impaired failover, a settlement deadline is at imminent risk, or rapid error-budget burn makes a material SLA breach likely. Assign all roles; notify internal stakeholders within 10 minutes; publish customer status within 15 minutes; update every 30 minutes; require a postmortem.
- **SEV2 — major bounded event:** approximately 1%–10% of payment attempts fail, a material customer subset or critical customer is down, degradation has a workaround, or contractual impact is likely. Assign an incident commander and responders; add communications and scribe roles for customer impact; publish status within 30 minutes; update every 60 minutes; require a postmortem for customer-visible events.
- **SEV3 — limited event:** localized impact, a safe workaround, and no financial-integrity, security, regulatory, or material contractual risk. The owning team leads; page only if immediate action is necessary; use a ticket otherwise.
- Treat the percentage thresholds as declaration guardrails, not reasons to under-classify integrity, settlement, security, or strategic-customer risk.
- Permit any employee to declare an incident. Only the incident commander may lower severity, with the rationale logged.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of active customer harm. Declare resolution only after stability, backlog recovery, and required ledger reconciliation.
6. Define roles, authority, and handoffs (depends on: 5)
Separate command, communication, recordkeeping, and technical repair. One named person must hold command at every moment of a major incident.
- **Incident Commander:** owns severity, priorities, role assignment, escalation, decision cadence, mitigation coordination, and closure. The commander does not act as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language. Legal or Compliance retains ownership of regulatory submissions.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, and status changes.
- **Subject-Matter Responders:** diagnose and mitigate services for which they have accepted ownership, access, training, and runbooks.
- **Executive Duty Officer:** removes organizational obstacles and approves exceptional business decisions without displacing the incident commander.
- Require separate commander, communications, scribe, and primary technical lead for SEV0 and SEV1.
- For customer-visible SEV2, keep the commander separate from the primary technical responder; communications and scribe may be combined if workload permits.
- Announce every role assignment and transfer in the incident channel. Require verbal and written handoff with current impact, decisions, risks, and next actions.
- Keep executives and account managers out of the technical command path; questions flow through the Executive Duty Officer or Communications Lead.
7. Create sustainable 24×7 coverage across the service estate (depends on: 3, 6)
Use central command coverage and service-domain responder coverage rather than creating 28 fragile night rotations. Engineers remain responsible only for services they own or have formally trained to support.
- Build a company incident-command pool of approximately 18–24 certified senior engineers and managers, with a primary and backup scheduled at all times.
- Build 12–16-person communications and scribe pools using Support, Customer Operations, Engineering Operations, and qualified engineering managers.
- Schedule an Engineering Director or equivalent as the 24×7 Executive Duty Officer.
- Group related service owners into approximately 8–12 coherent responder domains only where members have training, access, and explicit acceptance.
- Require Tier 0 and Tier 1 domains to provide 24×7 primary and secondary responders, normally with at least six trained people in each sustainable rotation.
- Give Tier 2 services business-hours coverage plus a maintained manager escalation path. Treat Tier 3 conditions as tickets unless their impact changes.
- Maintain distinct but coordinated coverage for the ledger application and the PostgreSQL platform.
- Route unknown-owner events to the duty commander and platform triage temporarily. Treat every such event as an ownership control defect.
- If a small team cannot staff a fair rotation, merge coverage only after training or provide headcount, service reassignment, or decommissioning.
8. Approve compensation and fatigue protections (depends on: 7)
End unpaid on-call before expanding mandatory coverage. Publish the policy through HR, Legal, Finance, and Payroll within 14 days.
- Pay fixed weekly stipends for primary and secondary service rotations.
- Pay separate stipends for duty commander, communications, and scribe assignments.
- Provide additional call-out compensation or equivalent paid recovery time for material after-hours work.
- Apply overtime and reporting rules correctly for non-exempt employees under federal and New York requirements.
- Pay higher rates for company holidays and provide a protected recovery day after qualifying overnight work, SEV0 events, or prolonged SEV1 response.
- Target no more than one primary week in six and prohibit simultaneous primary assignments.
- Avoid consecutive primary weeks and make all swaps visible in the paging system.
- Reduce sprint commitments for people carrying primary duty rather than expecting normal delivery capacity.
- Provide a documented accommodation path for health, disability, or caregiving constraints without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
9. Set the service readiness and runbook standard (depends on: 3, 5, 7)
A team cannot respond effectively at night without ownership, access, telemetry, and rehearsed recovery procedures. Apply a formal readiness gate to every Tier 0 and Tier 1 service.
- Require a current architecture diagram, dependency map, dashboards, SLOs, runbooks, rollback method, feature-control method, contacts, and tested escalation path.
- Record RTO, RPO, data-integrity requirements, regional mode, and customer-facing capabilities in the service catalog.
- Link every paging alert to the exact runbook step expected from the responder.
- Prevent new paging alerts for services that fail readiness review. Preserve existing critical detection through a documented exception until a safe replacement exists.
- Write priority playbooks for PostgreSQL failure, ledger-integrity investigation, regional failover, Kubernetes control-plane degradation, payment-processor failure, queue backlog, credential compromise, and payment suspension.
- For the ledger, document read-only or stop-processing modes, failover controls, replay and duplicate protections, backlog handling, and post-recovery reconciliation.
- Require runbook review after material incidents and at least twice per year through exercises.
10. Establish one paging and incident system of record (depends on: 4, 5, 6, 7)
Consolidate paging and incident coordination without requiring an unsafe big-bang replacement of every monitoring system. Monitoring sources may remain specialized, but all human pages must enter one controlled platform.
- Select one enterprise paging platform and one integrated incident record, with chat, telephone, SMS, conference, status-page, ticketing, and service-catalog integrations.
- Initially ingest events from all six current tools, then deduplicate, correlate, and route by service ownership.
- Automate incident channel and bridge creation, role paging, timeline capture, severity changes, communications reminders, and postmortem creation.
- Preserve immutable records of declarations, acknowledgements, role assignments, decisions, and messages.
- Use role-based access, multifactor authentication, break-glass controls, and periodic access reviews.
- Provide telephone and offline fallback procedures for loss of chat, identity, the paging vendor, or an AWS region.
- Test paging and fallback paths weekly.
- Retire a legacy paging route only after its signals have owners, quality review, successful end-to-end tests, and at least two weeks of verified operation in the new path.
11. Enforce alert quality and burn down noise safely (depends on: 3, 10)
Treat paging alerts as production products with owners and quality requirements. Do not reduce noise by silently disabling detection.
- Require every page to identify the service, owner, customer or SLO risk, urgency, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Define an actionable page as one requiring prompt human judgment or intervention that materially reduces customer, financial, security, or contractual risk.
- Route informational, capacity-planning, and non-urgent conditions to dashboards or ticket queues.
- Prefer symptom and error-budget burn alerts over raw CPU, memory, pod, or log-volume thresholds.
- Run new alerts in shadow mode for at least seven days unless an emergency exception is approved.
- Review alerts with less than 50% actionability, more than three firings in seven days, or repeated no-action acknowledgements within two business days.
- Require compensating detection and an owner before suppressing or removing an alert.
- Set a responder load target of no more than two after-hours pages per person per week. A breach creates a mandatory alert-remediation plan.
- Review alert actionability, duplication, missed detection, and page load monthly by domain.
- Prioritize the small number of rules producing most of the current 85% noise.
12. Detect payment and ledger failures before customers (depends on: 3, 11)
Shift detection from infrastructure symptoms to customer journeys and financial outcomes. Set internal objectives stricter than the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, webhooks, and reporting freshness.
- Run external synthetic transactions from outside the production boundary and through both regions at least every minute for critical paths.
- Monitor payment failure rates, latency, queue age, delayed value, settlement-window risk, regional asymmetry, and third-party response quality.
- Add continuous ledger controls for reconciliation breaks, unexpected balances, duplicate identifiers, replication lag, backup health, and failover readiness.
- Add tenant or cohort anomaly detection for high-value customers and common payment methods.
- Convert high-priority Support, account-manager, bank, and processor reports into incident candidates within five minutes.
- Review every customer-first incident as a missed-detection defect and create a corrective action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Codify escalation and live incident execution (depends on: 6, 10)
Create one time-bound path from signal to ownership and mitigation. Delivery of a notification does not count as acknowledgement.
- Page the owning Tier 0 or Tier 1 primary immediately; page the secondary after five unacknowledged minutes; page the domain manager at 10 minutes; escalate to the engineering director at 15 minutes.
- For SEV0–SEV2, page the duty commander immediately. Page the backup after five minutes and require the Engineering Director to assume command if no certified commander owns the event by 10 minutes.
- Automatically involve command for integrity concerns, security concerns, regional events, cross-team impact, critical customer journeys, or unresolved ownership.
- If impact remains materially unknown after 15 minutes, increase response posture rather than waiting for certainty.
- Open one incident channel, bridge, and system record. State severity, known impact, assigned roles, current objective, and next update time.
- Freeze unrelated changes during SEV0 and SEV1 unless the commander records an exception.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require controlled backlog processing and reconciliation before resolving payment or ledger incidents.
- Require a formal command handoff for incidents extending beyond four hours or when fatigue impairs a role holder.
14. Standardize internal and customer communications (depends on: 5, 6, 10)
Communicate known impact early without waiting for root cause. The Communications Lead uses approved facts and always states the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish customer status within 15 minutes when customer-visible and legally safe, then update at least every 30 minutes.
- For SEV1, brief internal stakeholders and publish status within 15 minutes, then update every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send notice within 30 minutes, then update every 60 minutes.
- For SEV3, communicate directly to affected customers only when impact or contract terms require it.
- Give account managers an approved statement and affected-customer list within 30 minutes for SEV0 or SEV1 and within 60 minutes for SEV2.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. Describe affected capabilities, symptoms, workarounds, and the next update.
- Do not speculate about root cause, blame, data integrity, security scope, or recovery time.
- Post a Monitoring update within 30 minutes of mitigation. Post Resolved only after stability and required reconciliation.
- Provide a customer-facing incident summary within five business days for qualifying incidents.
- Record and approve any delay or restriction of public details during an active security threat.
15. Operationalize regulatory, partner, contract, and credit decisions (depends on: 4, 5, 14)
Give Legal and Compliance a timed decision process while keeping technical command with the incident commander. Record both reportable and non-reportable determinations.
- Build a jurisdiction and obligation matrix covering applicable NYDFS requirements, state breach laws, GLBA or FTC requirements, PCI obligations, money-transmitter rules, sponsor banks, payment networks, cyber insurance, and customer contracts.
- Validate applicability and deadlines with counsel rather than assuming every operational incident is reportable.
- Begin reportability assessment immediately for every SEV0, security event, suspected ledger-integrity event, and relevant SEV1.
- Target a documented initial legal assessment within one hour for SEV0 and within two hours for other potentially reportable events.
- Record the decision, evidence, approver, legal deadline, submission owner, and confirmation of delivery.
- Maintain tested 24×7 contacts for regulators, banks, networks, insurers, outside counsel, and critical vendors.
- Encode customer-specific notification clocks and channels in the customer record used by the Communications Lead.
- Have Finance calculate affected minutes, likely credits, and contractual exposure from the incident record within five business days.
- Review whether proactive credits or claims-based handling applies by customer segment and contract.
16. Make postmortems mandatory, consistent, and blameless (depends on: 5, 6, 10)
Use one review standard to learn from incidents and test whether controls worked. Keep learning reviews separate from performance or misconduct processes.
- Require a postmortem for every SEV0 and SEV1.
- Require one for customer-visible SEV2, customer-first detection, incidents lasting more than two hours, contractual breaches, repeat failures, control gaps, and ledger-integrity near misses.
- Produce the factual draft within three business days, conduct the review within five, and publish the approved version within 10.
- Use one template covering summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, lessons, and actions.
- Analyze why detection was not earlier and why mitigation took as long as it did.
- Examine technical, organizational, process, testing, dependency, and incentive factors rather than forcing a single root cause.
- Use a trained facilitator and written blameless rules. Describe decisions in the context and information available at the time.
- Publish broadly useful findings internally while maintaining restricted versions for security, privacy, personnel, or privileged material.
- Review postmortem quality and recurring factors in a weekly Incident Review Board.
17. Enforce action ownership and effectiveness tracking (depends on: 16)
Treat corrective actions as risk commitments rather than suggestions. Closing a ticket is insufficient without evidence that the control or system behavior improved.
- Give every action one named individual owner, manager, priority, due date, expected risk reduction, verification method, and linked work item.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with intermediate milestones.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Place recurrence-prevention actions for SEV0 and SEV1 ahead of discretionary feature work unless an executive accepts the residual risk.
- Escalate overdue high-risk actions to the manager after seven days, director after 14 days, and CTO after 30 days.
- Require written residual-risk acceptance, compensating controls, and a new review date when high-risk work is deferred.
- Verify completed actions through tests, telemetry, drills, or production evidence.
- Triage the 53 currently open historical actions within 30 days. Complete, re-plan, or formally accept the risk, prioritizing ledger, regional, security, and detection items.
18. Measure outcomes, controls, business impact, and human load (depends on: 10, 11, 14, 17)
Use a balanced scorecard so teams are not rewarded for suppressing alerts or avoiding incident declarations. Report medians and 90th percentiles, not averages alone.
- Measure time from first impact to internal detection, declaration, acknowledgement, command assignment, mitigation, and resolution.
- Track customer-first detection, missed escalations, status-page timeliness, update-cadence compliance, and role conflicts.
- Track incident count, recurrence, availability, error-budget burn, affected payment value, delayed transactions, reconciliation breaks, and SLA credits.
- Track page volume, actionability, duplicates, after-hours pages, missed acknowledgements, and load per responder.
- Track postmortem timeliness, action completion, action age, verified effectiveness, and repeated contributing factors.
- Track rotation size, duty frequency, recovery days, swaps, attrition signals, and quarterly responder sentiment.
- Hold a weekly Incident Review Board for incidents, actions, missed controls, and noisy alerts.
- Hold a monthly executive reliability review for trends, investment decisions, contractual exposure, and accepted risks.
- Hold a quarterly resilience and controls review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile dashboard data against customer cases and sampled incident records monthly to identify missing incidents or metric gaming.
19. Train participants and address pager resistance (depends on: 6, 7, 8, 13, 14, 16)
Introduce the operating model as a fair exchange, not an audit mandate. The message is that people are paid, paged only for accepted services, supported by trained command, and given capacity to fix defects.
- Train all employees to recognize impact, declare an incident, and find the incident channel and customer status page.
- Give all 260 engineers role-based training in severity, escalation, evidence preservation, handoff, and financial-integrity precautions.
- Certify incident commanders through instruction, simulation, two shadowed events or exercises, and observed performance.
- Train Communications Leads in status writing, account-manager briefing, contractual clocks, and legal escalation.
- Train scribes in timeline quality, decision capture, fact-versus-hypothesis labeling, and evidence handling.
- Require responders to demonstrate access, dashboards, rollback, runbooks, and escalation competence before primary duty.
- Require two shadow shifts before independent primary on-call.
- Appoint one adoption champion in each team and hold weekly office hours during rollout.
- Publish compensation, fatigue protections, service boundaries, and the accommodation process before assigning shifts.
- Use paid working time for training, exercises, shadowing, runbook work, and certification.
- Survey engineers at baseline, day 60, day 120, and quarterly thereafter.
20. Pilot on the payment critical path (depends on: 8, 9, 11, 12, 13, 14, 17, 19)
Run a four-to-six-week pilot across the highest-risk customer journey. Use real incidents and exercises to correct the process before wider rollout.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, authentication or API edge, settlement, and Support intake.
- Activate paid primary and secondary rotations, central command, communications, the incident record, status templates, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new system authoritative.
- Review every pilot page within one business day for routing, actionability, responder load, and missing context.
- Have the program team coach incidents without silently taking command from assigned role holders.
- Hold a weekly pilot retrospective and correct critical process or tool defects within 48 hours.
- Require 95% command assignment within five minutes, 95% timely communications, no unpaid pages, complete postmortems, and at least a 50% page-noise reduction before expansion.
21. Roll out by customer journey and risk (depends on: 18, 20)
Expand in fixed waves rather than waiting for every team to become perfect. Apply explicit readiness gates and time-limited exceptions.
- Days 0–7: operate the interim declaration, command, and communications process.
- By day 30: complete Tier 0 ownership, activate the first certified command roster, approve compensation, and begin the pilot.
- By day 60: provide compensated 24×7 coverage for every Tier 0 and Tier 1 customer journey and route all critical pages through the new platform.
- By day 90: complete critical detection upgrades, customer communications, regulatory playbooks, and the first major alert-noise reduction.
- By day 120: assign every production service a response tier, owner, tested escalation, and appropriate coverage model.
- Roll out remaining teams in two-to-three-week waves ordered by customer and dependency risk.
- Require each wave to pass ownership, alert, runbook, access, training, compensation, and tabletop gates.
- Disable legacy paging paths after verified cutover rather than leaving ambiguous parallel obligations.
- Publish a weekly adoption dashboard by team and escalate failed gates as business risks.
- Never start mandatory night coverage before compensation, staffing, training, and access are ready.
22. Exercise command, regional resilience, and ledger recovery (depends on: 9, 13, 14, 15, 19)
Validate the process under realistic conditions before depending on it during a crisis. Use the same action-tracking rules for exercises and real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops covering payment failure, customer-first detection, third-party failure, and ambiguous ownership.
- Exercise loss of one AWS region, Kubernetes degradation, PostgreSQL failure, suspected duplicate payments, queue backlog, and settlement risk.
- Exercise simultaneous operational and security events to test command and disclosure boundaries.
- Exercise loss of chat, status-page, identity, or paging providers using telephone and offline fallbacks.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors.
- Validate backups, restoration, RPO, RTO, failover prerequisites, financial controls, and post-recovery reconciliation.
- Avoid uncontrolled production-ledger experiments; use staging, replicas, simulations, or tightly governed production tests.
- Complete at least two cross-company exercises before the audit, including one overnight or unannounced paging test.
23. Test SOC 2 operating effectiveness before the auditor (depends on: 4, 18, 21, 22)
Demonstrate that the controls operate consistently, not merely that policies exist. Correct failures through tracked remediation rather than rewriting historical records.
- Preserve policy approvals, service ownership, schedules, compensation activation, access reviews, training, certifications, incidents, communications, postmortems, actions, and exercises.
- Sample evidence monthly from initial signal through verified action closure.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Use exceptions to document missed acknowledgements, late communications, incomplete records, and compensating controls.
- Conduct a formal mock audit in month 7 using the populations and interview roles expected by the external auditor.
- Trace at least one SEV0 or SEV1, one SEV2, a customer-reported event, and an exercise end to end.
- Remediate evidence and operating gaps before external fieldwork.
- Brief commanders, responders, Support, and Compliance for auditor interviews without scripting inaccurate answers.
24. Institutionalize continuous improvement (depends on: 18, 21, 23)
Keep incident management active after the audit by assigning permanent owners, budgets, and review cycles. Use incident trends to drive architectural investment.
- Assign permanent owners for the policy, service catalog, paging platform, status page, training program, metrics, and evidence repository.
- Review severity thresholds, communications timing, staffing, and compensation annually and after material process failures.
- Recertify commanders and communications leads annually through observed exercises.
- Review recurring failure families quarterly and require executive action when remediation repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fatigue, fairness, tool friction, and psychological safety.
- Report severe incidents, SLA exposure, overdue risks, resilience investments, and customer-first detection to the board or risk committee quarterly.
- Prioritize reduction of the shared-ledger concentration risk, stronger regional independence, deployment safety, graceful degradation, and automated mitigation.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
--- PROPOSAL 3 (agent qwen3.8-max_refine_3, alibaba/qwen3.8-max) ---
Estimated complexity: high
Success metrics: - Median time to detect customer-impacting incidents falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median mitigation time falls from 3h10 to under 90 minutes by day 120 and under 60 minutes by month 9.
- A named incident commander is assigned within 5 minutes in at least 95% of SEV1 and SEV2 incidents; zero incidents remain unowned for more than 10 minutes.
- First status-page update occurs within 15 minutes of SEV1 declaration and 30 minutes of SEV2 declaration in at least 95% of qualifying incidents.
- Monthly paging volume falls from 3,400 to under 500 actionable pages within six months, with alert actionability above 80%.
- Out-of-hours pages average no more than 2 per responder per week; sustained breaches trigger mandatory alert remediation.
- Six legacy alerting tools are consolidated into one paging and incident platform, with legacy paging paths disabled by week 16.
- 100% of Tier 0 and Tier 1 services have a named owning team, escalation path, dashboard, and runbook by day 60.
- All 28 teams are onboarded with readiness gates by week 22; every 24x7 critical-path rotation has at least six trained responders.
- 30 or more certified incident commanders and 20 or more certified communications leads provide continuous primary and secondary coverage.
- Paid on-call is approved by HR, legal, and finance and active in payroll before any new mandatory night rotation begins.
- 100% of required SEV1 and SEV2 postmortems are drafted within 3 business days and published within 10 business days in the standard template.
- Postmortem action closure rises from 17% to at least 90% of high-priority actions completed by their due date within two quarters.
- All 53 open historical actions are triaged within 30 days; high-risk unaccepted items are completed or formally risk-accepted within 90 days.
- Repeat incidents from a known unaddressed contributing factor decline by at least 50% within six months.
- Annualized SLA credits fall from $1.3M to under $400k within 12 months.
- Monthly availability meets or exceeds 99.95% by month 6, with exceptions reviewed at the executive reliability meeting.
- At least two cross-company exercises, including regional and ledger scenarios, are completed before the SOC 2 audit, with critical findings tracked.
- The month-6 internal dry-run audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete evidence.
- SOC 2 Type II incident-response controls pass with zero exceptions at the month-8 audit.
- On-call sentiment improves quarter over quarter; fewer than 10% of engineers report unwillingness to participate in their owned-service rotation by month 6.
- No increase in attrition among engineers on active rotations compared with baseline.
Steps (27):
1. Executive mandate, program funding, and governance
Convert the CEO email into a **written mandate within 48 hours**. Name one accountable program owner and a small decision group. Time-box design to five weeks so rollout starts well before the audit.
- Appoint the CTO as executive sponsor and a Head of Reliability or Incident Management as program owner with full-time authority.
- Create an 8–10 person design working group: SRE/platform lead, payments and ledger engineering managers, support lead, compliance, legal, HR, finance, and two rotating engineering managers.
- Approve budget lines: tooling consolidation, on-call compensation, training, coaching, and 2–3 dedicated program staff.
- State the non-negotiables: paid on-call, named service ownership, one severity scale, one paging platform, mandatory postmortems, and protected engineering capacity for reliability actions.
- Fix the master timeline: interim controls week 1, design weeks 1–5, pilot weeks 6–11, full rollout weeks 12–22, internal audit rehearsal month 6, audit-ready month 7.
- Publish a one-page charter to the company that says incident response is a company operating process, not an optional team practice.
2. Evidence baseline from incidents, alerts, and coverage gaps (depends on: 1)
Before changing anything, create an auditable **before picture** from the last 12 months. This baseline drives the severity design, staffing model, and executive reporting.
- Reconstruct all 31 customer-impacting incidents: detection source, first owner, severity, mitigation time, credits paid, and whether command was clear.
- Write a specific case review of the two incidents with unclear ownership for more than one hour.
- Inventory the six alert tools, alert volume by team, noise rate, rules with no owner, and rules with no runbook.
- Map the 16 teams without on-call and the 12 teams with unpaid on-call.
- Review the 53 open postmortem actions and triage the highest-risk items first.
- Freeze baseline metrics: 22-minute detection, 40% customer-first detection, 3h10 mitigation, $1.3M credits, 3,400 alerts, 85% noise, and 17% action closure.
3. Stakeholder listening, resistance mapping, and co-design group (depends on: 1)
Treat engineer pushback as **design input, not an attitude problem**. The objection to carrying a pager for other teams' code must shape ownership, routing, compensation, and command staffing.
- Interview leads from all 28 teams plus support, customer success, sales, compliance, and HR within two weeks.
- Separate the real causes of resistance: unpaid night work, unfamiliar systems, poor runbooks, unfair load, fear of blame, or unclear authority.
- Recruit 10–15 respected engineers and managers as co-designers so the model is built with the teams.
- Survey baseline on-call sentiment, trust in alerts, and psychological safety; repeat at 90, 180, and 365 days.
- Document the central promise: responders are paged for services they own, trained incident commanders coordinate, and all on-call work is compensated.
4. Service ownership catalog, criticality tiers, and dependency map (depends on: 2)
No alert can page the right team until every service has a **named owner**. Build the catalog as the routing foundation for on-call, severity impact mapping, status-page components, and audit evidence.
- Assign one accountable team to each of the 180 services, with an engineering manager, Slack channel, escalation policy, dashboard, and runbook link.
- Tier services: Tier 0 for money movement, ledger integrity, authentication, settlement, and shared PostgreSQL; Tier 1 for customer-facing degradable services; Tier 2 for internal or batch services; Tier 3 for non-critical services.
- Map customer journeys to services, databases, regions, third parties, and contractual SLA components.
- Mark orphan and shared services; require ownership, reassignment, or decommissioning within 30 days.
- Document cross-region dependencies, failover constraints, and services that defeat nominal regional redundancy.
- Treat missing ownership for Tier 0 or Tier 1 services as an executive escalation and a release-blocking risk.
5. Severity scale, declaration rights, and automatic triggers (depends on: 2, 4)
Adopt one severity scale so declaration is a lookup, not a debate. **Anyone may declare; only the incident commander may downgrade.** When uncertain, start higher.
- SEV1: money movement stopped or incorrect, ledger integrity in doubt, confirmed security or data event, both regions impaired, or broad SLA-credit exposure.
- SEV2: major degradation, settlement window at risk, one or more critical customers fully down, or likely contractual breach.
- SEV3: limited impact with a workaround, no financial-integrity or security risk.
- SEV4: no customer impact; handled by ticket during business hours.
- Define fixed triggers for each level: pages, command staffing, bridge call, status-page timing, account-manager outreach, executive notification, and postmortem obligation.
- Add automatic escalation: SEV3 open more than 2 hours becomes SEV2; any incident touching the shared ledger or unclear ownership becomes SEV2 unless the commander documents otherwise.
- Publish a decision tree and 10–12 worked examples from the actual 31 incidents.
6. Incident roles, authority, and handover rules (depends on: 5)
Solve the **nobody-was-in-charge failure** by making command explicit, trained, and transferable. Separate coordination from technical remediation so responders are not asked to debug unfamiliar code.
- Incident Commander: owns severity, priorities, escalation, mitigation strategy, role assignment, handoffs, and closure. Does not write code during the incident. May freeze deploys, invoke pre-approved failover, and pull any on-call responder.
- Communications Lead: owns status page, internal updates, account-manager briefs, and coordination with legal or compliance.
- Scribe: maintains the timestamped record of decisions, actions, role changes, and customer communications.
- Subject-matter responders: engineers from owning teams who diagnose and remediate within their accepted domain.
- Executive liaison: required for SEV1; shields the commander from executive questions and owns board or regulator escalation.
- Require the commander to claim the role within 5 minutes of SEV1 or SEV2 declaration, state it in the incident channel, and record every handover.
- Allow role combination only for SEV3 or SEV4; require separate people for command, communications, and primary technical work at SEV1.
7. Paid on-call, fatigue safeguards, and HR compliance (depends on: 3)
Unpaid on-call is a retention, fairness, and legal risk in New York. **Pay must be approved before any new mandatory rotation starts.** This is the fastest way to reduce pager resistance.
- Create a weekly stipend for primary and secondary on-call, differentiated by tier and night responsibility.
- Add per-page or per-incident pay for out-of-hours activation, plus guaranteed recovery time after night work or major incidents.
- Create a separate stipend for duty incident commanders and communications leads because command is a heavier burden.
- Verify FLSA, New York wage-hour, overtime, holiday, and payroll treatment with HR, legal, and finance.
- Cap rotation frequency: no more than one primary week in four or one week in six depending on staffing; no simultaneous primary assignments.
- Require at least six trained responders for any 24x7 rotation; fund hiring or service reassignment where teams are too small.
- Publish the compensation package and payroll start date before asking engineers to join rotations.
8. On-call staffing model and rotation rules (depends on: 4, 6, 7)
Do not force 28 identical night rotations. Use a **layered model**: central command coverage, critical-path team coverage, and business-hours coverage for lower-tier services.
- Company incident-command and communications rotation: 24x7 pool of 25–35 trained volunteers and designated senior staff, giving primary plus secondary coverage at all times.
- Critical-path teams: 24x7 primary and secondary on-call for Tier 0 and Tier 1 domains, including payments, ledger, authentication, API edge, Kubernetes platform, and shared PostgreSQL.
- Other teams: business-hours on-call with a documented night escalation list owned by the engineering manager.
- Group services into 8–12 coherent domains so rotations are sustainable; no rotation with fewer than six people is approved without an executive exception.
- Require handoff overlap, shadow shifts before first primary duty, and no on-call during approved PTO.
- Define cross-team pull rules: the commander may page another team's on-call with a 10-minute acknowledgement obligation; this is coordinated by command, not pushed onto responders.
- Publish schedules, swap rules, and load limits in the paging tool.
9. Alert quality standard, page budget, and noise controls (depends on: 2, 4)
With 3,400 alerts and 85% noise, detection fails because people stop trusting pages. Make alert quality a **condition of paging anyone**.
- Every page must have a named owning team, customer or SLO impact, severity, dashboard, runbook, expected action, and escalation policy.
- Page on customer-visible symptoms: payment success rate, latency, settlement deadlines, error-budget burn, ledger integrity, and replication health.
- Demote cause-based CPU, memory, or infrastructure-only alerts to dashboards or tickets unless they map to a customer journey.
- Set a page budget: maximum 2 out-of-hours pages per person per week; breach triggers a mandatory alert-tuning sprint.
- Auto-flag alerts that fire repeatedly without action or have high no-action acknowledgement rates.
- Require shadow mode for new alerts before they page humans, except for documented emergencies.
- Review alert quality monthly by team and publish a noise leaderboard.
- Target fewer than 500 actionable pages per month and above 80% actionability within six months.
10. Detection uplift across payments, ledger, and customer signals (depends on: 4, 9)
Customers detected 40% of incidents first. Detection must shift to **payment outcomes, ledger integrity, and inbound customer signals**, not host metrics alone.
- Define SLIs and SLOs for payment initiation, authorization, settlement, reconciliation, refunds, API availability, and reporting freshness.
- Run synthetic end-to-end payment tests from outside the platform in both AWS regions every 60 seconds.
- Add continuous ledger assurance: double-entry reconciliation, replication lag, failover readiness, disk pressure, and settlement-window countdown alerts.
- Monitor per-customer anomalies for top accounts so a single-tenant outage is detected before the account manager calls.
- Convert high-priority support tickets and account-manager reports into incident candidates within 5 minutes.
- Track detection source for every incident; make customer-first detection a reviewed defect with a corrective action.
- Add partner, banking, and card-network notification intake into the same declaration path.
11. Single incident platform and alert-tool consolidation (depends on: 5, 8, 9)
Collapse six tools into **one paging and incident-management system** with one queue, one timeline, and one audit trail. Do not create a parallel audit process.
- Select an integrated stack: paging and on-call scheduling, incident workflow, Slack or chat integration, bridge calling, status-page API, and ticketing integration.
- Implement one-command declaration that creates the incident record, channel, bridge, severity label, role prompts, and clock.
- Migrate alert sources by wave; retire a legacy paging path only after routing, ownership, and acknowledgement tests pass.
- Automate evidence capture: timestamps, acknowledgements, role assignments, severity changes, communications, and postmortem links.
- Integrate with the service catalog, on-call schedules, Jira or equivalent action tracker, customer-success tooling, and the status page.
- Verify the platform works during a single-region failure, including SMS, phone, and offline fallback paths.
- Set a hard date after which pages outside the chosen platform are not valid on-call obligations.
12. Escalation paths and five-minute command rule (depends on: 6, 8, 11)
Create one path from **signal to named commander in under five minutes**, any hour of the day. Make escalation automatic and time-bound.
- Accept declarations from automated alerts, engineers, support, account managers, partners, and customers through the same command.
- For SEV1 and SEV2, page the duty incident commander and owning-team primary immediately.
- Acknowledgement ladder: primary 5 minutes, secondary 10 minutes, manager 15 minutes, director or executive 20 minutes.
- If no commander claims the incident within 5 minutes, the platform assigns and announces one; the assignee may hand over but cannot leave the incident unowned.
- For SEV3, require team acknowledgement within 30 minutes; otherwise create a tracked work item.
- Give the commander pre-approved authority to invoke regional failover, ledger read-only mode, feature kill switches, and partner notifications without waiting for executive sign-off.
- Record every missed acknowledgement and escalation failure for weekly review.
13. Internal communications protocol (depends on: 6, 11)
Separate the **working incident channel from the audience channel** so responders can work and executives, support, and sales stay informed without interrupting the commander.
- Create one incident channel and one bridge per SEV1–SEV3 incident; use a read-only broadcast channel for executives and support.
- Set update cadence: every 15 minutes for SEV1, 30 minutes for SEV2, and at state changes for SEV3.
- Use a fixed update template: impact, customer-visible symptoms, current action, ETA or next update, commander, and communications lead.
- Brief support and customer success with a live affected-customer list and approved holding statements within 15 minutes of SEV1 or SEV2.
- Require executives to route questions through the executive liaison; the commander is not interrupted.
- Define fatigue rules and formal handover for incidents lasting more than 4 hours.
14. Customer status page, account-manager outreach, and SLA credit workflow (depends on: 5, 13)
Replace ad-hoc status updates with a **timed, owned, template-driven customer communication process**. Link incidents to SLA credits so finance and customer success are not surprised.
- Publish status-page updates within 15 minutes of SEV1 declaration and 30 minutes for SEV2; update every 30 or 60 minutes until resolution.
- Assign the communications lead as the single author; use pre-approved templates reviewed by legal and communications.
- Map status-page components to customer capabilities: payments, settlement, reporting, API, onboarding, and regional availability.
- For top accounts, require direct account-manager outreach within 30 minutes of SEV1 with an approved briefing pack.
- Send proactive email or webhook notifications to subscribed customers for SEV1 and SEV2.
- Publish a resolution notice within 30 minutes of mitigation and a customer-facing incident summary within 5 business days for SEV1.
- Automate SLA credit calculation from incident duration and affected capability; review with finance and legal within 5 business days.
- Track credits by incident and root-cause family to guide reliability investment.
15. Regulatory, legal, and partner notification playbook (depends on: 5, 14)
In payments, some incidents are reportable and the clock starts at detection. Make regulatory assessment a **mandatory step in the incident process**, not an afterthought.
- Map obligations: NYDFS cybersecurity-event rules, state breach laws, GLBA safeguards, PCI DSS if in scope, money-transmitter duties, card-network and sponsor-bank contracts, and cyber-insurer notice.
- Require a legal or compliance reportability assessment within 2 hours for every SEV1 and every security-related SEV2, even when the answer is not reportable.
- Maintain a 24x7 contact matrix for regulators, sponsor banks, card networks, outside counsel, insurers, and law enforcement.
- Pre-draft notification templates and preserve legal review.
- Encode customer-contract notification deadlines into account tiers and the communications workflow.
- Record every reportability decision, approver, deadline, and submission confirmation in the incident record.
16. Postmortem standard and blameless review (depends on: 5, 6)
Replace inconsistent postmortems with a **mandatory, blameless, single-format process**. The discipline comes from deadlines, facilitation, and action tracking.
- Require postmortems for all SEV1 and SEV2 incidents, customer-detected incidents, incidents over 2 hours, repeat failures, and ledger near-misses.
- Draft within 3 business days, peer review within 5, publish within 10 for SEV1 and SEV2.
- Use one template: summary, impact, timeline, detection analysis, response analysis, contributing factors, what worked, what failed, and action items.
- Make blamelessness explicit: focus on systems and decisions, not individual fault; never use postmortems in performance discipline.
- Hold a weekly incident review board to review postmortems, ratify severity, and challenge weak actions.
- Maintain a searchable postmortem library and quarterly recurring-cause analysis.
- Require a trained facilitator for major reviews; the incident commander attends but does not facilitate.
17. Action-item tracking, ownership, and delivery gates (depends on: 16, 11)
Only 11 of 64 actions were closed. Give postmortem actions the same status as **customer commitments**, with named owners and visible escalation.
- Create every action as a ticket with one named individual owner, priority, due date, and verification method.
- Use delivery classes: containment within 7 days, corrective work within 30 days, strategic work within 90 days.
- Reserve 15–20% of team sprint capacity for reliability and incident actions.
- Escalate overdue items: manager at 7 days, director at 14 days, CTO dashboard at 30 days.
- Block related feature releases when overdue P0 actions prevent recurrence of a severe incident.
- Require director approval and documented residual risk acceptance for overdue high-risk items.
- Verify effectiveness after completion; closing a ticket without evidence does not close the action.
- Target 90% of high-priority actions completed on time within two quarters.
18. Runbooks, critical-incident playbooks, and readiness bar (depends on: 4, 8)
Poor runbooks are a real cause of pager resistance and slow mitigation. Define a **minimum readiness bar** before a service is allowed to page anyone at night.
- Require for every Tier 0 and Tier 1 service: architecture summary, dependencies, dashboards, alert-to-runbook map, rollback procedure, feature flags, escalation contacts, and customer-impact statement.
- Write major playbooks for shared PostgreSQL ledger failure, regional failover, Kubernetes control-plane loss, payment-processor outage, settlement-window breach, duplicate-payment suspicion, and security compromise.
- Define ledger recovery rules: failover procedure, read-only degraded mode, reconciliation, RPO/RTO, and data-loss tolerance approved by executives.
- Test runbooks in drills at least twice a year; mark untested runbooks stale.
- Prevent paging alerts for services without readiness sign-off unless the engineering manager accepts the gap in writing.
- Keep runbooks linked from every alert and incident template.
19. Training, certification, and role readiness (depends on: 6, 12, 13, 16)
Command and communications are skills. Build a **tiered certification path** so rotations are staffed by people who have practiced, not by whoever is around.
- All employees: 1-hour module on declaring incidents, finding the incident channel, and reading the status page.
- All responders: half-day training on severity, acknowledgement, escalation, runbooks, and evidence hygiene.
- Incident commanders: 2-day course plus two shadowed incidents and one simulation before certification.
- Communications leads: training on status-page writing, customer language, account-manager briefs, and regulatory triggers.
- Scribes: training on timeline discipline and audit evidence.
- Certify for 12 months; renew through a simulation.
- Require shadow shifts before independent primary duty; no new hire holds primary within 90 days.
- Publish the certification register as an audit artifact.
20. Simulation program and game days (depends on: 19, 11, 18)
Rehearse the process before it meets a real SEV1. Simulations build commander confidence, expose runbook gaps, and produce audit evidence.
- Run monthly 60-minute tabletops using real incidents from the 31-incident baseline.
- Run quarterly game days covering regional failover, ledger replica promotion, dependency failure, partner outage, and security event.
- Run twice-yearly unannounced paging drills to measure night acknowledgement times.
- Include support, account managers, legal, compliance, and executives in at least one exercise per quarter.
- Produce tracked action items from every exercise using the same board as real incidents.
- Measure time to commander, time to first status update, and time to mitigation decision.
21. Critical-path pilot and gate review (depends on: 7, 10, 11, 14, 18, 19)
Prove the model on the highest-risk services with willing teams before full rollout. Run a **six-week pilot with daily feedback and public exit criteria**.
- Pilot with payments, ledger/database, platform/Kubernetes, API edge, authentication, and support intake.
- Activate severity scale, duty commanders, paid rotations, single paging platform, alert budget, status-page policy, and postmortem process.
- Hold a weekly pilot retrospective and fix process defects quickly.
- Validate night acknowledgement, cross-team pull response, severity clarity, and compensation payroll.
- Exit gate: commander assigned within 5 minutes in 95% of incidents, status page on time, page noise down at least 50%, postmortems on time, and positive on-call sentiment.
- Publish pilot results to the whole company as the main adoption argument.
22. Wave rollout across all 28 teams (depends on: 21)
Roll out by criticality and dependency, not by calendar alone. Use **readiness gates** so teams are not forced live without coverage.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1 teams. Wave 3: Tier 2 teams. Wave 4: Tier 3 and internal platforms.
- Gate per team: catalog entry complete, alerts migrated, runbooks ready, six-person rotation staffed, one commander candidate nominated, compensation in payroll, and one tabletop passed.
- Assign a program coach to each wave for three weeks.
- Freeze legacy paging paths for each team after successful onboarding.
- Publish a live adoption scoreboard by team.
- Complete all teams by week 22, leaving months of operating evidence before the audit.
23. Metrics, dashboards, and review cadence (depends on: 11, 16, 21)
Instrument the process itself. Leadership must see whether the program is working, and auditors must see **operating evidence, not retrospective paperwork**.
- Track response metrics: MTTD, time to declare, time to commander, acknowledgement time, mitigation time, resolution time, and customer-first detection rate.
- Track quality metrics: page volume, alert actionability, missed pages, postmortem timeliness, action closure rate, and status-page compliance.
- Track business metrics: availability against 99.95%, SLA credits, repeat incidents, and error-budget consumption.
- Track people metrics: on-call load, night pages per person, recovery time usage, sentiment, and attrition signals.
- Hold weekly incident review, monthly reliability review, quarterly executive review, and annual policy review.
- Publish dashboards internally and give every metric a target and owner.
24. Change management, incentives, and culture (depends on: 3, 7, 21)
The process will be judged on fairness. Communicate the deal repeatedly and make participation **recognized, compensated, and safe**.
- Core message: you are paid, paged only for what you own, supported by a trained commander, and given real capacity for actions.
- Run CTO all-hands, team roadshows, office hours, FAQ, and an internal incident-management hub.
- Add incident response and reliability work to promotion criteria and manager objectives.
- Recognize good postmortems, alert-noise reduction, and calm incident leadership.
- Provide a written path for engineers who cannot do nights; cover those shifts with paid volunteers or adjusted staffing.
- Prohibit retaliation for good-faith declaration or escalation.
- Publish sentiment survey results, including bad news, to maintain credibility.
25. SOC 2 evidence design and internal dry-run audit (depends on: 16, 17, 22, 23)
Design the process so audit evidence is a **by-product of normal operations**. Test it internally before the external auditor does.
- Map the process to SOC 2 criteria: incident identification, response, recovery, monitoring, communication, control activities, availability, and corrective action.
- Approve versioned policies: incident response policy, severity standard, on-call policy, communications policy, postmortem standard, and exception process.
- Retain incident records, paging logs, status-page history, postmortems, action tickets, training records, drill records, and access reviews for the audit period.
- Record regulatory reportability decisions even when no notification is required.
- Run an internal dry-run audit at month 6: sample at least 15 incidents and trace evidence end to end.
- Fix gaps with at least four weeks before the audit.
- Maintain an exception and remediation log instead of claiming perfection.
26. Program risk register and contingency planning (depends on: 1)
Name the likely failure modes now and pre-commit responses. Review the register monthly with the executive sponsor.
- Commander volunteer shortfall: roster command duty for senior engineers and managers until the pool reaches 25 trained people.
- Compensation delay: use immediate time-off-in-lieu plus phased stipend, but do not start mandatory night rotations without pay.
- Tool migration slip: preserve single paging and routing scope; defer advanced automation if needed.
- Alert pruning causing missed incidents: move noisy alerts to ticket first, observe 30 days, then delete.
- Burnout in experienced teams: monitor load weekly and cap pages per person.
- Major SEV1 during rollout: program owner shifts to incident support, wave schedule slips one wave, sponsor is informed same day.
- Shared ledger concentration: track blast-radius reduction and failover improvements as top strategic actions.
27. Continuous improvement, maturity roadmap, and post-audit sustainability (depends on: 23, 25)
Prevent the classic failure where the process decays after the audit. Build the second-year operating plan before the first year ends.
- Hold quarterly process retrospectives with commanders, responders, support, and communications leads.
- Re-baseline targets every six months; tighten goals once current targets are met.
- Move from lagging metrics to leading indicators: error-budget burn, near-miss rate, drill performance, and action ageing.
- Plan year-two improvements: follow-the-sun coverage, automated mitigation, ledger blast-radius reduction, error-budget release gates, and per-customer real-time impact reporting.
- Keep annual policy review, certification renewal, drills, and board reporting on a permanent calendar independent of the audit cycle.
- Report quarterly to the board or risk committee on availability, credits, severe incidents, overdue actions, and resilience investment.
--- PROPOSAL 4 (agent grok4.6_refine_4, xai/grok-4.6) ---
Estimated complexity: high
Success metrics: - Median time to detect falls from 22 minutes to under 10 minutes by month 3 and under 5 minutes by month 9.
- Customer-first detection falls from 40% to under 20% by month 3 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by month 4 and under 60 minutes by month 12.
- A named Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes after week 4.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases by month 3, with required update cadence met in 95% of cases.
- Monthly pages fall from 3,400 to under 500 within 6 months, with actionability above 75%; out-of-hours pages at or under 2 per person per week.
- All six legacy alerting tools route through one paging platform by week 16; legacy paging paths disabled per wave.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by due date within two quarters; all 53 currently open historical actions triaged within 30 days of S2.
- SLA credits fall from $1.3M to under $400k in the first 12 months; customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known unaddressed cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar; every Layer B rotation has 6+ certified responders or a time-limited executive exception.
- Paid on-call policy approved by HR, Legal, and Finance and in payroll before any mandatory night rotation starts.
- All 28 teams onboarded by week 24 with 24x7 command cover from 25+ certified ICs and 15+ certified Comms Leads.
- Internal dry-run at month 6 passes a 15-incident evidence walkthrough; month-7 mock audit finds no unowned high-risk control gap; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment improves quarter over quarter; no increase in attrition among engineers on rotation; pulse at 60 and 120 days shows at least 70% agree rotations are fair and limited to services they own.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked actions; two cross-company exercises completed before the audit.
Steps (24):
1. Program charter, mandate, and funding
Convert the CEO email into a named program with one owner, a budget, and a deadline earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, 28 teams will opt out.
- Appoint a Program Lead (Head of Reliability) with direct CTO and CEO sponsorship.
- Form a small steering group: CTO, VP Eng, Head of CS/Support, CISO, Legal, Finance, HR. Not a 28-team committee.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Timeline: operating floor in 7 days; design weeks 1–6; pilot weeks 7–12; rollout weeks 13–24; mock audit week 28; SOC 2 at month 8.
- Fund tooling ($150–250k/yr), on-call pay (~$600k–1M/yr), 2–3 program FTEs, and reserved engineering capacity. Anchor the ask against $1.3M in credits plus unmeasured incident cost.
- Freeze baselines now: 31 incidents, MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, $1.3M credits, 3,400 alerts/month at 85% noise, 11 of 64 actions closed.
2. Immediate 7-day operating floor (depends on: 1)
Do not wait for tooling, compensation, or the audit. Put a minimum process in place this week so the next outage has a named commander.
Start retaining artifacts on day one. This week becomes the first audit evidence.
- Publish a one-page interim severity guide and a single declaration path through Slack, phone, and pager.
- Staff interim primary and backup Incident Commander 24x7 from existing on-call veterans and engineering managers. Compensate this duty retroactively.
- Require a named IC within 10 minutes for every suspected major incident. If nobody claims it, the duty manager is IC.
- Use one channel, one bridge, one timeline doc, and one naming convention for every major incident.
- Direct Support to escalate credible customer reports immediately. They do not wait for engineering confirmation.
- Triage all 53 open historical actions. Close, re-plan, or formally risk-accept. Ledger integrity, payment duplication, regional resilience, security, and detection come first.
- Hold a daily 15-minute ops review until the permanent process is live.
3. Forensic baseline of incidents and alerts (depends on: 1)
Rebuild the facts before locking design. This is the before picture for the CEO and the auditor.
Every later design choice should trace to this evidence.
- Re-code each of the 31 incidents: trigger, service, detection source, timestamps, who led, credits paid, root-cause family.
- Quantify the 40% customer-first detections and name the missing signal in each case.
- Reconstruct the two nobody-in-charge incidents minute by minute. Use them as the burning-platform story.
- Audit the six alerting tools: volume per tool and team, top 50 noisy rules, rules with no owner or runbook.
- Freeze the baseline numbers. Do not let them drift during design.
4. Listening tour and the fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the main delivery risk. Treat it as a design constraint, not an attitude problem.
The answer that will actually stick is **the deal**: you are paid; you are paged only for services you own; a trained commander runs the room; your postmortem actions get real sprint capacity.
- Interview all 28 teams plus Support, CS, and Sales in two weeks.
- Separate the objections: unpaid work, nights, unfamiliar code, bad runbooks, fear of blame. Each needs a different fix.
- Collect informal practices from the 12 teams already on-call. They are the pilot candidates and the veteran pool.
- Recruit 10–15 credible engineers as a design working group so the process is co-authored.
- Baseline sentiment on alert trust, on-call willingness, and burnout. Re-measure at 6 and 12 months.
5. Service ownership catalog and criticality tiers (depends on: 3)
You cannot page the right person across 180 services until each one has a named owner. This is the foundation of fairness, routing, and audit evidence.
Build a machine-readable catalog as the single source of truth.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, the shared PostgreSQL cluster, queues, partner banks, and customer-facing endpoints.
- Assign one owning team, named engineering manager, Slack channel, escalation policy, and dependency list per service.
- Tier 0: money movement, ledger, auth, shared Postgres, regional control plane. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map Tier 0/1 services to customer capabilities: initiation, authorization, settlement, reporting, onboarding.
- Assign coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
6. Severity scale and declaration rules (depends on: 3, 5)
Replace in-the-moment debate with a lookup. Four payments-specific levels, with worked examples drawn from the 31 real incidents.
Anyone may declare. Only the Incident Commander may downgrade, with recorded rationale. When unsure, start high.
- **SEV1**: money movement stopped or incorrect; ledger integrity in doubt; material security or data exposure; both regions impaired; or more than 10% of customers impacted. Pages IC, comms, scribe, SMEs, and exec liaison. Bridge in 5 minutes. Status page in 15 minutes. Mandatory postmortem and regulator assessment.
- SEV2: severe degradation; settlement window at risk; a strategic customer fully down; SLA breach likely. IC and SMEs paged. Status page in 30 minutes. Mandatory postmortem.
- SEV3: partial impact with a workaround; no credit exposure. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, or a repeat.
- SEV4: internal or minor. Ticket only. No page.
- Auto-escalate: any SEV3 open more than 2 hours, or any incident touching the shared ledger, becomes SEV2. Unknown impact after 15 minutes is raised, not sat on.
- Nobody is punished for over-declaring. Publish that rule in writing and repeat it.
7. Roles, authority, and ledger dual-control (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, transferable, and logged.
The Incident Commander owns the incident, not the fix, and **never types** in a production terminal.
- Incident Commander: declares severity, pulls anyone, freezes deploys, invokes failover, authorises spend. One IC at a time. Assumed within 5 minutes and announced in channel.
- Communications Lead: single voice to customers, status page, account managers, and the exec summary.
- Scribe: timestamped timeline, decisions, and open questions. Feeds the postmortem and the audit trail. Required for SEV1 and SEV2.
- Subject-matter responders: diagnose and mitigate only services they own, with access and runbooks.
- Executive Liaison (SEV1): shields the IC from exec questions; owns regulator and board escalation.
- The IC may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access rules.
- Handover is verbal and written, with time and new owner recorded. Roles may combine below SEV2; never at SEV1. The IC stays in charge if a VP joins.
8. Three-layer 24x7 coverage model (depends on: 5, 7)
Do not put 28 teams on night rotation. That is what engineers are rejecting.
Staff command centrally. Page engineers only for code their team owns.
- Layer A, company command: Duty IC plus secondary, Duty Comms, and a scribe pool. Target 25–35 certified ICs and 15–20 comms leads. Roughly one week per person every 6–8 months. Five-minute ack SLA, then secondary, then on-call director.
- Layer B, Tier 0/1 domains: group 28 teams into 8–12 product and platform domains (ledger app, Postgres platform, payments orchestration, auth, API edge, Kubernetes/platform, settlement). Primary plus secondary 24x7. Minimum six trained people. Target one week in six, never worse than one in four.
- Layer C, Tier 2/3: business-hours on-call. After hours the IC pages the EM, who holds a written escalation list.
- No engineer joins another team's responder pool without training, access, runbooks, shadow shifts, and both teams' acceptance.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner.
- Evaluate follow-the-sun coverage as a 12-month option, not a year-1 dependency.
9. Paid on-call, NY labor compliance, and fatigue rules (depends on: 4, 8)
Unpaid on-call is a retention risk and a New York legal exposure. Pay must be live in payroll before any mandatory night rotation starts.
Treat interrupted nights and recovery time as compensable work.
- Weekly stipend by layer, higher for 24x7 Tier 0/1 and Duty IC, lower for business-hours, with holiday and weekend premiums.
- After-hours call-out pay or TOIL. A paid recovery day after prolonged overnight work, a SEV1, or a qualifying SEV2. Managers cover the next day.
- HR, Legal, and Finance publish dollar amounts, FLSA exempt/non-exempt treatment, NY wage-hour rules, tax treatment, and payroll timing within 14 days of charter.
- Load rules: no primary on two rotations; no consecutive primary weeks; no on-call the week after a SEV1 you commanded.
- A person may declare temporarily unfit after overnight work with no performance penalty.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- If a rotation cannot staff six people, merge domains or hire. Do not run two-person 24x7.
- Model annual cost against $1.3M in credits and get it as a CFO/board line item.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts a month at 85% noise is why detection takes 22 minutes. Make quality a condition of paging a human.
A page is a product with a quality bar, not a dump of host metrics.
- Every paging alert must have a named owner, customer or SLO impact, runbook, tested threshold, severity mapping, dashboard, expected action, and dedup key. Fail any of these and it becomes a ticket or is deleted.
- Page on customer symptoms: SLO burn, error budget, settlement-queue age. Cause-based CPU and memory alerts become dashboards or tickets.
- Set a **page budget** of at most two out-of-hours pages per person per week. A breach triggers a mandatory tuning sprint and blocks new paging alerts for that team.
- Auto-quarantine alerts that fire more than five times a month with no action, or that have more than 70% no-action acks. Never silently disable without compensating detection and a recorded decision.
- Run new alerts in shadow for seven days unless an emergency exception is approved.
- Monthly per-team kill, tune, or keep review. Target under 500 pages a month and actionability above 75% within six months.
11. Customer-journey detection and ledger assurance (depends on: 5, 10)
Stop customers being the monitoring system. Detect on money-path outcomes, not host metrics.
Every postmortem will ask why a customer saw it first.
- Define SLOs per Tier 0/1 capability: initiation success, auth latency, settlement timeliness, API availability, reporting freshness. Set an internal target stricter than 99.95%.
- Run synthetic full-lifecycle payments from outside the platform, in both regions, every 60 seconds. Include a small real-value canary where legally feasible.
- Add ledger assurance: continuous double-entry reconciliation, replication lag, failover-readiness, unexpected balances, duplicate identifiers, settlement-window countdown.
- Add per-customer anomaly detection for the top 100 accounts.
- Auto-create a triage incident within five minutes when support tickets or AM reports match impact keywords.
12. Single incident and paging platform (depends on: 6, 8, 10)
Collapse six alerting tools into one queue, one timeline, and one audit record.
The platform must still work if one AWS region is down. Verify SMS and phone fallback, plus an offline runbook.
- Select one paging product, one incident layer, and one hosted status page.
- One Slack command declares the incident, creates the channel and bridge, pages the Duty IC, sets severity, and starts the clock.
- Ingest the six existing tools first. Deduplicate and route. Retire a legacy path only after named owners and two weeks of verified operation.
- Auto-capture timestamps, acknowledgements, role assignments, severity changes, comms, mitigated time, and resolved time. Export for SOC 2.
- Integrate the service catalog, Jira, Salesforce or CS tooling for affected-customer lists, and Zoom or Slack Huddle.
- Test paging, escalation, status publication, and conference access every week.
- After a team's wave, old paging paths are disabled, not left as a fallback.
13. Detection-to-command escalation path (depends on: 7, 8, 12)
Write the single path from something looks wrong to someone is in charge. Target: a commander in under five minutes, any hour.
If nobody claims IC in five minutes, the platform assigns the Duty IC. The assignee can hand over, not decline.
- Converge every entry point on the same declare command: alert, engineer, support, account manager, partner bank, SEV hotline.
- SEV1: page the owning primary immediately; secondary at 5 minutes unacked; domain manager and Duty IC at 10; exec liaison at 15.
- SEV2: primary ack in 10 minutes; IC assigned in 15.
- Human acknowledgement is required. Delivery to a device does not count.
- The IC can page any team's on-call, with a 10-minute ack obligation. This reciprocity makes single-team ownership viable.
- Unowned alerts go to Layer A command, then the missing owner record is a control defect.
- Pre-authorise regional failover, ledger read-only mode, and partner-bank notice so the IC does not wait for an executive. Dual-control still applies to ledger writes.
14. Live execution and major-incident playbooks (depends on: 7, 13)
Limit customer and financial harm before proving root cause. One procedure from the first minute to handback.
A service cannot page at night until it meets the readiness bar.
- Open channel, bridge, record, and timeline immediately for SEV1 and SEV2. The IC states severity, known impact, hypothesis, objective, roles, and next update time.
- Freeze unrelated production changes during SEV1. Record exceptions the IC approves.
- Prefer reversible mitigation: rollback, feature flag, traffic isolation, rate limit, partner reroute.
- Write playbooks first for Postgres ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank outage, settlement-window breach, suspected security compromise, and suspected duplicate payments.
- Guard against split-brain, replay, and duplication during regional or database recovery. Reconcile and process backlog before calling the incident resolved.
- Mitigation means customer impact ended. Resolution means stable, backlog processed, and ledger reconciled.
- Formal IC handover after 4 hours. Reopen if impact recurs during the stability window.
- Readiness bar per Tier 0/1 service: diagram, dependencies, dashboard, runbook, rollback, kill switch, escalation contacts, RPO/RTO. Untested runbooks are marked stale.
15. Internal, customer, and regulatory communications (depends on: 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved process. The Communications Lead is the single author.
State impact and the next update time. Never speculate on cause.
- Internal: one working channel, one bridge, one read-only broadcast for execs, Support, and Sales. SEV1 update every 30 minutes even if unchanged; SEV2 every 60 minutes.
- Executives ask questions only of the Executive Liaison. Publish this as a signed exec behaviour rule.
- Status page: SEV1 within 15 minutes, SEV2 within 30 minutes; then 30/60-minute updates; resolve notice within 30 minutes of mitigation. Templates pre-approved with Legal.
- Top 100 accounts: named AM contact within 30 minutes of SEV1, with a briefing pack from Comms. Long tail gets the status page plus email or webhook.
- Customer-facing summary within 5 business days for SEV1.
- Mandatory regulatory checkpoint on every SEV1 and every security SEV2, within 2 hours, recorded even when not reportable. Cover NYDFS Part 500 72-hour clock, state breach laws, GLBA/FTC, PCI if in scope, sponsor-bank and card-network windows, FinCEN/OFAC if relevant.
- Legal owns outbound regulatory letters. The IC owns facts. Security or Legal may limit public detail during an active threat, with the reason recorded.
- Encode bespoke customer-contract notice SLAs into account tiering.
16. SLA credit and financial-impact workflow (depends on: 6, 15)
Link incidents to money so severity, credits, and investment stay consistent. Finance should not learn about outages from invoices.
Make credit calculation an output of the incident record, not a negotiation.
- Agree availability measurement per contract and component with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and capability telemetry. Produce a proposed credit schedule within 5 business days of resolution.
- Posture: proactive credits for top-tier accounts; claims-based for the rest. Document the approval chain.
- Track credits per incident and root-cause family. Quarterly report which reliability investments would have prevented which credits.
- Target: cut credits from $1.3M to under $400k in 12 months. Use that delta as the ongoing business case.
17. Blameless postmortems and action tracking (depends on: 6, 12)
Eleven of 64 actions closed is the clearest process failure. Make learning mandatory and actions as binding as customer commitments.
Closing a ticket without evidence of effectiveness does not close the action.
- Mandatory for every SEV1 and SEV2, every customer-first detection, every incident over 2 hours, every repeat of a known cause, and every ledger near-miss.
- Draft within 3 business days, review within 5, published internally within 10. The IC owns the draft. The owning EM is accountable.
- One template: timeline, customer and financial impact, why detection was late, why mitigation took that long, contributing factors, what went well, actions.
- Blameless in writing. No individual named as a cause. HR and management commit that postmortems are never used in performance reviews.
- Every action gets a named person, priority, due date, Jira ticket, and verification method. P0 (prevents SEV1 recurrence) due in 30 days and committed into the next sprint before roadmap work. P1 in 60 days. P2 in 90 days.
- Teams reserve 15–20% of sprint capacity for reliability and incident actions.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Weekly Incident Review Board with engineering directors. Target 90% of P0/P1 actions closed on time within two quarters.
18. Training, certification, and commander academy (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Nobody holds independent duty untrained. Training happens in paid working time.
The certification register is an audit artefact.
- All employees, 1 hour: recognise impact, declare, find the channel and status page.
- Responder, half day: severity, escalation, runbooks, comms hygiene. Mandatory before joining a rotation.
- Scribe, 2 hours: timeline discipline. The entry point.
- Incident Commander, two days plus shadowing: command presence, decisions under uncertainty, severity, handover, exec management. Two shadowed incidents and one simulated SEV1 before certification.
- Communications Lead, one day: status writing, customer tiering, legal boundaries, regulator triggers.
- Certification valid 12 months, renewed via simulation.
- New joiners shadow two shifts and do not hold primary in their first 90 days.
- Appoint a champion in each of the 28 teams.
19. Simulations and game days (depends on: 14, 15, 18)
Rehearse before the next real SEV1. Use a standing calendar, not a one-off exercise.
Do not inject uncontrolled changes into the production ledger.
- Monthly 60-minute tabletop per engineering group, using a real incident from the 31.
- Quarterly full-scale game day: regional failover, ledger replica promotion, dependency failure. Whole role structure, timed.
- Twice-yearly unannounced paging drill, including nights, to measure real acknowledgement times.
- One security incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Also exercise status-page failure and loss of the primary chat or pager.
- Use replicas, staging, or tightly governed tests for ledger scenarios.
- Every exercise produces actions in the same tracker as real incidents.
- Complete two cross-company exercises before the SOC 2 audit.
20. Critical-path pilot (depends on: 9, 11, 12, 18, 19)
Prove the process on the highest-risk surface with willing teams before asking 28 teams to adopt it.
Publish a one-page result company-wide. That result is the adoption argument.
- Six-week Wave 0: ledger app, Postgres platform, payments orchestration, Kubernetes/platform, API gateway, Support intake, plus two of the 12 teams already on-call.
- Activate the full stack: severity scale, Duty IC, one pager, page budget, status page, mandatory postmortems, paid on-call.
- Parallel-run old paths for one week, then cut over. The program lead coaches every SEV2+ and does not secretly take command.
- Weekly retro. Expect 20–40 process defects and fix them in the standard before rollout.
- Exit gate: MTTD under 10 minutes on pilot services; IC assigned within 5 minutes in 95% of incidents; page volume down 50%; all postmortems on time; no unpaid pages; sentiment not worse.
21. Metrics, reviews, and error budgets (depends on: 12, 17, 20)
Instrument the process itself. Do not reward hiding incidents or suppressing pages.
Every metric has a target and a named owner. Dashboards are public inside the company.
- Response: MTTD, time to declare, time to IC, MTTA, MTTM, MTTR, incidents by severity, percent customer-first. Report median and p90, by tier, journey, and region.
- Quality: pages per person per week, actionability, budget breaches, postmortem on-time rate, action closure and age, missed acknowledgements.
- Business: credits, 99.95% per capability, error-budget burn, failed-payment volume, reconciliation breaks, repeat-incident rate.
- People: rotation size, frequency, after-hours pages, recovery days, sentiment, attrition among on-call staff.
- Cadence: weekly Incident Review Board; monthly Reliability Review; quarterly exec and board review; annual policy review.
- Error budgets on Tier 0/1 SLOs. Burn too fast and the team pauses features to pay down reliability.
- Reconcile dashboards monthly against a sample of incident records and customer cases so missing incidents cannot hide.
22. Wave rollout with readiness gates (depends on: 20, 21)
Roll out in four waves by criticality, every three weeks. Gates keep the standard credible. A missed gate is rescheduled, not waived.
Finish all 28 teams by week 24 so roughly three months of operating evidence remain before audit fieldwork.
- Wave 1: remaining Tier 0. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team kit: catalog complete, alerts migrated and inside budget, runbooks at the readiness bar, rotation of 6+ certified responders, one IC nominee, one drill passed, paid rotation in payroll.
- Named coach for three weeks. The director signs the gate.
- Freeze legacy paging per wave.
- Publish a live adoption scoreboard.
- Any production service that cannot provide sustainable ownership gets an executive-reviewed deadline, compensating control, and expiry date. No indefinite verbal exceptions.
23. Change management, incentives, and the deal (depends on: 1, 4, 9)
Run this from day one in parallel with design. Engineers will judge fairness. Executives will judge visible results.
Repeat the deal until it is muscle memory: paid on-call, paged only for what you own, trained commander, real sprint capacity for actions.
- Launch via CTO all-hands, per-team roadshows, a one-page laptop card, an internal wiki, and a Slack help channel with a 4-hour answer SLA.
- Put incident-response contribution into promotion criteria. Award best postmortem and biggest noise cut quarterly. Thank people publicly after every SEV1.
- Put adoption, alert hygiene, action closure, and on-call load fairness into every engineering manager's quarterly objectives.
- Write an exception path for engineers who cannot do nights because of caring responsibilities or health, covered by stipended volunteers.
- Prohibit retaliation for good-faith declaration or escalation.
- Pulse-survey at 60 and 120 days. If fairness or load is red, pause expansion until fixed.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
24. SOC 2 evidence, mock audit, and sustainability (depends on: 17, 21, 22)
Evidence is a by-product of doing the work, not a reconstruction the week before the auditor. Protect the process after SOC 2 is signed.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to CC7.2–CC7.4, CC2.2/CC2.3, CC5, and availability A1.2. Confirm interpretation early.
- Maintain versioned, signed policies for incident response, severity, on-call, communications, and postmortems, reviewed annually.
- Automate evidence: incident records, paging and ack logs, status history, postmortem library, action closure, training register, drill records, and reportability decisions including not reportable.
- Internal dry-run at month 6 on a 15-incident evidence walkthrough. Mock audit at month 7. Fix with weeks to spare.
- Assign permanent owners for policy, pager, status page, catalog, training, and metrics.
- Year-two roadmap: follow-the-sun cell, self-healing for the top recurring causes, error-budget release gates, and blast-radius reduction on the shared ledger.
- Quarterly board summary of severe incidents, credits, overdue P0s, and resilience investment so attention does not die after the audit.
--- PROPOSAL 5 (agent deepseek-v4-pro_refine_5, deepseek/deepseek-v4-pro) ---
Estimated complexity: high
Success metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 6 months of full rollout.
- Customer-first detection falls from 40% to below 10% within 6 months and below 5% within 12.
- Median time to mitigate SEV0/SEV1 falls from 3h10 to under 60 minutes within 12 months; SEV2 under 2 hours.
- Incident Commander assigned and announced within 5 minutes for 95% of SEV0/SEV1 incidents; zero incidents with unclear command beyond 10 minutes.
- Status page updated within policy time for 95% of SEV0/SEV1/SEV2 (10/15/30 minutes).
- Monthly pages fall from 3,400 to under 500 with actionability above 75%.
- All six legacy alerting tools decommissioned by week 16.
- 100% of SEV0/SEV1 incidents have a blameless postmortem published within 10 business days.
- Postmortem action closure rises from 17% to 90% for P0/P1 actions on time.
- SLA credits fall from $1.3M to under $400K in first 12 months.
- Customer-impacting incidents decline to under 15 per year; repeat root causes under 10%.
- 100% of 180 services have a named owning team and criticality tier.
- All 28 teams onboarded by week 24; every Tier 0/1 team has 24x7 primary+secondary coverage with 6+ certified responders.
- At least 30 certified Incident Commanders and 20 certified Communications Leads active.
- Paid on-call policy is approved and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among on-call engineers.
- Quarterly game days and twice-yearly unannounced paging drills executed on schedule, each with tracked action items.
Steps (26):
1. Executive mandate, program governance, and interim incident command
Convert the CEO's outage complaint into a company-level improvement program with a named owner, budget, and deadlines. In the first week, establish an interim command process so no incident remains unowned while the permanent process is designed.
- Appoint Head of Reliability as program owner and CTO/CEO as executive sponsor.
- Form steering group with Engineering, SRE, Support, CS, Legal, Compliance, HR, Finance, Security.
- Approve non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Fund tooling, on-call compensation, training, and 3 dedicated program FTEs; anchor to $1.3M credits.
- Set timeline: design weeks 1-4, tooling and pilot weeks 5-10, rollout weeks 11-20, audit evidence collection from week 8, dry-run month 6.
- Stand up interim 24x7 duty officer and single declaration path within 48 hours, compensated retroactively.
2. Baseline data and alert estate analysis (depends on: 1)
Re-open the last 31 incidents and profile the current alert estate so every later design decision is evidence-based.
- Re-code each incident: detection source, timestamps, owner, severity, credits, root-cause family.
- Quantify customer-first detections and missing signals.
- Analyze two 'nobody in charge' incidents minute-by-minute.
- Inventory six alerting tools: volume, noise, owner, runbook coverage, top 50 noisy rules.
- Freeze baseline metrics: MTTD 22m, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed.
3. Stakeholder listening and resistance mapping (depends on: 1, 2)
Treat engineer pushback as design input. Interview all 28 teams plus Support, CS, and Sales to understand real objections and find early adopters.
- Test objection: unpaid work, night load, unfamiliar code, poor runbooks, or blame.
- Document current informal practices from 12 on-call teams.
- Recruit 10-15 credible engineers as co-design working group.
- Survey baseline sentiment on on-call and alerts.
- Publish promise: paid on-call, page only for owned services, trained commander, reserved reliability capacity.
4. Service ownership catalog and criticality tiering (depends on: 2, 3)
Build a machine-readable service catalog that assigns each of the 180 services a single owning team, escalation policy, and criticality tier.
- Define fields: owner team, manager, Slack channel, escalation policy, dependencies, dashboards, runbooks.
- Tier 0: ledger, money movement, auth, shared PostgreSQL; Tier 1 customer-facing; Tier 2 internal; Tier 3 non-critical.
- Map customer-visible capabilities and dependencies across regions.
- Identify orphan and shared services; force ownership decision within 30 days or schedule decommissioning.
- Publish coverage gaps by tier; Tier 0 gaps become executive escalations.
5. Define severity levels and declaration triggers (depends on: 4)
Adopt a five-level severity model with objective payment-specific triggers, so declaration is a lookup rather than debate. Anyone may declare; only the Incident Commander may downgrade.
- SEV0: unauthorized, lost, duplicated, or corrupted money movement; ledger integrity loss; confirmed data breach; both regions failed. Page all roles, exec, legal, and consider payment pause.
- SEV1: widespread payment failure, no workaround, severe SLA breach, or single-region total loss. Full role activation and public status page.
- SEV2: significant degradation, multiple customers, or workaround available but SLA risk. IC and SMEs paged; status page if customer-visible.
- SEV3: limited impact with workaround; team-led, business-hours response.
- SEV4: internal/no customer impact; ticket only.
- Auto-escalation: unresolved SEV3 >2h becomes SEV2; unresolved SEV2 >1h becomes SEV1; ledger or security always at least SEV2.
- Provide decision tree and 12 worked examples from actual incidents.
6. Define incident roles, decision authority, and handover rules (depends on: 5)
Codify five roles with written responsibilities and explicit authority, so there is never confusion about who is in charge.
- Incident Commander: owns severity, priorities, cross-team pulls, mitigation decisions; does not code.
- Communications Lead: owns status page, internal updates, account manager briefs, regulator coordination.
- Scribe: maintains timeline, decisions, and evidence for postmortem/audit.
- Subject-Matter Responders: diagnose and remediate only their owned services.
- Executive Liaison (SEV0/SEV1): handles exec communications and external stakeholders.
- Rule: IC identified within 5 minutes and announced in channel; handover announced and logged; roles may combine below SEV2, never at SEV0/SEV1.
7. Design 24x7 incident command and comms staffing model (depends on: 6, 4)
Create a central, trained incident command rotation instead of making each of 28 teams field its own commander.
- Recruit 24-30 certified ICs and 12-16 Comms Leads from a cross-team volunteer pool with manager approval.
- Weekly rotations, primary and secondary; 5-minute acknowledgement SLA with auto-escalation.
- Scribe pool as entry-level rotation.
- Night coverage: paid command rotation in New York timezone initially; evaluate follow-the-sun coverage later.
- Eligibility: certification required; commanders can leave with 30 days notice.
- Ensure distinct persons for IC, CL, and primary SME on SEV0/SEV1.
8. Design team on-call rotations and cross-team escalation policies (depends on: 4, 6)
Define tiered team on-call obligations so engineers are only paged for services they own, and cross-team pages go through the IC.
- Tier 0/1 services: 24x7 primary+secondary, at least 6 trained responders per rotation, one-week shifts.
- Tier 2/3: business-hours on-call; after-hours escalation list to engineering manager.
- Platform, infrastructure, database: 24x7 due shared ledger and Kubernetes.
- Routing: every page resolves via service catalog to owning team's escalation policy; cross-team pages only by IC.
- Guardrails: max one primary week in four, no consecutive weeks, no on-call in week after leading SEV0/1, protected recovery time after night work.
9. Design paid on-call compensation and fatigue safeguards (depends on: 8, 3)
Make on-call paid and legally compliant before any new rotation starts, and convert unpaid pager culture into a fair employment condition.
- Weekly stipends: primary 24x7 $800-1,200, secondary 30-50%, business-hours $300-500; holidays premium.
- Out-of-hours incident pay: $150 per night page plus hourly beyond one hour; time-off-in-lieu after overnight work.
- Additional command rotation stipend and SEV0/1 bonus for active responders.
- Verify FLSA/NY wage rules with HR, Legal, Finance; document exemption treatment.
- Base annual budget against current $1.3M credits.
- Track load and trigger staffing review if >2 after-hours pages per person per week sustained.
10. Enforce alert quality standards and noise budget (depends on: 2, 4, 8)
Replace 3,400 monthly alerts at 85% noise with a contractual paging standard that makes on-call sustainable.
- Every paging alert must have owning team, customer impact statement, runbook link, severity mapping, and tested threshold.
- Page only on customer-visible symptoms or SLO burn; cause-based alerts become tickets or dashboards.
- Set page budget: max 2 out-of-hours pages per person per week; breach triggers mandatory alert-tuning sprint.
- Auto-quarantine alerts with >5 firings/month without action or >70% no-action acknowledgements.
- Target <500 actionable pages/month and >75% actionability within 6 months.
- Weekly per-team alert review, monthly cross-team review.
11. Build detection uplift: synthetics, SLOs, and support intake (depends on: 4, 10)
Shift detection from host metrics to customer outcomes so the company stops hearing about outages from clients first.
- Define SLOs per Tier 0/1 capability: payment initiation, auth, settlement timeliness, API availability, ledger consistency.
- Deploy external synthetic transactions from both regions every 60 seconds, covering full payment flow and ledger write.
- Add ledger assurance checks: replication lag, double-entry balance, settlement window countdown.
- Top-100 customer anomaly detection to catch single-tenant outages.
- Auto-create triage incident from support tickets or account manager keywords within 5 minutes.
- Track customer-detected-first as a defect and require a postmortem action.
12. Consolidate alerting and incident tooling (depends on: 5, 8, 10, 11)
Collapse six alerting tools into one integrated paging and incident management platform to create a single system of record for people and audit.
- Select paging/on-call platform and incident management layer (e.g., PagerDuty + incident.io/FireHydrant).
- Implement one-command Slack declaration that auto-creates channel, bridge, pages roles, sets severity, starts timeline.
- Migrate all monitoring sources into the one tool; decommission legacy paging only after two weeks verified.
- Integrate service catalog, status page, Jira action tracking, Salesforce/CS customer lists, and conference bridge.
- Ensure out-of-band paging and offline fallback if a region or chat tool is down.
- Automate evidence capture for SOC2: timestamps, role assignments, severity changes, comms sent.
13. Define acknowledgement and escalation paths (depends on: 12, 6, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts at first signal.
- For SEV0/SEV1: page primary on-call; after 5 minutes unacked page secondary; at 10 page manager and duty IC; at 15 page executive officer.
- For SEV2: primary ack within 10 minutes, IC assigned within 15; escalate on miss.
- For SEV3: ack within 30 minutes or create work item.
- Automatic duty IC page for ledger/security, cross-team, or unresolved ownership.
- If impact unknown after 15 minutes, raise severity.
- Cross-team responders summoned by IC have 10-minute acknowledgement obligation.
- Route alerts with no owner to command rotation, then treat missing ownership as control defect.
- Human acknowledgement required; delivery confirmation is not sufficient.
14. Standardize internal communications (depends on: 6, 5)
Separate the working war room from executive and stakeholder updates, with fixed cadence and pre-approved templates.
- Auto-create incident channel and one read-only broadcast channel for execs, support, sales.
- SEV0: internal update every 15 minutes; SEV1 every 30; SEV2 every 60; SEV3 on state change.
- Template: impact, what we know, what we are doing, ETA/next update, current IC and CL.
- Executives questions through Executive Liaison only; IC is not interrupted.
- Support/CS receives affected-customer list and holding statement within 15 minutes for SEV0/SEV1 and 30 for SEV2.
- Handover protocol for incidents lasting >4 hours: formal IC handover and fatigue check.
15. Standardize customer and status page communications (depends on: 5, 14)
Replace ad-hoc status page updates with a timed, role-owned, template-driven process, including account manager outreach.
- Status page timing: SEV0 initial post within 10 minutes, SEV1 within 15, SEV2 within 30; updates every 15/30/60 minutes until resolved.
- Resolution notice within 30 minutes of mitigation; customer-facing summary in 5 business days for SEV0/SEV1.
- Pre-approve 12-15 templates with Legal and Comms.
- Tiered outreach: top 100 accounts get direct call/email from AM within 30 minutes of SEV0/SEV1; long tail via subscription.
- Use factual language: state impact and next update; never speculate cause or blame vendor.
- Comms Lead is sole author for customer language.
16. Regulatory, legal, and account manager notification playbook (depends on: 5, 15)
Build a notification decision tree and contact matrix so legal/regulatory obligations are assessed early and never forgotten.
- Map obligations: NYDFS Part 500 72-hour cybersecurity event notification, state breach laws, GLBA/FTC, PCI, sponsor bank/card network contractual windows, FinCEN/OFAC if relevant.
- Add regulatory assessment checkpoint for every SEV0 and security SEV1 within 2 hours, even if not reportable.
- Maintain 24x7 contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel with named backups.
- Pre-draft notification templates and test quarterly.
- Encode customer-specific notification SLAs from enterprise contracts into customer tiering.
- Account managers receive legal-approved script and affected-customer list.
17. SLA credit workflow and financial impact measurement (depends on: 5, 15)
Link every incident to money automatically so severity, credits, and prioritization stay consistent and Finance is never surprised.
- Define availability measurement per contract and capability with Legal/Finance.
- Auto-compute affected minutes per customer from incident record and telemetry; generate credit proposal within 5 business days.
- Decide proactive credits for top tier vs claims-based for others; document approval chain.
- Track credits by root-cause family and incident; feed quarterly reliability investment decisions.
- Target reducing annual credits from $1.3M to under $400K in first year.
18. Standardize blameless postmortems (depends on: 5, 6)
Make postmortems mandatory with fixed deadlines, a single format, and a blameless review forum, replacing 'various formats'.
- Mandatory: all SEV0/SEV1; SEV2 with customer impact, credits, >2h, repeat cause, or detected by customer; near-miss involving ledger.
- Draft within 3 business days, peer review within 5, publish within 10.
- One template: timeline, customer/financial impact, detection gap, response gap, contributing factors, what went well.
- Blameless rules: context-based, no individual blame, never used in performance reviews.
- Weekly Incident Review Board reviews all postmortems and challenges quality.
- Publish searchable postmortem library and quarterly recurring causes report.
19. Track postmortem actions with owner and due date (depends on: 18, 12)
Fix the 11/64 completion rate by giving every action the same status as customer commitments, with capacity and escalation.
- Every action gets named owner, priority, due date, and Jira ticket auto-created from postmortem.
- P0 actions prevent SEV0 recurrence, due 30 days; P1 60; P2 90.
- Reserve 15-20% team sprint capacity for incident actions.
- Escalation ladder: manager at 7 days overdue, director at 14, CTO dashboard at 30; P0 overdue blocks features.
- Monthly reporting of closure rate in engineering leadership.
- Target 90% closure of P0/P1 on time in two quarters.
20. Create runbooks and service readiness bar (depends on: 4, 8, 11)
Ensure every service is prepared for 3 a.m. response before it is allowed to page anyone.
- Readiness checklist for Tier 0/1: architecture diagram, dependencies, dashboards, rollback, feature flags, escalation contacts, data-loss statement.
- Write major-incident playbooks: shared PostgreSQL failure, cross-region failover, Kubernetes control plane loss, partner outage, settlement breach, security compromise.
- Prioritize ledger playbooks: documented failover, read-only mode, reconciliation, signed RPO/RTO.
- Runbooks must be tested twice yearly; stale runbooks marked in catalog.
- No paging alerts without readiness sign-off; gap reported to director.
21. Train, certify, and simulate incident response (depends on: 6, 13, 14, 15, 18, 20)
Build a certification path so 24x7 roles are staffed by people who have practiced, and validate the process through drills.
- Scribe course (2 hrs); Responder (half-day); Communications Lead (1 day); Incident Commander (2 days + shadowing/tabletop).
- Certify ICs and CLs; recertify annually.
- New responders shadow two shifts before primary; no primary in first 90 days.
- Monthly tabletop using real past incidents; quarterly game day including regional failover or ledger scenario.
- Twice-yearly unannounced paging drill; one regulator/legal exercise annually.
- Track drill metrics and action items.
22. Pilot on critical services and iterate (depends on: 12, 13, 15, 19, 21)
Prove the process on the highest-risk surface before full rollout. Run a six-week pilot with tight measurement and public results.
- Select 5-6 teams: payments core, ledger/database, platform/Kubernetes, API gateway, plus two existing on-call teams.
- Activate full stack: severity, roles, command rotation, paid on-call, single tooling, alert budget, status page, postmortems.
- Weekly pilot retro; fix defects within 48 hours.
- Exit criteria: MTTD <10 min, IC assigned within 5 min in 95% incidents, pages down 50%, postmortems on time, positive sentiment.
- Publish one-page pilot result as adoption argument.
23. Wave rollout to all teams and decommission legacy paths (depends on: 22, 20)
Roll out to all 28 teams in four waves by criticality, with explicit readiness gates and legacy tool shutdown.
- Wave 1 remaining Tier 0; Wave 2 Tier 1; Wave 3 Tier 2; Wave 4 Tier 3/internal.
- Per-team onboarding: catalog complete, alerts pruned, runbooks ready, rotation staffed with 6+ certified responders, one IC candidate, one drill passed.
- Coach per wave 3 weeks.
- Gate signed by director; failures rescheduled, not waived.
- After onboarding, freeze legacy alerting paths; no fallback to old tools.
- Publish adoption scoreboard.
24. Establish metrics, dashboards, and review cadence (depends on: 12, 19, 22)
Make process performance visible with a small set of metrics and fixed review meetings, so the program is owned by data.
- Response: MTTD, time to IC, MTTA, MTTM, MTTR, % customer-detected-first.
- Quality: page volume per person, alert actionability, postmortem on-time, action closure rate.
- Business: SLA credits, availability vs 99.95, repeat incidents.
- People: on-call load, page per engineer, sentiment, attrition.
- Cadence: weekly Incident Review Board, monthly Reliability Review, quarterly Executive/Board review.
- Dashboards self-serve with targets and named owners.
25. SOC 2 readiness and internal dry-run audit (depends on: 19, 23, 24)
Design evidence as a by-product and test it with an internal walkthrough before the external auditor arrives.
- Map process to SOC2 CC7.3/CC7.4, CC7.2, CC2.2/2.3, CC5, availability criteria.
- Publish versioned policy documents: Incident Response, Severity, On-Call, Communication, Postmortem.
- Automate evidence: incident records with timestamps, role assignments, paging logs, status history, postmortems, action board, training register, drill records.
- Ensure process operates at least 3 months before fieldwork.
- Run internal dry-run month 6; sample 15 incidents; fix gaps with 8 weeks to spare.
- Keep remediation log for process deviations.
26. Continuous improvement, culture, and sustainability (depends on: 23, 25)
Prevent the process from decaying after the audit by embedding review, feedback, and roadmap ownership.
- Quarterly process retrospective with IC pool and responders.
- Re-baseline metrics every six months; raise targets.
- Year-two roadmap: follow-the-sun coverage, self-healing top 3 causes, error budgets gating releases, blast-radius reduction for shared ledger.
- Annual policy review and certification renewal.
- Quarterly on-call sentiment survey with published actions.
- Board quarterly report on availability, credits, and incident trends.
Your answer has these parts:
- "round_summary": one or two sentences on how the round went as a whole.
- "converging": true if the proposals of this round are more similar to each other than those of the previous round, false otherwise.
- "differences": a short list (four items at most) of what still separates the proposals of this round, naming the proposals and the steps concerned.
- "influences": a short list (five items at most) of the flows of ideas of the round: who took what from whom, and what nobody took, naming proposals and steps ("P1 and P3 took P2's CDC sync (steps 2 and 13)").
- "proposals": one entry per proposal of round 1, each with:
- "proposal": its number,
- "assessment": "improved", "worsened", "mixed" or "unchanged" with respect to its previous version ("no_previous_version" if that agent produced nothing in the previous round),
- "what_changed": how it improved or worsened and why: two or three short sentences, or a list,
- "improvements": a list of concrete gains (specific steps, metrics, structure),
- "regressions": a list of concrete losses (dropped steps, vaguer metrics, broken dependencies...),
- "taken": the ideas this proposal visibly adopted from the OTHER proposals of round 0 (not from its own previous version): one entry per idea with "from_proposal" (the number of the proposal it came from), "steps" (the numbers of the steps of that proposal where the idea lives, as listed above; empty if it is not tied to specific steps), "what" (the idea, one sentence) and "why" (how it was used or adapted, one sentence),
- "rejected": the ideas of the OTHER proposals of round 0 that this proposal visibly declined: an explicit contradiction, or a prominent idea it saw and left out while taking the opposite approach. Same fields; "why" gives the evidence (what the proposal does instead). Do not list mere omissions without evidence; an empty list is a valid answer.
[ROUND 2]
[SYSTEM]
You are an expert reviewer of multi-agent planning processes.
Several LLM agents drafted plans for a task, refined them over a number of rounds while seeing each other's proposals, and finally voted for the best one.
Be exhaustive but precise: name concrete steps, ideas and metrics, never generalities. Judge plans by their fitness for the task as stated, their realism, their completeness, the soundness of their order and dependencies, how measurable their success is and how they handle things going wrong.
You are an impartial evaluator, not a chronicler: assess the proposals and the process on their merits, never rationalise what happened or assume that the outcome was right.
After your analysis, answer in the requested structure.
Every text field you write will be read by a busy person who skims. Make it easy to skim: short sentences and short paragraphs; when you name several things, prefer a list to a paragraph, with sub-items when an item has parts, but keep a single fact as a sentence; lead with the point and then the evidence; name proposals and steps by number (P2, step 4); no preamble, no repetition of the question, no closing summary; bold at most one key phrase per item or paragraph. Text fields accept Markdown: a blank line between paragraphs, "- " for lists, **bold**.
[HUMAN]
Task given to the agents: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
This is round 2, a refinement round: every agent received ALL the proposals of round 1 and wrote a new plan, improving on them or taking a different approach. By convention, the previous version of proposal N is proposal N of round 1, written by the same model.
PROPOSALS OF ROUND 1 (the previous versions):
--- PROPOSAL 1 (agent opus5_refine_1, anthropic/claude-opus-5) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (32):
1. Executive mandate, single owner, funding and non-negotiables
Convert the CEO email into a chartered program with one accountable owner and authority over all 28 teams. Incident response becomes a **company operating process**, not a per-team choice.
- Name an executive sponsor (CTO) and one accountable owner (Head of Reliability / Incident Management) with a small permanent office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group with Engineering, SRE/Platform, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours. It is not a 28-person committee.
- Fix the non-negotiables now: one severity scale, one paging tool, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.
- Set the clock deliberately earlier than the audit: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready week 28.
- Approve budget against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises.
2. Seven-day interim command bridge (depends on: 1)
Do not let design work leave the company unprotected for six weeks. Put a crude but real process in place within seven days and improve it later.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster from engineering managers and senior SREs, primary plus backup, 24x7. Pay it retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Tell Support to escalate credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items: complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This becomes both the design input and the "before" picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start / detect / declare / commander-assigned / mitigate / resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. This list drives the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome. Identify the top 50 rules producing most of the noise and every rule with no owner or runbook.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS and Sales in two weeks. Test what the objection actually is: unpaid work, lost sleep, unfamiliar code, missing runbooks, or fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey (fairness, trust in alerts, willingness, burnout) to re-measure at 60, 120 and 365 days.
5. Service catalog, ownership and money-path tiering (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that is the single source of truth for paging, impact and audit.
- One accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not by technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions and third parties, including sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Every orphan service gets an owner in 30 days or a decommission date approved by the sponsor. Tier 0 without an owner is an executive escalation.
6. Severity scale, declaration rules and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- **SEV3 (major, contained):** narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- **SEV4:** no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring; only the commander may downgrade, with the evidence recorded.
- Lifecycle: Detected → Declared → Triaged → **Mitigated** (customer impact ends) → Monitoring → **Resolved** (backlog processed and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Incident roles, authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, roles, cadence and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit spend. Retains command when a VP joins.
- **Communications Lead:** single voice for status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates for SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team; they mitigate, they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance and Vendor Management join on defined triggers.
- Rules: command claimed within 5 minutes, stated in channel ("I am IC"), distinct people for command, comms and technical lead at SEV1/SEV2, and every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation or privileged-access rules.
8. 24x7 coverage model: central command corps, local expertise (depends on: 5, 7)
Do not create 28 night rotations. Centralise coordination in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- **Layer A — Incident Command corps:** ~30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of ~18 from Support, CS and engineering management; a scribe pool used as the training entry point.
- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- **Layer C — everyone else:** business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing, with cost and timeline: a US-only paid night rotation now, a Lisbon or APAC follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
9. On-call compensation, labour compliance and fatigue safeguards (depends on: 8, 4)
Unpaid on-call in New York is both a retention problem and a legal exposure. Pay for it before asking anyone to sign up, and publish the numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 for business-hours rotations, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover rather than expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, tax treatment and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a ~15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
10. Alert quality standard, page budget and noise burn-down (depends on: 3, 5)
3,400 alerts at 85% noise is the reason detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human, and burn the backlog down deliberately rather than by mass silencing.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a **page budget** of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements; return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
- Target 3,400 → under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
11. Detection uplift on the money path (depends on: 10, 5)
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls; five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrating without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack or the identity provider is unavailable. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live incident execution doctrine (depends on: 7, 12, 13)
Give responders one short operating procedure for the first minutes through closure. Priority is limiting customer and financial harm, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; controlled backlog drain; **reconciliation completed before any payment or ledger incident is declared resolved**.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and Customer Success, and reopening if impact recurs.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook and RSS; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack; the long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity or blame; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded **even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review; Legal owns outbound regulatory text, the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
18. SLA credit and financial impact workflow (depends on: 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one, and use the delta as the standing business case for on-call pay and reliability capacity.
19. Blameless postmortem standard and Incident Review Board (depends on: 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first?** and **why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, capacity reservation and enforcement (depends on: 19, 12)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
21. Runbooks, readiness bar and ledger blast-radius reduction (depends on: 5, 8)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the single largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel workstream on blast-radius reduction — tenant or function partitioning, read replicas, and isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (depends on: 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days + shadowing):** command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence for their own domain before independent primary duty; two shadow shifts minimum, never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (depends on: 22, 12, 21)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 11, 12, 21, 22, 9)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect (from first impact), time to declare, time to commander, acknowledgement, mitigate, resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; team scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Wave rollout to all 28 teams with readiness gates (depends on: 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director; failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
28. Alert noise burn-down campaign (depends on: 10, 12, 25)
Run the noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
29. Change management, fairness and pager culture (depends on: 4, 9, 25)
Run this from day one in parallel. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (depends on: 24, 26, 27)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and contingencies (depends on: 1)
Name the ways this programme fails and pre-commit the response. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 26, 27, 30)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 2 (agent gpt5.6-sol_refine_2, openai/gpt-5.6-sol) ---
Estimated complexity: high
Success metrics: - By day 7, every suspected SEV0–SEV2 uses one incident record, one coordination channel, and one named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named commander for more than 10 minutes; at least 95% are assigned within 5 minutes.
- By day 30, 100% of Tier 0 services have a named owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, every Tier 0 and Tier 1 customer journey has compensated 24×7 command and subject-matter coverage.
- By day 120, all 180 production services have an owner, response tier, tested escalation path, and appropriate coverage model.
- No mandatory night rotation starts before its compensation, training, access, and staffing controls are active.
- Every critical primary rotation has at least six qualified responders or a documented, expiring executive exception by day 120.
- No responder is routinely scheduled for primary duty more often than one week in six by day 120.
- At least 95% of critical pages are acknowledged within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by day 120 and under 60 minutes by month 8.
- At least 95% of SEV0 and SEV1 customer notices are issued within 15 minutes of declaration by month 3.
- At least 95% of customer-visible SEV2 notices are issued within 30 minutes by month 3.
- At least 95% of incidents meet their required update cadence by month 3.
- Monthly paging volume falls from 3,400 to no more than 1,500 by day 90 and no more than 700 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6 without loss of critical detection coverage.
- All new paging alerts satisfy the owner, impact, action, dashboard, runbook, deduplication, and escalation standard by day 60.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 currently open historical actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date by month 6, with effectiveness evidence.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Two end-to-end cross-company exercises, including regional failure and ledger recovery, are completed before the audit.
- Monthly availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- SLA credits decline by at least 50% on an annualized trailing basis within 12 months.
- Quarterly on-call surveys show improving fairness and sustainability, with at least 75% favorable responses by month 6.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (24):
1. Create the mandate, ownership, and funding
Launch incident management as a company operating program within 48 hours. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable program owner.
- Form a small steering group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Fund two to three implementation staff, paging and incident tooling, observability work, training, exercises, and on-call compensation.
- Reserve 10%–15% of engineering capacity for alert remediation, runbooks, and incident actions.
- Authorize incident commanders to freeze deployments, order rollback, disable features, shift traffic, and invoke continuity plans.
- Preserve financial controls. Incident commanders may coordinate ledger recovery but may not bypass dual approval, privileged-access controls, or reconciliation.
- Set the delivery target: critical controls operational within 60 days, enterprise rollout within 120 days, and a mock audit in month 7.
2. Install an interim process in seven days (depends on: 1)
Do not wait for new tools or the final policy. Put a minimum viable incident process into operation immediately and start collecting evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communications schedule.
- Establish one monitored declaration path through chat, telephone, and the existing paging tools.
- Create a standard incident channel, bridge, incident document, and naming convention.
- Staff temporary primary and backup incident commanders 24×7 from the existing on-call teams and engineering leadership.
- Compensate interim duties retroactively under the final compensation policy.
- Require a named incident commander within 10 minutes for every suspected major incident.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Hold a daily 15-minute control review until the permanent process is live.
3. Build the baseline, ownership catalog, and risk map (depends on: 1)
Establish the facts behind the current failures and assign every production component an owner. Use the resulting catalog as the source for routing, escalation, and audit evidence.
- Reconstruct all 31 customer-impacting incidents, including first impact, detection source, declaration, command assignment, mitigation, resolution, customer communications, and credits.
- Analyze the two incidents with no clear leader and every case detected first by customers.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, databases, queues, regional dependencies, payment processors, and banking partners.
- Assign one accountable team, engineering manager, product owner, primary escalation, secondary escalation, dashboard, and runbook to each service.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 based on financial integrity, customer impact, contractual exposure, and dependency centrality.
- Map critical customer journeys to their application, PostgreSQL, Kubernetes, regional, and third-party dependencies.
- Inventory the six alert sources and all 3,400 monthly alerts by owner, volume, actionability, and duplication.
- Interview representatives from all 28 teams and baseline on-call sentiment, fatigue, and objections.
- Give orphan services an owner or decommissioning decision within 30 days.
4. Design compliance and evidence controls from day one (depends on: 1)
Map the operating process to audit, legal, contractual, and record-retention requirements before finalizing it. Confirm the expected SOC 2 observation period with the auditor immediately.
- Map controls to applicable SOC 2 criteria for monitoring, incident identification, response, recovery, communications, corrective action, access, and availability.
- Define evidence required for declarations, pages, acknowledgements, role assignments, decisions, status updates, postmortems, actions, training, drills, and exceptions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve the Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, and Postmortem Standard.
- Record control exceptions with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing evidence before the audit.
5. Adopt one severity and lifecycle standard (depends on: 2, 3, 4)
Use four impact-based severity levels across operational, security, data, and third-party incidents. Start at the highest credible severity when facts are uncertain, then downgrade with recorded evidence.
- **SEV0 — financial or security crisis:** suspected ledger corruption, unauthorized or duplicated funds movement, material data compromise, both-region loss, or a decision to suspend payment processing. Page all roles immediately; engage executives, Security, Legal, Compliance, and Risk; assess regulatory duties within one hour; use distinct role holders; require a postmortem.
- **SEV1 — critical availability event:** a core payment journey is unavailable, payment failures exceed an initial 10% guardrail for five minutes, regional loss has impaired failover, a settlement deadline is at imminent risk, or rapid error-budget burn makes a material SLA breach likely. Assign all roles; notify internal stakeholders within 10 minutes; publish customer status within 15 minutes; update every 30 minutes; require a postmortem.
- **SEV2 — major bounded event:** approximately 1%–10% of payment attempts fail, a material customer subset or critical customer is down, degradation has a workaround, or contractual impact is likely. Assign an incident commander and responders; add communications and scribe roles for customer impact; publish status within 30 minutes; update every 60 minutes; require a postmortem for customer-visible events.
- **SEV3 — limited event:** localized impact, a safe workaround, and no financial-integrity, security, regulatory, or material contractual risk. The owning team leads; page only if immediate action is necessary; use a ticket otherwise.
- Treat the percentage thresholds as declaration guardrails, not reasons to under-classify integrity, settlement, security, or strategic-customer risk.
- Permit any employee to declare an incident. Only the incident commander may lower severity, with the rationale logged.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of active customer harm. Declare resolution only after stability, backlog recovery, and required ledger reconciliation.
6. Define roles, authority, and handoffs (depends on: 5)
Separate command, communication, recordkeeping, and technical repair. One named person must hold command at every moment of a major incident.
- **Incident Commander:** owns severity, priorities, role assignment, escalation, decision cadence, mitigation coordination, and closure. The commander does not act as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language. Legal or Compliance retains ownership of regulatory submissions.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, and status changes.
- **Subject-Matter Responders:** diagnose and mitigate services for which they have accepted ownership, access, training, and runbooks.
- **Executive Duty Officer:** removes organizational obstacles and approves exceptional business decisions without displacing the incident commander.
- Require separate commander, communications, scribe, and primary technical lead for SEV0 and SEV1.
- For customer-visible SEV2, keep the commander separate from the primary technical responder; communications and scribe may be combined if workload permits.
- Announce every role assignment and transfer in the incident channel. Require verbal and written handoff with current impact, decisions, risks, and next actions.
- Keep executives and account managers out of the technical command path; questions flow through the Executive Duty Officer or Communications Lead.
7. Create sustainable 24×7 coverage across the service estate (depends on: 3, 6)
Use central command coverage and service-domain responder coverage rather than creating 28 fragile night rotations. Engineers remain responsible only for services they own or have formally trained to support.
- Build a company incident-command pool of approximately 18–24 certified senior engineers and managers, with a primary and backup scheduled at all times.
- Build 12–16-person communications and scribe pools using Support, Customer Operations, Engineering Operations, and qualified engineering managers.
- Schedule an Engineering Director or equivalent as the 24×7 Executive Duty Officer.
- Group related service owners into approximately 8–12 coherent responder domains only where members have training, access, and explicit acceptance.
- Require Tier 0 and Tier 1 domains to provide 24×7 primary and secondary responders, normally with at least six trained people in each sustainable rotation.
- Give Tier 2 services business-hours coverage plus a maintained manager escalation path. Treat Tier 3 conditions as tickets unless their impact changes.
- Maintain distinct but coordinated coverage for the ledger application and the PostgreSQL platform.
- Route unknown-owner events to the duty commander and platform triage temporarily. Treat every such event as an ownership control defect.
- If a small team cannot staff a fair rotation, merge coverage only after training or provide headcount, service reassignment, or decommissioning.
8. Approve compensation and fatigue protections (depends on: 7)
End unpaid on-call before expanding mandatory coverage. Publish the policy through HR, Legal, Finance, and Payroll within 14 days.
- Pay fixed weekly stipends for primary and secondary service rotations.
- Pay separate stipends for duty commander, communications, and scribe assignments.
- Provide additional call-out compensation or equivalent paid recovery time for material after-hours work.
- Apply overtime and reporting rules correctly for non-exempt employees under federal and New York requirements.
- Pay higher rates for company holidays and provide a protected recovery day after qualifying overnight work, SEV0 events, or prolonged SEV1 response.
- Target no more than one primary week in six and prohibit simultaneous primary assignments.
- Avoid consecutive primary weeks and make all swaps visible in the paging system.
- Reduce sprint commitments for people carrying primary duty rather than expecting normal delivery capacity.
- Provide a documented accommodation path for health, disability, or caregiving constraints without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
9. Set the service readiness and runbook standard (depends on: 3, 5, 7)
A team cannot respond effectively at night without ownership, access, telemetry, and rehearsed recovery procedures. Apply a formal readiness gate to every Tier 0 and Tier 1 service.
- Require a current architecture diagram, dependency map, dashboards, SLOs, runbooks, rollback method, feature-control method, contacts, and tested escalation path.
- Record RTO, RPO, data-integrity requirements, regional mode, and customer-facing capabilities in the service catalog.
- Link every paging alert to the exact runbook step expected from the responder.
- Prevent new paging alerts for services that fail readiness review. Preserve existing critical detection through a documented exception until a safe replacement exists.
- Write priority playbooks for PostgreSQL failure, ledger-integrity investigation, regional failover, Kubernetes control-plane degradation, payment-processor failure, queue backlog, credential compromise, and payment suspension.
- For the ledger, document read-only or stop-processing modes, failover controls, replay and duplicate protections, backlog handling, and post-recovery reconciliation.
- Require runbook review after material incidents and at least twice per year through exercises.
10. Establish one paging and incident system of record (depends on: 4, 5, 6, 7)
Consolidate paging and incident coordination without requiring an unsafe big-bang replacement of every monitoring system. Monitoring sources may remain specialized, but all human pages must enter one controlled platform.
- Select one enterprise paging platform and one integrated incident record, with chat, telephone, SMS, conference, status-page, ticketing, and service-catalog integrations.
- Initially ingest events from all six current tools, then deduplicate, correlate, and route by service ownership.
- Automate incident channel and bridge creation, role paging, timeline capture, severity changes, communications reminders, and postmortem creation.
- Preserve immutable records of declarations, acknowledgements, role assignments, decisions, and messages.
- Use role-based access, multifactor authentication, break-glass controls, and periodic access reviews.
- Provide telephone and offline fallback procedures for loss of chat, identity, the paging vendor, or an AWS region.
- Test paging and fallback paths weekly.
- Retire a legacy paging route only after its signals have owners, quality review, successful end-to-end tests, and at least two weeks of verified operation in the new path.
11. Enforce alert quality and burn down noise safely (depends on: 3, 10)
Treat paging alerts as production products with owners and quality requirements. Do not reduce noise by silently disabling detection.
- Require every page to identify the service, owner, customer or SLO risk, urgency, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Define an actionable page as one requiring prompt human judgment or intervention that materially reduces customer, financial, security, or contractual risk.
- Route informational, capacity-planning, and non-urgent conditions to dashboards or ticket queues.
- Prefer symptom and error-budget burn alerts over raw CPU, memory, pod, or log-volume thresholds.
- Run new alerts in shadow mode for at least seven days unless an emergency exception is approved.
- Review alerts with less than 50% actionability, more than three firings in seven days, or repeated no-action acknowledgements within two business days.
- Require compensating detection and an owner before suppressing or removing an alert.
- Set a responder load target of no more than two after-hours pages per person per week. A breach creates a mandatory alert-remediation plan.
- Review alert actionability, duplication, missed detection, and page load monthly by domain.
- Prioritize the small number of rules producing most of the current 85% noise.
12. Detect payment and ledger failures before customers (depends on: 3, 11)
Shift detection from infrastructure symptoms to customer journeys and financial outcomes. Set internal objectives stricter than the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, webhooks, and reporting freshness.
- Run external synthetic transactions from outside the production boundary and through both regions at least every minute for critical paths.
- Monitor payment failure rates, latency, queue age, delayed value, settlement-window risk, regional asymmetry, and third-party response quality.
- Add continuous ledger controls for reconciliation breaks, unexpected balances, duplicate identifiers, replication lag, backup health, and failover readiness.
- Add tenant or cohort anomaly detection for high-value customers and common payment methods.
- Convert high-priority Support, account-manager, bank, and processor reports into incident candidates within five minutes.
- Review every customer-first incident as a missed-detection defect and create a corrective action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Codify escalation and live incident execution (depends on: 6, 10)
Create one time-bound path from signal to ownership and mitigation. Delivery of a notification does not count as acknowledgement.
- Page the owning Tier 0 or Tier 1 primary immediately; page the secondary after five unacknowledged minutes; page the domain manager at 10 minutes; escalate to the engineering director at 15 minutes.
- For SEV0–SEV2, page the duty commander immediately. Page the backup after five minutes and require the Engineering Director to assume command if no certified commander owns the event by 10 minutes.
- Automatically involve command for integrity concerns, security concerns, regional events, cross-team impact, critical customer journeys, or unresolved ownership.
- If impact remains materially unknown after 15 minutes, increase response posture rather than waiting for certainty.
- Open one incident channel, bridge, and system record. State severity, known impact, assigned roles, current objective, and next update time.
- Freeze unrelated changes during SEV0 and SEV1 unless the commander records an exception.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require controlled backlog processing and reconciliation before resolving payment or ledger incidents.
- Require a formal command handoff for incidents extending beyond four hours or when fatigue impairs a role holder.
14. Standardize internal and customer communications (depends on: 5, 6, 10)
Communicate known impact early without waiting for root cause. The Communications Lead uses approved facts and always states the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish customer status within 15 minutes when customer-visible and legally safe, then update at least every 30 minutes.
- For SEV1, brief internal stakeholders and publish status within 15 minutes, then update every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send notice within 30 minutes, then update every 60 minutes.
- For SEV3, communicate directly to affected customers only when impact or contract terms require it.
- Give account managers an approved statement and affected-customer list within 30 minutes for SEV0 or SEV1 and within 60 minutes for SEV2.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. Describe affected capabilities, symptoms, workarounds, and the next update.
- Do not speculate about root cause, blame, data integrity, security scope, or recovery time.
- Post a Monitoring update within 30 minutes of mitigation. Post Resolved only after stability and required reconciliation.
- Provide a customer-facing incident summary within five business days for qualifying incidents.
- Record and approve any delay or restriction of public details during an active security threat.
15. Operationalize regulatory, partner, contract, and credit decisions (depends on: 4, 5, 14)
Give Legal and Compliance a timed decision process while keeping technical command with the incident commander. Record both reportable and non-reportable determinations.
- Build a jurisdiction and obligation matrix covering applicable NYDFS requirements, state breach laws, GLBA or FTC requirements, PCI obligations, money-transmitter rules, sponsor banks, payment networks, cyber insurance, and customer contracts.
- Validate applicability and deadlines with counsel rather than assuming every operational incident is reportable.
- Begin reportability assessment immediately for every SEV0, security event, suspected ledger-integrity event, and relevant SEV1.
- Target a documented initial legal assessment within one hour for SEV0 and within two hours for other potentially reportable events.
- Record the decision, evidence, approver, legal deadline, submission owner, and confirmation of delivery.
- Maintain tested 24×7 contacts for regulators, banks, networks, insurers, outside counsel, and critical vendors.
- Encode customer-specific notification clocks and channels in the customer record used by the Communications Lead.
- Have Finance calculate affected minutes, likely credits, and contractual exposure from the incident record within five business days.
- Review whether proactive credits or claims-based handling applies by customer segment and contract.
16. Make postmortems mandatory, consistent, and blameless (depends on: 5, 6, 10)
Use one review standard to learn from incidents and test whether controls worked. Keep learning reviews separate from performance or misconduct processes.
- Require a postmortem for every SEV0 and SEV1.
- Require one for customer-visible SEV2, customer-first detection, incidents lasting more than two hours, contractual breaches, repeat failures, control gaps, and ledger-integrity near misses.
- Produce the factual draft within three business days, conduct the review within five, and publish the approved version within 10.
- Use one template covering summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, lessons, and actions.
- Analyze why detection was not earlier and why mitigation took as long as it did.
- Examine technical, organizational, process, testing, dependency, and incentive factors rather than forcing a single root cause.
- Use a trained facilitator and written blameless rules. Describe decisions in the context and information available at the time.
- Publish broadly useful findings internally while maintaining restricted versions for security, privacy, personnel, or privileged material.
- Review postmortem quality and recurring factors in a weekly Incident Review Board.
17. Enforce action ownership and effectiveness tracking (depends on: 16)
Treat corrective actions as risk commitments rather than suggestions. Closing a ticket is insufficient without evidence that the control or system behavior improved.
- Give every action one named individual owner, manager, priority, due date, expected risk reduction, verification method, and linked work item.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with intermediate milestones.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Place recurrence-prevention actions for SEV0 and SEV1 ahead of discretionary feature work unless an executive accepts the residual risk.
- Escalate overdue high-risk actions to the manager after seven days, director after 14 days, and CTO after 30 days.
- Require written residual-risk acceptance, compensating controls, and a new review date when high-risk work is deferred.
- Verify completed actions through tests, telemetry, drills, or production evidence.
- Triage the 53 currently open historical actions within 30 days. Complete, re-plan, or formally accept the risk, prioritizing ledger, regional, security, and detection items.
18. Measure outcomes, controls, business impact, and human load (depends on: 10, 11, 14, 17)
Use a balanced scorecard so teams are not rewarded for suppressing alerts or avoiding incident declarations. Report medians and 90th percentiles, not averages alone.
- Measure time from first impact to internal detection, declaration, acknowledgement, command assignment, mitigation, and resolution.
- Track customer-first detection, missed escalations, status-page timeliness, update-cadence compliance, and role conflicts.
- Track incident count, recurrence, availability, error-budget burn, affected payment value, delayed transactions, reconciliation breaks, and SLA credits.
- Track page volume, actionability, duplicates, after-hours pages, missed acknowledgements, and load per responder.
- Track postmortem timeliness, action completion, action age, verified effectiveness, and repeated contributing factors.
- Track rotation size, duty frequency, recovery days, swaps, attrition signals, and quarterly responder sentiment.
- Hold a weekly Incident Review Board for incidents, actions, missed controls, and noisy alerts.
- Hold a monthly executive reliability review for trends, investment decisions, contractual exposure, and accepted risks.
- Hold a quarterly resilience and controls review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile dashboard data against customer cases and sampled incident records monthly to identify missing incidents or metric gaming.
19. Train participants and address pager resistance (depends on: 6, 7, 8, 13, 14, 16)
Introduce the operating model as a fair exchange, not an audit mandate. The message is that people are paid, paged only for accepted services, supported by trained command, and given capacity to fix defects.
- Train all employees to recognize impact, declare an incident, and find the incident channel and customer status page.
- Give all 260 engineers role-based training in severity, escalation, evidence preservation, handoff, and financial-integrity precautions.
- Certify incident commanders through instruction, simulation, two shadowed events or exercises, and observed performance.
- Train Communications Leads in status writing, account-manager briefing, contractual clocks, and legal escalation.
- Train scribes in timeline quality, decision capture, fact-versus-hypothesis labeling, and evidence handling.
- Require responders to demonstrate access, dashboards, rollback, runbooks, and escalation competence before primary duty.
- Require two shadow shifts before independent primary on-call.
- Appoint one adoption champion in each team and hold weekly office hours during rollout.
- Publish compensation, fatigue protections, service boundaries, and the accommodation process before assigning shifts.
- Use paid working time for training, exercises, shadowing, runbook work, and certification.
- Survey engineers at baseline, day 60, day 120, and quarterly thereafter.
20. Pilot on the payment critical path (depends on: 8, 9, 11, 12, 13, 14, 17, 19)
Run a four-to-six-week pilot across the highest-risk customer journey. Use real incidents and exercises to correct the process before wider rollout.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, authentication or API edge, settlement, and Support intake.
- Activate paid primary and secondary rotations, central command, communications, the incident record, status templates, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new system authoritative.
- Review every pilot page within one business day for routing, actionability, responder load, and missing context.
- Have the program team coach incidents without silently taking command from assigned role holders.
- Hold a weekly pilot retrospective and correct critical process or tool defects within 48 hours.
- Require 95% command assignment within five minutes, 95% timely communications, no unpaid pages, complete postmortems, and at least a 50% page-noise reduction before expansion.
21. Roll out by customer journey and risk (depends on: 18, 20)
Expand in fixed waves rather than waiting for every team to become perfect. Apply explicit readiness gates and time-limited exceptions.
- Days 0–7: operate the interim declaration, command, and communications process.
- By day 30: complete Tier 0 ownership, activate the first certified command roster, approve compensation, and begin the pilot.
- By day 60: provide compensated 24×7 coverage for every Tier 0 and Tier 1 customer journey and route all critical pages through the new platform.
- By day 90: complete critical detection upgrades, customer communications, regulatory playbooks, and the first major alert-noise reduction.
- By day 120: assign every production service a response tier, owner, tested escalation, and appropriate coverage model.
- Roll out remaining teams in two-to-three-week waves ordered by customer and dependency risk.
- Require each wave to pass ownership, alert, runbook, access, training, compensation, and tabletop gates.
- Disable legacy paging paths after verified cutover rather than leaving ambiguous parallel obligations.
- Publish a weekly adoption dashboard by team and escalate failed gates as business risks.
- Never start mandatory night coverage before compensation, staffing, training, and access are ready.
22. Exercise command, regional resilience, and ledger recovery (depends on: 9, 13, 14, 15, 19)
Validate the process under realistic conditions before depending on it during a crisis. Use the same action-tracking rules for exercises and real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops covering payment failure, customer-first detection, third-party failure, and ambiguous ownership.
- Exercise loss of one AWS region, Kubernetes degradation, PostgreSQL failure, suspected duplicate payments, queue backlog, and settlement risk.
- Exercise simultaneous operational and security events to test command and disclosure boundaries.
- Exercise loss of chat, status-page, identity, or paging providers using telephone and offline fallbacks.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors.
- Validate backups, restoration, RPO, RTO, failover prerequisites, financial controls, and post-recovery reconciliation.
- Avoid uncontrolled production-ledger experiments; use staging, replicas, simulations, or tightly governed production tests.
- Complete at least two cross-company exercises before the audit, including one overnight or unannounced paging test.
23. Test SOC 2 operating effectiveness before the auditor (depends on: 4, 18, 21, 22)
Demonstrate that the controls operate consistently, not merely that policies exist. Correct failures through tracked remediation rather than rewriting historical records.
- Preserve policy approvals, service ownership, schedules, compensation activation, access reviews, training, certifications, incidents, communications, postmortems, actions, and exercises.
- Sample evidence monthly from initial signal through verified action closure.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Use exceptions to document missed acknowledgements, late communications, incomplete records, and compensating controls.
- Conduct a formal mock audit in month 7 using the populations and interview roles expected by the external auditor.
- Trace at least one SEV0 or SEV1, one SEV2, a customer-reported event, and an exercise end to end.
- Remediate evidence and operating gaps before external fieldwork.
- Brief commanders, responders, Support, and Compliance for auditor interviews without scripting inaccurate answers.
24. Institutionalize continuous improvement (depends on: 18, 21, 23)
Keep incident management active after the audit by assigning permanent owners, budgets, and review cycles. Use incident trends to drive architectural investment.
- Assign permanent owners for the policy, service catalog, paging platform, status page, training program, metrics, and evidence repository.
- Review severity thresholds, communications timing, staffing, and compensation annually and after material process failures.
- Recertify commanders and communications leads annually through observed exercises.
- Review recurring failure families quarterly and require executive action when remediation repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fatigue, fairness, tool friction, and psychological safety.
- Report severe incidents, SLA exposure, overdue risks, resilience investments, and customer-first detection to the board or risk committee quarterly.
- Prioritize reduction of the shared-ledger concentration risk, stronger regional independence, deployment safety, graceful degradation, and automated mitigation.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
--- PROPOSAL 3 (agent qwen3.8-max_refine_3, alibaba/qwen3.8-max) ---
Estimated complexity: high
Success metrics: - Median time to detect customer-impacting incidents falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median mitigation time falls from 3h10 to under 90 minutes by day 120 and under 60 minutes by month 9.
- A named incident commander is assigned within 5 minutes in at least 95% of SEV1 and SEV2 incidents; zero incidents remain unowned for more than 10 minutes.
- First status-page update occurs within 15 minutes of SEV1 declaration and 30 minutes of SEV2 declaration in at least 95% of qualifying incidents.
- Monthly paging volume falls from 3,400 to under 500 actionable pages within six months, with alert actionability above 80%.
- Out-of-hours pages average no more than 2 per responder per week; sustained breaches trigger mandatory alert remediation.
- Six legacy alerting tools are consolidated into one paging and incident platform, with legacy paging paths disabled by week 16.
- 100% of Tier 0 and Tier 1 services have a named owning team, escalation path, dashboard, and runbook by day 60.
- All 28 teams are onboarded with readiness gates by week 22; every 24x7 critical-path rotation has at least six trained responders.
- 30 or more certified incident commanders and 20 or more certified communications leads provide continuous primary and secondary coverage.
- Paid on-call is approved by HR, legal, and finance and active in payroll before any new mandatory night rotation begins.
- 100% of required SEV1 and SEV2 postmortems are drafted within 3 business days and published within 10 business days in the standard template.
- Postmortem action closure rises from 17% to at least 90% of high-priority actions completed by their due date within two quarters.
- All 53 open historical actions are triaged within 30 days; high-risk unaccepted items are completed or formally risk-accepted within 90 days.
- Repeat incidents from a known unaddressed contributing factor decline by at least 50% within six months.
- Annualized SLA credits fall from $1.3M to under $400k within 12 months.
- Monthly availability meets or exceeds 99.95% by month 6, with exceptions reviewed at the executive reliability meeting.
- At least two cross-company exercises, including regional and ledger scenarios, are completed before the SOC 2 audit, with critical findings tracked.
- The month-6 internal dry-run audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete evidence.
- SOC 2 Type II incident-response controls pass with zero exceptions at the month-8 audit.
- On-call sentiment improves quarter over quarter; fewer than 10% of engineers report unwillingness to participate in their owned-service rotation by month 6.
- No increase in attrition among engineers on active rotations compared with baseline.
Steps (27):
1. Executive mandate, program funding, and governance
Convert the CEO email into a **written mandate within 48 hours**. Name one accountable program owner and a small decision group. Time-box design to five weeks so rollout starts well before the audit.
- Appoint the CTO as executive sponsor and a Head of Reliability or Incident Management as program owner with full-time authority.
- Create an 8–10 person design working group: SRE/platform lead, payments and ledger engineering managers, support lead, compliance, legal, HR, finance, and two rotating engineering managers.
- Approve budget lines: tooling consolidation, on-call compensation, training, coaching, and 2–3 dedicated program staff.
- State the non-negotiables: paid on-call, named service ownership, one severity scale, one paging platform, mandatory postmortems, and protected engineering capacity for reliability actions.
- Fix the master timeline: interim controls week 1, design weeks 1–5, pilot weeks 6–11, full rollout weeks 12–22, internal audit rehearsal month 6, audit-ready month 7.
- Publish a one-page charter to the company that says incident response is a company operating process, not an optional team practice.
2. Evidence baseline from incidents, alerts, and coverage gaps (depends on: 1)
Before changing anything, create an auditable **before picture** from the last 12 months. This baseline drives the severity design, staffing model, and executive reporting.
- Reconstruct all 31 customer-impacting incidents: detection source, first owner, severity, mitigation time, credits paid, and whether command was clear.
- Write a specific case review of the two incidents with unclear ownership for more than one hour.
- Inventory the six alert tools, alert volume by team, noise rate, rules with no owner, and rules with no runbook.
- Map the 16 teams without on-call and the 12 teams with unpaid on-call.
- Review the 53 open postmortem actions and triage the highest-risk items first.
- Freeze baseline metrics: 22-minute detection, 40% customer-first detection, 3h10 mitigation, $1.3M credits, 3,400 alerts, 85% noise, and 17% action closure.
3. Stakeholder listening, resistance mapping, and co-design group (depends on: 1)
Treat engineer pushback as **design input, not an attitude problem**. The objection to carrying a pager for other teams' code must shape ownership, routing, compensation, and command staffing.
- Interview leads from all 28 teams plus support, customer success, sales, compliance, and HR within two weeks.
- Separate the real causes of resistance: unpaid night work, unfamiliar systems, poor runbooks, unfair load, fear of blame, or unclear authority.
- Recruit 10–15 respected engineers and managers as co-designers so the model is built with the teams.
- Survey baseline on-call sentiment, trust in alerts, and psychological safety; repeat at 90, 180, and 365 days.
- Document the central promise: responders are paged for services they own, trained incident commanders coordinate, and all on-call work is compensated.
4. Service ownership catalog, criticality tiers, and dependency map (depends on: 2)
No alert can page the right team until every service has a **named owner**. Build the catalog as the routing foundation for on-call, severity impact mapping, status-page components, and audit evidence.
- Assign one accountable team to each of the 180 services, with an engineering manager, Slack channel, escalation policy, dashboard, and runbook link.
- Tier services: Tier 0 for money movement, ledger integrity, authentication, settlement, and shared PostgreSQL; Tier 1 for customer-facing degradable services; Tier 2 for internal or batch services; Tier 3 for non-critical services.
- Map customer journeys to services, databases, regions, third parties, and contractual SLA components.
- Mark orphan and shared services; require ownership, reassignment, or decommissioning within 30 days.
- Document cross-region dependencies, failover constraints, and services that defeat nominal regional redundancy.
- Treat missing ownership for Tier 0 or Tier 1 services as an executive escalation and a release-blocking risk.
5. Severity scale, declaration rights, and automatic triggers (depends on: 2, 4)
Adopt one severity scale so declaration is a lookup, not a debate. **Anyone may declare; only the incident commander may downgrade.** When uncertain, start higher.
- SEV1: money movement stopped or incorrect, ledger integrity in doubt, confirmed security or data event, both regions impaired, or broad SLA-credit exposure.
- SEV2: major degradation, settlement window at risk, one or more critical customers fully down, or likely contractual breach.
- SEV3: limited impact with a workaround, no financial-integrity or security risk.
- SEV4: no customer impact; handled by ticket during business hours.
- Define fixed triggers for each level: pages, command staffing, bridge call, status-page timing, account-manager outreach, executive notification, and postmortem obligation.
- Add automatic escalation: SEV3 open more than 2 hours becomes SEV2; any incident touching the shared ledger or unclear ownership becomes SEV2 unless the commander documents otherwise.
- Publish a decision tree and 10–12 worked examples from the actual 31 incidents.
6. Incident roles, authority, and handover rules (depends on: 5)
Solve the **nobody-was-in-charge failure** by making command explicit, trained, and transferable. Separate coordination from technical remediation so responders are not asked to debug unfamiliar code.
- Incident Commander: owns severity, priorities, escalation, mitigation strategy, role assignment, handoffs, and closure. Does not write code during the incident. May freeze deploys, invoke pre-approved failover, and pull any on-call responder.
- Communications Lead: owns status page, internal updates, account-manager briefs, and coordination with legal or compliance.
- Scribe: maintains the timestamped record of decisions, actions, role changes, and customer communications.
- Subject-matter responders: engineers from owning teams who diagnose and remediate within their accepted domain.
- Executive liaison: required for SEV1; shields the commander from executive questions and owns board or regulator escalation.
- Require the commander to claim the role within 5 minutes of SEV1 or SEV2 declaration, state it in the incident channel, and record every handover.
- Allow role combination only for SEV3 or SEV4; require separate people for command, communications, and primary technical work at SEV1.
7. Paid on-call, fatigue safeguards, and HR compliance (depends on: 3)
Unpaid on-call is a retention, fairness, and legal risk in New York. **Pay must be approved before any new mandatory rotation starts.** This is the fastest way to reduce pager resistance.
- Create a weekly stipend for primary and secondary on-call, differentiated by tier and night responsibility.
- Add per-page or per-incident pay for out-of-hours activation, plus guaranteed recovery time after night work or major incidents.
- Create a separate stipend for duty incident commanders and communications leads because command is a heavier burden.
- Verify FLSA, New York wage-hour, overtime, holiday, and payroll treatment with HR, legal, and finance.
- Cap rotation frequency: no more than one primary week in four or one week in six depending on staffing; no simultaneous primary assignments.
- Require at least six trained responders for any 24x7 rotation; fund hiring or service reassignment where teams are too small.
- Publish the compensation package and payroll start date before asking engineers to join rotations.
8. On-call staffing model and rotation rules (depends on: 4, 6, 7)
Do not force 28 identical night rotations. Use a **layered model**: central command coverage, critical-path team coverage, and business-hours coverage for lower-tier services.
- Company incident-command and communications rotation: 24x7 pool of 25–35 trained volunteers and designated senior staff, giving primary plus secondary coverage at all times.
- Critical-path teams: 24x7 primary and secondary on-call for Tier 0 and Tier 1 domains, including payments, ledger, authentication, API edge, Kubernetes platform, and shared PostgreSQL.
- Other teams: business-hours on-call with a documented night escalation list owned by the engineering manager.
- Group services into 8–12 coherent domains so rotations are sustainable; no rotation with fewer than six people is approved without an executive exception.
- Require handoff overlap, shadow shifts before first primary duty, and no on-call during approved PTO.
- Define cross-team pull rules: the commander may page another team's on-call with a 10-minute acknowledgement obligation; this is coordinated by command, not pushed onto responders.
- Publish schedules, swap rules, and load limits in the paging tool.
9. Alert quality standard, page budget, and noise controls (depends on: 2, 4)
With 3,400 alerts and 85% noise, detection fails because people stop trusting pages. Make alert quality a **condition of paging anyone**.
- Every page must have a named owning team, customer or SLO impact, severity, dashboard, runbook, expected action, and escalation policy.
- Page on customer-visible symptoms: payment success rate, latency, settlement deadlines, error-budget burn, ledger integrity, and replication health.
- Demote cause-based CPU, memory, or infrastructure-only alerts to dashboards or tickets unless they map to a customer journey.
- Set a page budget: maximum 2 out-of-hours pages per person per week; breach triggers a mandatory alert-tuning sprint.
- Auto-flag alerts that fire repeatedly without action or have high no-action acknowledgement rates.
- Require shadow mode for new alerts before they page humans, except for documented emergencies.
- Review alert quality monthly by team and publish a noise leaderboard.
- Target fewer than 500 actionable pages per month and above 80% actionability within six months.
10. Detection uplift across payments, ledger, and customer signals (depends on: 4, 9)
Customers detected 40% of incidents first. Detection must shift to **payment outcomes, ledger integrity, and inbound customer signals**, not host metrics alone.
- Define SLIs and SLOs for payment initiation, authorization, settlement, reconciliation, refunds, API availability, and reporting freshness.
- Run synthetic end-to-end payment tests from outside the platform in both AWS regions every 60 seconds.
- Add continuous ledger assurance: double-entry reconciliation, replication lag, failover readiness, disk pressure, and settlement-window countdown alerts.
- Monitor per-customer anomalies for top accounts so a single-tenant outage is detected before the account manager calls.
- Convert high-priority support tickets and account-manager reports into incident candidates within 5 minutes.
- Track detection source for every incident; make customer-first detection a reviewed defect with a corrective action.
- Add partner, banking, and card-network notification intake into the same declaration path.
11. Single incident platform and alert-tool consolidation (depends on: 5, 8, 9)
Collapse six tools into **one paging and incident-management system** with one queue, one timeline, and one audit trail. Do not create a parallel audit process.
- Select an integrated stack: paging and on-call scheduling, incident workflow, Slack or chat integration, bridge calling, status-page API, and ticketing integration.
- Implement one-command declaration that creates the incident record, channel, bridge, severity label, role prompts, and clock.
- Migrate alert sources by wave; retire a legacy paging path only after routing, ownership, and acknowledgement tests pass.
- Automate evidence capture: timestamps, acknowledgements, role assignments, severity changes, communications, and postmortem links.
- Integrate with the service catalog, on-call schedules, Jira or equivalent action tracker, customer-success tooling, and the status page.
- Verify the platform works during a single-region failure, including SMS, phone, and offline fallback paths.
- Set a hard date after which pages outside the chosen platform are not valid on-call obligations.
12. Escalation paths and five-minute command rule (depends on: 6, 8, 11)
Create one path from **signal to named commander in under five minutes**, any hour of the day. Make escalation automatic and time-bound.
- Accept declarations from automated alerts, engineers, support, account managers, partners, and customers through the same command.
- For SEV1 and SEV2, page the duty incident commander and owning-team primary immediately.
- Acknowledgement ladder: primary 5 minutes, secondary 10 minutes, manager 15 minutes, director or executive 20 minutes.
- If no commander claims the incident within 5 minutes, the platform assigns and announces one; the assignee may hand over but cannot leave the incident unowned.
- For SEV3, require team acknowledgement within 30 minutes; otherwise create a tracked work item.
- Give the commander pre-approved authority to invoke regional failover, ledger read-only mode, feature kill switches, and partner notifications without waiting for executive sign-off.
- Record every missed acknowledgement and escalation failure for weekly review.
13. Internal communications protocol (depends on: 6, 11)
Separate the **working incident channel from the audience channel** so responders can work and executives, support, and sales stay informed without interrupting the commander.
- Create one incident channel and one bridge per SEV1–SEV3 incident; use a read-only broadcast channel for executives and support.
- Set update cadence: every 15 minutes for SEV1, 30 minutes for SEV2, and at state changes for SEV3.
- Use a fixed update template: impact, customer-visible symptoms, current action, ETA or next update, commander, and communications lead.
- Brief support and customer success with a live affected-customer list and approved holding statements within 15 minutes of SEV1 or SEV2.
- Require executives to route questions through the executive liaison; the commander is not interrupted.
- Define fatigue rules and formal handover for incidents lasting more than 4 hours.
14. Customer status page, account-manager outreach, and SLA credit workflow (depends on: 5, 13)
Replace ad-hoc status updates with a **timed, owned, template-driven customer communication process**. Link incidents to SLA credits so finance and customer success are not surprised.
- Publish status-page updates within 15 minutes of SEV1 declaration and 30 minutes for SEV2; update every 30 or 60 minutes until resolution.
- Assign the communications lead as the single author; use pre-approved templates reviewed by legal and communications.
- Map status-page components to customer capabilities: payments, settlement, reporting, API, onboarding, and regional availability.
- For top accounts, require direct account-manager outreach within 30 minutes of SEV1 with an approved briefing pack.
- Send proactive email or webhook notifications to subscribed customers for SEV1 and SEV2.
- Publish a resolution notice within 30 minutes of mitigation and a customer-facing incident summary within 5 business days for SEV1.
- Automate SLA credit calculation from incident duration and affected capability; review with finance and legal within 5 business days.
- Track credits by incident and root-cause family to guide reliability investment.
15. Regulatory, legal, and partner notification playbook (depends on: 5, 14)
In payments, some incidents are reportable and the clock starts at detection. Make regulatory assessment a **mandatory step in the incident process**, not an afterthought.
- Map obligations: NYDFS cybersecurity-event rules, state breach laws, GLBA safeguards, PCI DSS if in scope, money-transmitter duties, card-network and sponsor-bank contracts, and cyber-insurer notice.
- Require a legal or compliance reportability assessment within 2 hours for every SEV1 and every security-related SEV2, even when the answer is not reportable.
- Maintain a 24x7 contact matrix for regulators, sponsor banks, card networks, outside counsel, insurers, and law enforcement.
- Pre-draft notification templates and preserve legal review.
- Encode customer-contract notification deadlines into account tiers and the communications workflow.
- Record every reportability decision, approver, deadline, and submission confirmation in the incident record.
16. Postmortem standard and blameless review (depends on: 5, 6)
Replace inconsistent postmortems with a **mandatory, blameless, single-format process**. The discipline comes from deadlines, facilitation, and action tracking.
- Require postmortems for all SEV1 and SEV2 incidents, customer-detected incidents, incidents over 2 hours, repeat failures, and ledger near-misses.
- Draft within 3 business days, peer review within 5, publish within 10 for SEV1 and SEV2.
- Use one template: summary, impact, timeline, detection analysis, response analysis, contributing factors, what worked, what failed, and action items.
- Make blamelessness explicit: focus on systems and decisions, not individual fault; never use postmortems in performance discipline.
- Hold a weekly incident review board to review postmortems, ratify severity, and challenge weak actions.
- Maintain a searchable postmortem library and quarterly recurring-cause analysis.
- Require a trained facilitator for major reviews; the incident commander attends but does not facilitate.
17. Action-item tracking, ownership, and delivery gates (depends on: 16, 11)
Only 11 of 64 actions were closed. Give postmortem actions the same status as **customer commitments**, with named owners and visible escalation.
- Create every action as a ticket with one named individual owner, priority, due date, and verification method.
- Use delivery classes: containment within 7 days, corrective work within 30 days, strategic work within 90 days.
- Reserve 15–20% of team sprint capacity for reliability and incident actions.
- Escalate overdue items: manager at 7 days, director at 14 days, CTO dashboard at 30 days.
- Block related feature releases when overdue P0 actions prevent recurrence of a severe incident.
- Require director approval and documented residual risk acceptance for overdue high-risk items.
- Verify effectiveness after completion; closing a ticket without evidence does not close the action.
- Target 90% of high-priority actions completed on time within two quarters.
18. Runbooks, critical-incident playbooks, and readiness bar (depends on: 4, 8)
Poor runbooks are a real cause of pager resistance and slow mitigation. Define a **minimum readiness bar** before a service is allowed to page anyone at night.
- Require for every Tier 0 and Tier 1 service: architecture summary, dependencies, dashboards, alert-to-runbook map, rollback procedure, feature flags, escalation contacts, and customer-impact statement.
- Write major playbooks for shared PostgreSQL ledger failure, regional failover, Kubernetes control-plane loss, payment-processor outage, settlement-window breach, duplicate-payment suspicion, and security compromise.
- Define ledger recovery rules: failover procedure, read-only degraded mode, reconciliation, RPO/RTO, and data-loss tolerance approved by executives.
- Test runbooks in drills at least twice a year; mark untested runbooks stale.
- Prevent paging alerts for services without readiness sign-off unless the engineering manager accepts the gap in writing.
- Keep runbooks linked from every alert and incident template.
19. Training, certification, and role readiness (depends on: 6, 12, 13, 16)
Command and communications are skills. Build a **tiered certification path** so rotations are staffed by people who have practiced, not by whoever is around.
- All employees: 1-hour module on declaring incidents, finding the incident channel, and reading the status page.
- All responders: half-day training on severity, acknowledgement, escalation, runbooks, and evidence hygiene.
- Incident commanders: 2-day course plus two shadowed incidents and one simulation before certification.
- Communications leads: training on status-page writing, customer language, account-manager briefs, and regulatory triggers.
- Scribes: training on timeline discipline and audit evidence.
- Certify for 12 months; renew through a simulation.
- Require shadow shifts before independent primary duty; no new hire holds primary within 90 days.
- Publish the certification register as an audit artifact.
20. Simulation program and game days (depends on: 19, 11, 18)
Rehearse the process before it meets a real SEV1. Simulations build commander confidence, expose runbook gaps, and produce audit evidence.
- Run monthly 60-minute tabletops using real incidents from the 31-incident baseline.
- Run quarterly game days covering regional failover, ledger replica promotion, dependency failure, partner outage, and security event.
- Run twice-yearly unannounced paging drills to measure night acknowledgement times.
- Include support, account managers, legal, compliance, and executives in at least one exercise per quarter.
- Produce tracked action items from every exercise using the same board as real incidents.
- Measure time to commander, time to first status update, and time to mitigation decision.
21. Critical-path pilot and gate review (depends on: 7, 10, 11, 14, 18, 19)
Prove the model on the highest-risk services with willing teams before full rollout. Run a **six-week pilot with daily feedback and public exit criteria**.
- Pilot with payments, ledger/database, platform/Kubernetes, API edge, authentication, and support intake.
- Activate severity scale, duty commanders, paid rotations, single paging platform, alert budget, status-page policy, and postmortem process.
- Hold a weekly pilot retrospective and fix process defects quickly.
- Validate night acknowledgement, cross-team pull response, severity clarity, and compensation payroll.
- Exit gate: commander assigned within 5 minutes in 95% of incidents, status page on time, page noise down at least 50%, postmortems on time, and positive on-call sentiment.
- Publish pilot results to the whole company as the main adoption argument.
22. Wave rollout across all 28 teams (depends on: 21)
Roll out by criticality and dependency, not by calendar alone. Use **readiness gates** so teams are not forced live without coverage.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1 teams. Wave 3: Tier 2 teams. Wave 4: Tier 3 and internal platforms.
- Gate per team: catalog entry complete, alerts migrated, runbooks ready, six-person rotation staffed, one commander candidate nominated, compensation in payroll, and one tabletop passed.
- Assign a program coach to each wave for three weeks.
- Freeze legacy paging paths for each team after successful onboarding.
- Publish a live adoption scoreboard by team.
- Complete all teams by week 22, leaving months of operating evidence before the audit.
23. Metrics, dashboards, and review cadence (depends on: 11, 16, 21)
Instrument the process itself. Leadership must see whether the program is working, and auditors must see **operating evidence, not retrospective paperwork**.
- Track response metrics: MTTD, time to declare, time to commander, acknowledgement time, mitigation time, resolution time, and customer-first detection rate.
- Track quality metrics: page volume, alert actionability, missed pages, postmortem timeliness, action closure rate, and status-page compliance.
- Track business metrics: availability against 99.95%, SLA credits, repeat incidents, and error-budget consumption.
- Track people metrics: on-call load, night pages per person, recovery time usage, sentiment, and attrition signals.
- Hold weekly incident review, monthly reliability review, quarterly executive review, and annual policy review.
- Publish dashboards internally and give every metric a target and owner.
24. Change management, incentives, and culture (depends on: 3, 7, 21)
The process will be judged on fairness. Communicate the deal repeatedly and make participation **recognized, compensated, and safe**.
- Core message: you are paid, paged only for what you own, supported by a trained commander, and given real capacity for actions.
- Run CTO all-hands, team roadshows, office hours, FAQ, and an internal incident-management hub.
- Add incident response and reliability work to promotion criteria and manager objectives.
- Recognize good postmortems, alert-noise reduction, and calm incident leadership.
- Provide a written path for engineers who cannot do nights; cover those shifts with paid volunteers or adjusted staffing.
- Prohibit retaliation for good-faith declaration or escalation.
- Publish sentiment survey results, including bad news, to maintain credibility.
25. SOC 2 evidence design and internal dry-run audit (depends on: 16, 17, 22, 23)
Design the process so audit evidence is a **by-product of normal operations**. Test it internally before the external auditor does.
- Map the process to SOC 2 criteria: incident identification, response, recovery, monitoring, communication, control activities, availability, and corrective action.
- Approve versioned policies: incident response policy, severity standard, on-call policy, communications policy, postmortem standard, and exception process.
- Retain incident records, paging logs, status-page history, postmortems, action tickets, training records, drill records, and access reviews for the audit period.
- Record regulatory reportability decisions even when no notification is required.
- Run an internal dry-run audit at month 6: sample at least 15 incidents and trace evidence end to end.
- Fix gaps with at least four weeks before the audit.
- Maintain an exception and remediation log instead of claiming perfection.
26. Program risk register and contingency planning (depends on: 1)
Name the likely failure modes now and pre-commit responses. Review the register monthly with the executive sponsor.
- Commander volunteer shortfall: roster command duty for senior engineers and managers until the pool reaches 25 trained people.
- Compensation delay: use immediate time-off-in-lieu plus phased stipend, but do not start mandatory night rotations without pay.
- Tool migration slip: preserve single paging and routing scope; defer advanced automation if needed.
- Alert pruning causing missed incidents: move noisy alerts to ticket first, observe 30 days, then delete.
- Burnout in experienced teams: monitor load weekly and cap pages per person.
- Major SEV1 during rollout: program owner shifts to incident support, wave schedule slips one wave, sponsor is informed same day.
- Shared ledger concentration: track blast-radius reduction and failover improvements as top strategic actions.
27. Continuous improvement, maturity roadmap, and post-audit sustainability (depends on: 23, 25)
Prevent the classic failure where the process decays after the audit. Build the second-year operating plan before the first year ends.
- Hold quarterly process retrospectives with commanders, responders, support, and communications leads.
- Re-baseline targets every six months; tighten goals once current targets are met.
- Move from lagging metrics to leading indicators: error-budget burn, near-miss rate, drill performance, and action ageing.
- Plan year-two improvements: follow-the-sun coverage, automated mitigation, ledger blast-radius reduction, error-budget release gates, and per-customer real-time impact reporting.
- Keep annual policy review, certification renewal, drills, and board reporting on a permanent calendar independent of the audit cycle.
- Report quarterly to the board or risk committee on availability, credits, severe incidents, overdue actions, and resilience investment.
--- PROPOSAL 4 (agent grok4.6_refine_4, xai/grok-4.6) ---
Estimated complexity: high
Success metrics: - Median time to detect falls from 22 minutes to under 10 minutes by month 3 and under 5 minutes by month 9.
- Customer-first detection falls from 40% to under 20% by month 3 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by month 4 and under 60 minutes by month 12.
- A named Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes after week 4.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases by month 3, with required update cadence met in 95% of cases.
- Monthly pages fall from 3,400 to under 500 within 6 months, with actionability above 75%; out-of-hours pages at or under 2 per person per week.
- All six legacy alerting tools route through one paging platform by week 16; legacy paging paths disabled per wave.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by due date within two quarters; all 53 currently open historical actions triaged within 30 days of S2.
- SLA credits fall from $1.3M to under $400k in the first 12 months; customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known unaddressed cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar; every Layer B rotation has 6+ certified responders or a time-limited executive exception.
- Paid on-call policy approved by HR, Legal, and Finance and in payroll before any mandatory night rotation starts.
- All 28 teams onboarded by week 24 with 24x7 command cover from 25+ certified ICs and 15+ certified Comms Leads.
- Internal dry-run at month 6 passes a 15-incident evidence walkthrough; month-7 mock audit finds no unowned high-risk control gap; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment improves quarter over quarter; no increase in attrition among engineers on rotation; pulse at 60 and 120 days shows at least 70% agree rotations are fair and limited to services they own.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked actions; two cross-company exercises completed before the audit.
Steps (24):
1. Program charter, mandate, and funding
Convert the CEO email into a named program with one owner, a budget, and a deadline earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, 28 teams will opt out.
- Appoint a Program Lead (Head of Reliability) with direct CTO and CEO sponsorship.
- Form a small steering group: CTO, VP Eng, Head of CS/Support, CISO, Legal, Finance, HR. Not a 28-team committee.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Timeline: operating floor in 7 days; design weeks 1–6; pilot weeks 7–12; rollout weeks 13–24; mock audit week 28; SOC 2 at month 8.
- Fund tooling ($150–250k/yr), on-call pay (~$600k–1M/yr), 2–3 program FTEs, and reserved engineering capacity. Anchor the ask against $1.3M in credits plus unmeasured incident cost.
- Freeze baselines now: 31 incidents, MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, $1.3M credits, 3,400 alerts/month at 85% noise, 11 of 64 actions closed.
2. Immediate 7-day operating floor (depends on: 1)
Do not wait for tooling, compensation, or the audit. Put a minimum process in place this week so the next outage has a named commander.
Start retaining artifacts on day one. This week becomes the first audit evidence.
- Publish a one-page interim severity guide and a single declaration path through Slack, phone, and pager.
- Staff interim primary and backup Incident Commander 24x7 from existing on-call veterans and engineering managers. Compensate this duty retroactively.
- Require a named IC within 10 minutes for every suspected major incident. If nobody claims it, the duty manager is IC.
- Use one channel, one bridge, one timeline doc, and one naming convention for every major incident.
- Direct Support to escalate credible customer reports immediately. They do not wait for engineering confirmation.
- Triage all 53 open historical actions. Close, re-plan, or formally risk-accept. Ledger integrity, payment duplication, regional resilience, security, and detection come first.
- Hold a daily 15-minute ops review until the permanent process is live.
3. Forensic baseline of incidents and alerts (depends on: 1)
Rebuild the facts before locking design. This is the before picture for the CEO and the auditor.
Every later design choice should trace to this evidence.
- Re-code each of the 31 incidents: trigger, service, detection source, timestamps, who led, credits paid, root-cause family.
- Quantify the 40% customer-first detections and name the missing signal in each case.
- Reconstruct the two nobody-in-charge incidents minute by minute. Use them as the burning-platform story.
- Audit the six alerting tools: volume per tool and team, top 50 noisy rules, rules with no owner or runbook.
- Freeze the baseline numbers. Do not let them drift during design.
4. Listening tour and the fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the main delivery risk. Treat it as a design constraint, not an attitude problem.
The answer that will actually stick is **the deal**: you are paid; you are paged only for services you own; a trained commander runs the room; your postmortem actions get real sprint capacity.
- Interview all 28 teams plus Support, CS, and Sales in two weeks.
- Separate the objections: unpaid work, nights, unfamiliar code, bad runbooks, fear of blame. Each needs a different fix.
- Collect informal practices from the 12 teams already on-call. They are the pilot candidates and the veteran pool.
- Recruit 10–15 credible engineers as a design working group so the process is co-authored.
- Baseline sentiment on alert trust, on-call willingness, and burnout. Re-measure at 6 and 12 months.
5. Service ownership catalog and criticality tiers (depends on: 3)
You cannot page the right person across 180 services until each one has a named owner. This is the foundation of fairness, routing, and audit evidence.
Build a machine-readable catalog as the single source of truth.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, the shared PostgreSQL cluster, queues, partner banks, and customer-facing endpoints.
- Assign one owning team, named engineering manager, Slack channel, escalation policy, and dependency list per service.
- Tier 0: money movement, ledger, auth, shared Postgres, regional control plane. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map Tier 0/1 services to customer capabilities: initiation, authorization, settlement, reporting, onboarding.
- Assign coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
6. Severity scale and declaration rules (depends on: 3, 5)
Replace in-the-moment debate with a lookup. Four payments-specific levels, with worked examples drawn from the 31 real incidents.
Anyone may declare. Only the Incident Commander may downgrade, with recorded rationale. When unsure, start high.
- **SEV1**: money movement stopped or incorrect; ledger integrity in doubt; material security or data exposure; both regions impaired; or more than 10% of customers impacted. Pages IC, comms, scribe, SMEs, and exec liaison. Bridge in 5 minutes. Status page in 15 minutes. Mandatory postmortem and regulator assessment.
- SEV2: severe degradation; settlement window at risk; a strategic customer fully down; SLA breach likely. IC and SMEs paged. Status page in 30 minutes. Mandatory postmortem.
- SEV3: partial impact with a workaround; no credit exposure. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, or a repeat.
- SEV4: internal or minor. Ticket only. No page.
- Auto-escalate: any SEV3 open more than 2 hours, or any incident touching the shared ledger, becomes SEV2. Unknown impact after 15 minutes is raised, not sat on.
- Nobody is punished for over-declaring. Publish that rule in writing and repeat it.
7. Roles, authority, and ledger dual-control (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, transferable, and logged.
The Incident Commander owns the incident, not the fix, and **never types** in a production terminal.
- Incident Commander: declares severity, pulls anyone, freezes deploys, invokes failover, authorises spend. One IC at a time. Assumed within 5 minutes and announced in channel.
- Communications Lead: single voice to customers, status page, account managers, and the exec summary.
- Scribe: timestamped timeline, decisions, and open questions. Feeds the postmortem and the audit trail. Required for SEV1 and SEV2.
- Subject-matter responders: diagnose and mitigate only services they own, with access and runbooks.
- Executive Liaison (SEV1): shields the IC from exec questions; owns regulator and board escalation.
- The IC may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access rules.
- Handover is verbal and written, with time and new owner recorded. Roles may combine below SEV2; never at SEV1. The IC stays in charge if a VP joins.
8. Three-layer 24x7 coverage model (depends on: 5, 7)
Do not put 28 teams on night rotation. That is what engineers are rejecting.
Staff command centrally. Page engineers only for code their team owns.
- Layer A, company command: Duty IC plus secondary, Duty Comms, and a scribe pool. Target 25–35 certified ICs and 15–20 comms leads. Roughly one week per person every 6–8 months. Five-minute ack SLA, then secondary, then on-call director.
- Layer B, Tier 0/1 domains: group 28 teams into 8–12 product and platform domains (ledger app, Postgres platform, payments orchestration, auth, API edge, Kubernetes/platform, settlement). Primary plus secondary 24x7. Minimum six trained people. Target one week in six, never worse than one in four.
- Layer C, Tier 2/3: business-hours on-call. After hours the IC pages the EM, who holds a written escalation list.
- No engineer joins another team's responder pool without training, access, runbooks, shadow shifts, and both teams' acceptance.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner.
- Evaluate follow-the-sun coverage as a 12-month option, not a year-1 dependency.
9. Paid on-call, NY labor compliance, and fatigue rules (depends on: 4, 8)
Unpaid on-call is a retention risk and a New York legal exposure. Pay must be live in payroll before any mandatory night rotation starts.
Treat interrupted nights and recovery time as compensable work.
- Weekly stipend by layer, higher for 24x7 Tier 0/1 and Duty IC, lower for business-hours, with holiday and weekend premiums.
- After-hours call-out pay or TOIL. A paid recovery day after prolonged overnight work, a SEV1, or a qualifying SEV2. Managers cover the next day.
- HR, Legal, and Finance publish dollar amounts, FLSA exempt/non-exempt treatment, NY wage-hour rules, tax treatment, and payroll timing within 14 days of charter.
- Load rules: no primary on two rotations; no consecutive primary weeks; no on-call the week after a SEV1 you commanded.
- A person may declare temporarily unfit after overnight work with no performance penalty.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- If a rotation cannot staff six people, merge domains or hire. Do not run two-person 24x7.
- Model annual cost against $1.3M in credits and get it as a CFO/board line item.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts a month at 85% noise is why detection takes 22 minutes. Make quality a condition of paging a human.
A page is a product with a quality bar, not a dump of host metrics.
- Every paging alert must have a named owner, customer or SLO impact, runbook, tested threshold, severity mapping, dashboard, expected action, and dedup key. Fail any of these and it becomes a ticket or is deleted.
- Page on customer symptoms: SLO burn, error budget, settlement-queue age. Cause-based CPU and memory alerts become dashboards or tickets.
- Set a **page budget** of at most two out-of-hours pages per person per week. A breach triggers a mandatory tuning sprint and blocks new paging alerts for that team.
- Auto-quarantine alerts that fire more than five times a month with no action, or that have more than 70% no-action acks. Never silently disable without compensating detection and a recorded decision.
- Run new alerts in shadow for seven days unless an emergency exception is approved.
- Monthly per-team kill, tune, or keep review. Target under 500 pages a month and actionability above 75% within six months.
11. Customer-journey detection and ledger assurance (depends on: 5, 10)
Stop customers being the monitoring system. Detect on money-path outcomes, not host metrics.
Every postmortem will ask why a customer saw it first.
- Define SLOs per Tier 0/1 capability: initiation success, auth latency, settlement timeliness, API availability, reporting freshness. Set an internal target stricter than 99.95%.
- Run synthetic full-lifecycle payments from outside the platform, in both regions, every 60 seconds. Include a small real-value canary where legally feasible.
- Add ledger assurance: continuous double-entry reconciliation, replication lag, failover-readiness, unexpected balances, duplicate identifiers, settlement-window countdown.
- Add per-customer anomaly detection for the top 100 accounts.
- Auto-create a triage incident within five minutes when support tickets or AM reports match impact keywords.
12. Single incident and paging platform (depends on: 6, 8, 10)
Collapse six alerting tools into one queue, one timeline, and one audit record.
The platform must still work if one AWS region is down. Verify SMS and phone fallback, plus an offline runbook.
- Select one paging product, one incident layer, and one hosted status page.
- One Slack command declares the incident, creates the channel and bridge, pages the Duty IC, sets severity, and starts the clock.
- Ingest the six existing tools first. Deduplicate and route. Retire a legacy path only after named owners and two weeks of verified operation.
- Auto-capture timestamps, acknowledgements, role assignments, severity changes, comms, mitigated time, and resolved time. Export for SOC 2.
- Integrate the service catalog, Jira, Salesforce or CS tooling for affected-customer lists, and Zoom or Slack Huddle.
- Test paging, escalation, status publication, and conference access every week.
- After a team's wave, old paging paths are disabled, not left as a fallback.
13. Detection-to-command escalation path (depends on: 7, 8, 12)
Write the single path from something looks wrong to someone is in charge. Target: a commander in under five minutes, any hour.
If nobody claims IC in five minutes, the platform assigns the Duty IC. The assignee can hand over, not decline.
- Converge every entry point on the same declare command: alert, engineer, support, account manager, partner bank, SEV hotline.
- SEV1: page the owning primary immediately; secondary at 5 minutes unacked; domain manager and Duty IC at 10; exec liaison at 15.
- SEV2: primary ack in 10 minutes; IC assigned in 15.
- Human acknowledgement is required. Delivery to a device does not count.
- The IC can page any team's on-call, with a 10-minute ack obligation. This reciprocity makes single-team ownership viable.
- Unowned alerts go to Layer A command, then the missing owner record is a control defect.
- Pre-authorise regional failover, ledger read-only mode, and partner-bank notice so the IC does not wait for an executive. Dual-control still applies to ledger writes.
14. Live execution and major-incident playbooks (depends on: 7, 13)
Limit customer and financial harm before proving root cause. One procedure from the first minute to handback.
A service cannot page at night until it meets the readiness bar.
- Open channel, bridge, record, and timeline immediately for SEV1 and SEV2. The IC states severity, known impact, hypothesis, objective, roles, and next update time.
- Freeze unrelated production changes during SEV1. Record exceptions the IC approves.
- Prefer reversible mitigation: rollback, feature flag, traffic isolation, rate limit, partner reroute.
- Write playbooks first for Postgres ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank outage, settlement-window breach, suspected security compromise, and suspected duplicate payments.
- Guard against split-brain, replay, and duplication during regional or database recovery. Reconcile and process backlog before calling the incident resolved.
- Mitigation means customer impact ended. Resolution means stable, backlog processed, and ledger reconciled.
- Formal IC handover after 4 hours. Reopen if impact recurs during the stability window.
- Readiness bar per Tier 0/1 service: diagram, dependencies, dashboard, runbook, rollback, kill switch, escalation contacts, RPO/RTO. Untested runbooks are marked stale.
15. Internal, customer, and regulatory communications (depends on: 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved process. The Communications Lead is the single author.
State impact and the next update time. Never speculate on cause.
- Internal: one working channel, one bridge, one read-only broadcast for execs, Support, and Sales. SEV1 update every 30 minutes even if unchanged; SEV2 every 60 minutes.
- Executives ask questions only of the Executive Liaison. Publish this as a signed exec behaviour rule.
- Status page: SEV1 within 15 minutes, SEV2 within 30 minutes; then 30/60-minute updates; resolve notice within 30 minutes of mitigation. Templates pre-approved with Legal.
- Top 100 accounts: named AM contact within 30 minutes of SEV1, with a briefing pack from Comms. Long tail gets the status page plus email or webhook.
- Customer-facing summary within 5 business days for SEV1.
- Mandatory regulatory checkpoint on every SEV1 and every security SEV2, within 2 hours, recorded even when not reportable. Cover NYDFS Part 500 72-hour clock, state breach laws, GLBA/FTC, PCI if in scope, sponsor-bank and card-network windows, FinCEN/OFAC if relevant.
- Legal owns outbound regulatory letters. The IC owns facts. Security or Legal may limit public detail during an active threat, with the reason recorded.
- Encode bespoke customer-contract notice SLAs into account tiering.
16. SLA credit and financial-impact workflow (depends on: 6, 15)
Link incidents to money so severity, credits, and investment stay consistent. Finance should not learn about outages from invoices.
Make credit calculation an output of the incident record, not a negotiation.
- Agree availability measurement per contract and component with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and capability telemetry. Produce a proposed credit schedule within 5 business days of resolution.
- Posture: proactive credits for top-tier accounts; claims-based for the rest. Document the approval chain.
- Track credits per incident and root-cause family. Quarterly report which reliability investments would have prevented which credits.
- Target: cut credits from $1.3M to under $400k in 12 months. Use that delta as the ongoing business case.
17. Blameless postmortems and action tracking (depends on: 6, 12)
Eleven of 64 actions closed is the clearest process failure. Make learning mandatory and actions as binding as customer commitments.
Closing a ticket without evidence of effectiveness does not close the action.
- Mandatory for every SEV1 and SEV2, every customer-first detection, every incident over 2 hours, every repeat of a known cause, and every ledger near-miss.
- Draft within 3 business days, review within 5, published internally within 10. The IC owns the draft. The owning EM is accountable.
- One template: timeline, customer and financial impact, why detection was late, why mitigation took that long, contributing factors, what went well, actions.
- Blameless in writing. No individual named as a cause. HR and management commit that postmortems are never used in performance reviews.
- Every action gets a named person, priority, due date, Jira ticket, and verification method. P0 (prevents SEV1 recurrence) due in 30 days and committed into the next sprint before roadmap work. P1 in 60 days. P2 in 90 days.
- Teams reserve 15–20% of sprint capacity for reliability and incident actions.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Weekly Incident Review Board with engineering directors. Target 90% of P0/P1 actions closed on time within two quarters.
18. Training, certification, and commander academy (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Nobody holds independent duty untrained. Training happens in paid working time.
The certification register is an audit artefact.
- All employees, 1 hour: recognise impact, declare, find the channel and status page.
- Responder, half day: severity, escalation, runbooks, comms hygiene. Mandatory before joining a rotation.
- Scribe, 2 hours: timeline discipline. The entry point.
- Incident Commander, two days plus shadowing: command presence, decisions under uncertainty, severity, handover, exec management. Two shadowed incidents and one simulated SEV1 before certification.
- Communications Lead, one day: status writing, customer tiering, legal boundaries, regulator triggers.
- Certification valid 12 months, renewed via simulation.
- New joiners shadow two shifts and do not hold primary in their first 90 days.
- Appoint a champion in each of the 28 teams.
19. Simulations and game days (depends on: 14, 15, 18)
Rehearse before the next real SEV1. Use a standing calendar, not a one-off exercise.
Do not inject uncontrolled changes into the production ledger.
- Monthly 60-minute tabletop per engineering group, using a real incident from the 31.
- Quarterly full-scale game day: regional failover, ledger replica promotion, dependency failure. Whole role structure, timed.
- Twice-yearly unannounced paging drill, including nights, to measure real acknowledgement times.
- One security incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Also exercise status-page failure and loss of the primary chat or pager.
- Use replicas, staging, or tightly governed tests for ledger scenarios.
- Every exercise produces actions in the same tracker as real incidents.
- Complete two cross-company exercises before the SOC 2 audit.
20. Critical-path pilot (depends on: 9, 11, 12, 18, 19)
Prove the process on the highest-risk surface with willing teams before asking 28 teams to adopt it.
Publish a one-page result company-wide. That result is the adoption argument.
- Six-week Wave 0: ledger app, Postgres platform, payments orchestration, Kubernetes/platform, API gateway, Support intake, plus two of the 12 teams already on-call.
- Activate the full stack: severity scale, Duty IC, one pager, page budget, status page, mandatory postmortems, paid on-call.
- Parallel-run old paths for one week, then cut over. The program lead coaches every SEV2+ and does not secretly take command.
- Weekly retro. Expect 20–40 process defects and fix them in the standard before rollout.
- Exit gate: MTTD under 10 minutes on pilot services; IC assigned within 5 minutes in 95% of incidents; page volume down 50%; all postmortems on time; no unpaid pages; sentiment not worse.
21. Metrics, reviews, and error budgets (depends on: 12, 17, 20)
Instrument the process itself. Do not reward hiding incidents or suppressing pages.
Every metric has a target and a named owner. Dashboards are public inside the company.
- Response: MTTD, time to declare, time to IC, MTTA, MTTM, MTTR, incidents by severity, percent customer-first. Report median and p90, by tier, journey, and region.
- Quality: pages per person per week, actionability, budget breaches, postmortem on-time rate, action closure and age, missed acknowledgements.
- Business: credits, 99.95% per capability, error-budget burn, failed-payment volume, reconciliation breaks, repeat-incident rate.
- People: rotation size, frequency, after-hours pages, recovery days, sentiment, attrition among on-call staff.
- Cadence: weekly Incident Review Board; monthly Reliability Review; quarterly exec and board review; annual policy review.
- Error budgets on Tier 0/1 SLOs. Burn too fast and the team pauses features to pay down reliability.
- Reconcile dashboards monthly against a sample of incident records and customer cases so missing incidents cannot hide.
22. Wave rollout with readiness gates (depends on: 20, 21)
Roll out in four waves by criticality, every three weeks. Gates keep the standard credible. A missed gate is rescheduled, not waived.
Finish all 28 teams by week 24 so roughly three months of operating evidence remain before audit fieldwork.
- Wave 1: remaining Tier 0. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team kit: catalog complete, alerts migrated and inside budget, runbooks at the readiness bar, rotation of 6+ certified responders, one IC nominee, one drill passed, paid rotation in payroll.
- Named coach for three weeks. The director signs the gate.
- Freeze legacy paging per wave.
- Publish a live adoption scoreboard.
- Any production service that cannot provide sustainable ownership gets an executive-reviewed deadline, compensating control, and expiry date. No indefinite verbal exceptions.
23. Change management, incentives, and the deal (depends on: 1, 4, 9)
Run this from day one in parallel with design. Engineers will judge fairness. Executives will judge visible results.
Repeat the deal until it is muscle memory: paid on-call, paged only for what you own, trained commander, real sprint capacity for actions.
- Launch via CTO all-hands, per-team roadshows, a one-page laptop card, an internal wiki, and a Slack help channel with a 4-hour answer SLA.
- Put incident-response contribution into promotion criteria. Award best postmortem and biggest noise cut quarterly. Thank people publicly after every SEV1.
- Put adoption, alert hygiene, action closure, and on-call load fairness into every engineering manager's quarterly objectives.
- Write an exception path for engineers who cannot do nights because of caring responsibilities or health, covered by stipended volunteers.
- Prohibit retaliation for good-faith declaration or escalation.
- Pulse-survey at 60 and 120 days. If fairness or load is red, pause expansion until fixed.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
24. SOC 2 evidence, mock audit, and sustainability (depends on: 17, 21, 22)
Evidence is a by-product of doing the work, not a reconstruction the week before the auditor. Protect the process after SOC 2 is signed.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to CC7.2–CC7.4, CC2.2/CC2.3, CC5, and availability A1.2. Confirm interpretation early.
- Maintain versioned, signed policies for incident response, severity, on-call, communications, and postmortems, reviewed annually.
- Automate evidence: incident records, paging and ack logs, status history, postmortem library, action closure, training register, drill records, and reportability decisions including not reportable.
- Internal dry-run at month 6 on a 15-incident evidence walkthrough. Mock audit at month 7. Fix with weeks to spare.
- Assign permanent owners for policy, pager, status page, catalog, training, and metrics.
- Year-two roadmap: follow-the-sun cell, self-healing for the top recurring causes, error-budget release gates, and blast-radius reduction on the shared ledger.
- Quarterly board summary of severe incidents, credits, overdue P0s, and resilience investment so attention does not die after the audit.
--- PROPOSAL 5 (agent deepseek-v4-pro_refine_5, deepseek/deepseek-v4-pro) ---
Estimated complexity: high
Success metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 6 months of full rollout.
- Customer-first detection falls from 40% to below 10% within 6 months and below 5% within 12.
- Median time to mitigate SEV0/SEV1 falls from 3h10 to under 60 minutes within 12 months; SEV2 under 2 hours.
- Incident Commander assigned and announced within 5 minutes for 95% of SEV0/SEV1 incidents; zero incidents with unclear command beyond 10 minutes.
- Status page updated within policy time for 95% of SEV0/SEV1/SEV2 (10/15/30 minutes).
- Monthly pages fall from 3,400 to under 500 with actionability above 75%.
- All six legacy alerting tools decommissioned by week 16.
- 100% of SEV0/SEV1 incidents have a blameless postmortem published within 10 business days.
- Postmortem action closure rises from 17% to 90% for P0/P1 actions on time.
- SLA credits fall from $1.3M to under $400K in first 12 months.
- Customer-impacting incidents decline to under 15 per year; repeat root causes under 10%.
- 100% of 180 services have a named owning team and criticality tier.
- All 28 teams onboarded by week 24; every Tier 0/1 team has 24x7 primary+secondary coverage with 6+ certified responders.
- At least 30 certified Incident Commanders and 20 certified Communications Leads active.
- Paid on-call policy is approved and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among on-call engineers.
- Quarterly game days and twice-yearly unannounced paging drills executed on schedule, each with tracked action items.
Steps (26):
1. Executive mandate, program governance, and interim incident command
Convert the CEO's outage complaint into a company-level improvement program with a named owner, budget, and deadlines. In the first week, establish an interim command process so no incident remains unowned while the permanent process is designed.
- Appoint Head of Reliability as program owner and CTO/CEO as executive sponsor.
- Form steering group with Engineering, SRE, Support, CS, Legal, Compliance, HR, Finance, Security.
- Approve non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking.
- Fund tooling, on-call compensation, training, and 3 dedicated program FTEs; anchor to $1.3M credits.
- Set timeline: design weeks 1-4, tooling and pilot weeks 5-10, rollout weeks 11-20, audit evidence collection from week 8, dry-run month 6.
- Stand up interim 24x7 duty officer and single declaration path within 48 hours, compensated retroactively.
2. Baseline data and alert estate analysis (depends on: 1)
Re-open the last 31 incidents and profile the current alert estate so every later design decision is evidence-based.
- Re-code each incident: detection source, timestamps, owner, severity, credits, root-cause family.
- Quantify customer-first detections and missing signals.
- Analyze two 'nobody in charge' incidents minute-by-minute.
- Inventory six alerting tools: volume, noise, owner, runbook coverage, top 50 noisy rules.
- Freeze baseline metrics: MTTD 22m, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed.
3. Stakeholder listening and resistance mapping (depends on: 1, 2)
Treat engineer pushback as design input. Interview all 28 teams plus Support, CS, and Sales to understand real objections and find early adopters.
- Test objection: unpaid work, night load, unfamiliar code, poor runbooks, or blame.
- Document current informal practices from 12 on-call teams.
- Recruit 10-15 credible engineers as co-design working group.
- Survey baseline sentiment on on-call and alerts.
- Publish promise: paid on-call, page only for owned services, trained commander, reserved reliability capacity.
4. Service ownership catalog and criticality tiering (depends on: 2, 3)
Build a machine-readable service catalog that assigns each of the 180 services a single owning team, escalation policy, and criticality tier.
- Define fields: owner team, manager, Slack channel, escalation policy, dependencies, dashboards, runbooks.
- Tier 0: ledger, money movement, auth, shared PostgreSQL; Tier 1 customer-facing; Tier 2 internal; Tier 3 non-critical.
- Map customer-visible capabilities and dependencies across regions.
- Identify orphan and shared services; force ownership decision within 30 days or schedule decommissioning.
- Publish coverage gaps by tier; Tier 0 gaps become executive escalations.
5. Define severity levels and declaration triggers (depends on: 4)
Adopt a five-level severity model with objective payment-specific triggers, so declaration is a lookup rather than debate. Anyone may declare; only the Incident Commander may downgrade.
- SEV0: unauthorized, lost, duplicated, or corrupted money movement; ledger integrity loss; confirmed data breach; both regions failed. Page all roles, exec, legal, and consider payment pause.
- SEV1: widespread payment failure, no workaround, severe SLA breach, or single-region total loss. Full role activation and public status page.
- SEV2: significant degradation, multiple customers, or workaround available but SLA risk. IC and SMEs paged; status page if customer-visible.
- SEV3: limited impact with workaround; team-led, business-hours response.
- SEV4: internal/no customer impact; ticket only.
- Auto-escalation: unresolved SEV3 >2h becomes SEV2; unresolved SEV2 >1h becomes SEV1; ledger or security always at least SEV2.
- Provide decision tree and 12 worked examples from actual incidents.
6. Define incident roles, decision authority, and handover rules (depends on: 5)
Codify five roles with written responsibilities and explicit authority, so there is never confusion about who is in charge.
- Incident Commander: owns severity, priorities, cross-team pulls, mitigation decisions; does not code.
- Communications Lead: owns status page, internal updates, account manager briefs, regulator coordination.
- Scribe: maintains timeline, decisions, and evidence for postmortem/audit.
- Subject-Matter Responders: diagnose and remediate only their owned services.
- Executive Liaison (SEV0/SEV1): handles exec communications and external stakeholders.
- Rule: IC identified within 5 minutes and announced in channel; handover announced and logged; roles may combine below SEV2, never at SEV0/SEV1.
7. Design 24x7 incident command and comms staffing model (depends on: 6, 4)
Create a central, trained incident command rotation instead of making each of 28 teams field its own commander.
- Recruit 24-30 certified ICs and 12-16 Comms Leads from a cross-team volunteer pool with manager approval.
- Weekly rotations, primary and secondary; 5-minute acknowledgement SLA with auto-escalation.
- Scribe pool as entry-level rotation.
- Night coverage: paid command rotation in New York timezone initially; evaluate follow-the-sun coverage later.
- Eligibility: certification required; commanders can leave with 30 days notice.
- Ensure distinct persons for IC, CL, and primary SME on SEV0/SEV1.
8. Design team on-call rotations and cross-team escalation policies (depends on: 4, 6)
Define tiered team on-call obligations so engineers are only paged for services they own, and cross-team pages go through the IC.
- Tier 0/1 services: 24x7 primary+secondary, at least 6 trained responders per rotation, one-week shifts.
- Tier 2/3: business-hours on-call; after-hours escalation list to engineering manager.
- Platform, infrastructure, database: 24x7 due shared ledger and Kubernetes.
- Routing: every page resolves via service catalog to owning team's escalation policy; cross-team pages only by IC.
- Guardrails: max one primary week in four, no consecutive weeks, no on-call in week after leading SEV0/1, protected recovery time after night work.
9. Design paid on-call compensation and fatigue safeguards (depends on: 8, 3)
Make on-call paid and legally compliant before any new rotation starts, and convert unpaid pager culture into a fair employment condition.
- Weekly stipends: primary 24x7 $800-1,200, secondary 30-50%, business-hours $300-500; holidays premium.
- Out-of-hours incident pay: $150 per night page plus hourly beyond one hour; time-off-in-lieu after overnight work.
- Additional command rotation stipend and SEV0/1 bonus for active responders.
- Verify FLSA/NY wage rules with HR, Legal, Finance; document exemption treatment.
- Base annual budget against current $1.3M credits.
- Track load and trigger staffing review if >2 after-hours pages per person per week sustained.
10. Enforce alert quality standards and noise budget (depends on: 2, 4, 8)
Replace 3,400 monthly alerts at 85% noise with a contractual paging standard that makes on-call sustainable.
- Every paging alert must have owning team, customer impact statement, runbook link, severity mapping, and tested threshold.
- Page only on customer-visible symptoms or SLO burn; cause-based alerts become tickets or dashboards.
- Set page budget: max 2 out-of-hours pages per person per week; breach triggers mandatory alert-tuning sprint.
- Auto-quarantine alerts with >5 firings/month without action or >70% no-action acknowledgements.
- Target <500 actionable pages/month and >75% actionability within 6 months.
- Weekly per-team alert review, monthly cross-team review.
11. Build detection uplift: synthetics, SLOs, and support intake (depends on: 4, 10)
Shift detection from host metrics to customer outcomes so the company stops hearing about outages from clients first.
- Define SLOs per Tier 0/1 capability: payment initiation, auth, settlement timeliness, API availability, ledger consistency.
- Deploy external synthetic transactions from both regions every 60 seconds, covering full payment flow and ledger write.
- Add ledger assurance checks: replication lag, double-entry balance, settlement window countdown.
- Top-100 customer anomaly detection to catch single-tenant outages.
- Auto-create triage incident from support tickets or account manager keywords within 5 minutes.
- Track customer-detected-first as a defect and require a postmortem action.
12. Consolidate alerting and incident tooling (depends on: 5, 8, 10, 11)
Collapse six alerting tools into one integrated paging and incident management platform to create a single system of record for people and audit.
- Select paging/on-call platform and incident management layer (e.g., PagerDuty + incident.io/FireHydrant).
- Implement one-command Slack declaration that auto-creates channel, bridge, pages roles, sets severity, starts timeline.
- Migrate all monitoring sources into the one tool; decommission legacy paging only after two weeks verified.
- Integrate service catalog, status page, Jira action tracking, Salesforce/CS customer lists, and conference bridge.
- Ensure out-of-band paging and offline fallback if a region or chat tool is down.
- Automate evidence capture for SOC2: timestamps, role assignments, severity changes, comms sent.
13. Define acknowledgement and escalation paths (depends on: 12, 6, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts at first signal.
- For SEV0/SEV1: page primary on-call; after 5 minutes unacked page secondary; at 10 page manager and duty IC; at 15 page executive officer.
- For SEV2: primary ack within 10 minutes, IC assigned within 15; escalate on miss.
- For SEV3: ack within 30 minutes or create work item.
- Automatic duty IC page for ledger/security, cross-team, or unresolved ownership.
- If impact unknown after 15 minutes, raise severity.
- Cross-team responders summoned by IC have 10-minute acknowledgement obligation.
- Route alerts with no owner to command rotation, then treat missing ownership as control defect.
- Human acknowledgement required; delivery confirmation is not sufficient.
14. Standardize internal communications (depends on: 6, 5)
Separate the working war room from executive and stakeholder updates, with fixed cadence and pre-approved templates.
- Auto-create incident channel and one read-only broadcast channel for execs, support, sales.
- SEV0: internal update every 15 minutes; SEV1 every 30; SEV2 every 60; SEV3 on state change.
- Template: impact, what we know, what we are doing, ETA/next update, current IC and CL.
- Executives questions through Executive Liaison only; IC is not interrupted.
- Support/CS receives affected-customer list and holding statement within 15 minutes for SEV0/SEV1 and 30 for SEV2.
- Handover protocol for incidents lasting >4 hours: formal IC handover and fatigue check.
15. Standardize customer and status page communications (depends on: 5, 14)
Replace ad-hoc status page updates with a timed, role-owned, template-driven process, including account manager outreach.
- Status page timing: SEV0 initial post within 10 minutes, SEV1 within 15, SEV2 within 30; updates every 15/30/60 minutes until resolved.
- Resolution notice within 30 minutes of mitigation; customer-facing summary in 5 business days for SEV0/SEV1.
- Pre-approve 12-15 templates with Legal and Comms.
- Tiered outreach: top 100 accounts get direct call/email from AM within 30 minutes of SEV0/SEV1; long tail via subscription.
- Use factual language: state impact and next update; never speculate cause or blame vendor.
- Comms Lead is sole author for customer language.
16. Regulatory, legal, and account manager notification playbook (depends on: 5, 15)
Build a notification decision tree and contact matrix so legal/regulatory obligations are assessed early and never forgotten.
- Map obligations: NYDFS Part 500 72-hour cybersecurity event notification, state breach laws, GLBA/FTC, PCI, sponsor bank/card network contractual windows, FinCEN/OFAC if relevant.
- Add regulatory assessment checkpoint for every SEV0 and security SEV1 within 2 hours, even if not reportable.
- Maintain 24x7 contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel with named backups.
- Pre-draft notification templates and test quarterly.
- Encode customer-specific notification SLAs from enterprise contracts into customer tiering.
- Account managers receive legal-approved script and affected-customer list.
17. SLA credit workflow and financial impact measurement (depends on: 5, 15)
Link every incident to money automatically so severity, credits, and prioritization stay consistent and Finance is never surprised.
- Define availability measurement per contract and capability with Legal/Finance.
- Auto-compute affected minutes per customer from incident record and telemetry; generate credit proposal within 5 business days.
- Decide proactive credits for top tier vs claims-based for others; document approval chain.
- Track credits by root-cause family and incident; feed quarterly reliability investment decisions.
- Target reducing annual credits from $1.3M to under $400K in first year.
18. Standardize blameless postmortems (depends on: 5, 6)
Make postmortems mandatory with fixed deadlines, a single format, and a blameless review forum, replacing 'various formats'.
- Mandatory: all SEV0/SEV1; SEV2 with customer impact, credits, >2h, repeat cause, or detected by customer; near-miss involving ledger.
- Draft within 3 business days, peer review within 5, publish within 10.
- One template: timeline, customer/financial impact, detection gap, response gap, contributing factors, what went well.
- Blameless rules: context-based, no individual blame, never used in performance reviews.
- Weekly Incident Review Board reviews all postmortems and challenges quality.
- Publish searchable postmortem library and quarterly recurring causes report.
19. Track postmortem actions with owner and due date (depends on: 18, 12)
Fix the 11/64 completion rate by giving every action the same status as customer commitments, with capacity and escalation.
- Every action gets named owner, priority, due date, and Jira ticket auto-created from postmortem.
- P0 actions prevent SEV0 recurrence, due 30 days; P1 60; P2 90.
- Reserve 15-20% team sprint capacity for incident actions.
- Escalation ladder: manager at 7 days overdue, director at 14, CTO dashboard at 30; P0 overdue blocks features.
- Monthly reporting of closure rate in engineering leadership.
- Target 90% closure of P0/P1 on time in two quarters.
20. Create runbooks and service readiness bar (depends on: 4, 8, 11)
Ensure every service is prepared for 3 a.m. response before it is allowed to page anyone.
- Readiness checklist for Tier 0/1: architecture diagram, dependencies, dashboards, rollback, feature flags, escalation contacts, data-loss statement.
- Write major-incident playbooks: shared PostgreSQL failure, cross-region failover, Kubernetes control plane loss, partner outage, settlement breach, security compromise.
- Prioritize ledger playbooks: documented failover, read-only mode, reconciliation, signed RPO/RTO.
- Runbooks must be tested twice yearly; stale runbooks marked in catalog.
- No paging alerts without readiness sign-off; gap reported to director.
21. Train, certify, and simulate incident response (depends on: 6, 13, 14, 15, 18, 20)
Build a certification path so 24x7 roles are staffed by people who have practiced, and validate the process through drills.
- Scribe course (2 hrs); Responder (half-day); Communications Lead (1 day); Incident Commander (2 days + shadowing/tabletop).
- Certify ICs and CLs; recertify annually.
- New responders shadow two shifts before primary; no primary in first 90 days.
- Monthly tabletop using real past incidents; quarterly game day including regional failover or ledger scenario.
- Twice-yearly unannounced paging drill; one regulator/legal exercise annually.
- Track drill metrics and action items.
22. Pilot on critical services and iterate (depends on: 12, 13, 15, 19, 21)
Prove the process on the highest-risk surface before full rollout. Run a six-week pilot with tight measurement and public results.
- Select 5-6 teams: payments core, ledger/database, platform/Kubernetes, API gateway, plus two existing on-call teams.
- Activate full stack: severity, roles, command rotation, paid on-call, single tooling, alert budget, status page, postmortems.
- Weekly pilot retro; fix defects within 48 hours.
- Exit criteria: MTTD <10 min, IC assigned within 5 min in 95% incidents, pages down 50%, postmortems on time, positive sentiment.
- Publish one-page pilot result as adoption argument.
23. Wave rollout to all teams and decommission legacy paths (depends on: 22, 20)
Roll out to all 28 teams in four waves by criticality, with explicit readiness gates and legacy tool shutdown.
- Wave 1 remaining Tier 0; Wave 2 Tier 1; Wave 3 Tier 2; Wave 4 Tier 3/internal.
- Per-team onboarding: catalog complete, alerts pruned, runbooks ready, rotation staffed with 6+ certified responders, one IC candidate, one drill passed.
- Coach per wave 3 weeks.
- Gate signed by director; failures rescheduled, not waived.
- After onboarding, freeze legacy alerting paths; no fallback to old tools.
- Publish adoption scoreboard.
24. Establish metrics, dashboards, and review cadence (depends on: 12, 19, 22)
Make process performance visible with a small set of metrics and fixed review meetings, so the program is owned by data.
- Response: MTTD, time to IC, MTTA, MTTM, MTTR, % customer-detected-first.
- Quality: page volume per person, alert actionability, postmortem on-time, action closure rate.
- Business: SLA credits, availability vs 99.95, repeat incidents.
- People: on-call load, page per engineer, sentiment, attrition.
- Cadence: weekly Incident Review Board, monthly Reliability Review, quarterly Executive/Board review.
- Dashboards self-serve with targets and named owners.
25. SOC 2 readiness and internal dry-run audit (depends on: 19, 23, 24)
Design evidence as a by-product and test it with an internal walkthrough before the external auditor arrives.
- Map process to SOC2 CC7.3/CC7.4, CC7.2, CC2.2/2.3, CC5, availability criteria.
- Publish versioned policy documents: Incident Response, Severity, On-Call, Communication, Postmortem.
- Automate evidence: incident records with timestamps, role assignments, paging logs, status history, postmortems, action board, training register, drill records.
- Ensure process operates at least 3 months before fieldwork.
- Run internal dry-run month 6; sample 15 incidents; fix gaps with 8 weeks to spare.
- Keep remediation log for process deviations.
26. Continuous improvement, culture, and sustainability (depends on: 23, 25)
Prevent the process from decaying after the audit by embedding review, feedback, and roadmap ownership.
- Quarterly process retrospective with IC pool and responders.
- Re-baseline metrics every six months; raise targets.
- Year-two roadmap: follow-the-sun coverage, self-healing top 3 causes, error budgets gating releases, blast-radius reduction for shared ledger.
- Annual policy review and certification renewal.
- Quarterly on-call sentiment survey with published actions.
- Board quarterly report on availability, credits, and incident trends.
Step-level differences computed by the tool:
Proposal 1 vs the previous-round proposal it resembles most (opus5_refine_1): 28 steps kept, added ['Charter the program: one owner, one mandate, funded, dated before the audit', 'Seven-day operating floor so the next outage already has an owner', 'Three-layer 24x7 coverage: command corps, domain rotations, triage desk', 'Live execution doctrine and payments safety rules'], removed ['Executive mandate, single owner, funding and non-negotiables', 'Seven-day interim command bridge', '24x7 coverage model: central command corps, local expertise', 'Live incident execution doctrine']
Proposal 2 vs the previous-round proposal it resembles most (gpt5.6-sol_refine_2): 17 steps kept, added ['Install a seven-day incident-response floor', 'Create the factual, legal, and audit baseline', 'Co-design the fairness contract with engineers', 'Enforce an alert-quality contract and page budget', 'Standardize status-page and customer communications', 'Give corrective actions enforceable ownership', 'Publish policy and certify every response role', 'Measure performance and operate fixed review forums', 'Institutionalize improvement and reduce structural risk'], removed ['Install an interim process in seven days', 'Design compliance and evidence controls from day one', 'Enforce alert quality and burn down noise safely', 'Enforce action ownership and effectiveness tracking', 'Measure outcomes, controls, business impact, and human load', 'Train participants and address pager resistance', 'Institutionalize continuous improvement']
Proposal 3 vs the previous-round proposal it resembles most (opus5_refine_1): 27 steps kept, added ['Establish executive mandate, program office, and funding', 'Map compliance, evidence, and audit requirements from day one', 'Build the service ownership catalog and criticality tiers', 'Design the three-layer 24x7 coverage model', 'Build payment-outcome detection and ledger assurance'], removed ['Executive mandate, single owner, funding and non-negotiables', 'Service catalog, ownership and money-path tiering', '24x7 coverage model: central command corps, local expertise', 'Detection uplift on the money path', 'Runbooks, readiness bar and ledger blast-radius reduction']
Proposal 4 vs the previous-round proposal it resembles most (opus5_refine_1): 9 steps kept, added ['Charter the program and start the audit clock', 'Install a seven-day operating floor', 'Rebuild the forensic baseline', 'Map resistance and publish the fairness contract', 'Lock severity levels and what each one triggers', 'Staff 24x7 with a command corps, not 28 night rotations', 'Pay on-call, meet New York labor rules, and cap fatigue', 'Detect payment and ledger failures before customers do', 'Write the five-minute path from signal to command', 'Codify live execution, readiness, and ledger playbooks', 'Run one communications clock for internals, customers, and regulators', 'Tie incidents to SLA credits and true financial cost', 'Track actions as risk commitments with reserved capacity', 'Train and certify every role before independent duty', 'Publish Policy v1 and the signed fairness contract', 'Roll out in risk-ordered waves with readiness gates', 'Inspect at 90 days and lock year-two ownership'], removed ['Executive mandate, single owner, funding and non-negotiables', 'Seven-day interim command bridge', 'Forensic baseline of incidents, alerts and money lost', 'Listening tour, resistance map and the on-call deal', 'Severity scale, declaration rules and incident lifecycle', '24x7 coverage model: central command corps, local expertise', 'On-call compensation, labour compliance and fatigue safeguards', 'Detection uplift on the money path', 'Escalation ladder and the five-minute command rule', 'Live incident execution doctrine', 'Internal communications protocol', 'Customer communications and status page policy', 'Regulatory, partner and legal notification playbook', 'SLA credit and financial impact workflow', 'Action ownership, capacity reservation and enforcement', 'Runbooks, readiness bar and ledger blast-radius reduction', 'Training and certification academy', 'Publish Incident Management Policy v1', 'Wave rollout to all 28 teams with readiness gates', 'Alert noise burn-down campaign', 'Change management, fairness and pager culture', 'Risk register and contingencies', 'Ninety-day inspect-and-adapt, then year-two sustainability']
Proposal 5 vs the previous-round proposal it resembles most (opus5_refine_1): 17 steps kept, added ['Interim 7-day command', 'Listening tour and fairness contract', 'Service ownership catalog and criticality tiers', '24x7 command and responder coverage model', 'Service readiness bar and runbooks', 'Escalation paths and acknowledgement SLAs', 'Standardize internal and customer communications', 'Simulations and game days', 'Culture, fairness, and continuous feedback'], removed ['Seven-day interim command bridge', 'Listening tour, resistance map and the on-call deal', 'Service catalog, ownership and money-path tiering', '24x7 coverage model: central command corps, local expertise', 'Escalation ladder and the five-minute command rule', 'Live incident execution doctrine', 'Internal communications protocol', 'Customer communications and status page policy', 'Regulatory, partner and legal notification playbook', 'Action ownership, capacity reservation and enforcement', 'Runbooks, readiness bar and ledger blast-radius reduction', 'Exercise programme: tabletops, game days and unannounced drills', 'Publish Incident Management Policy v1', 'Alert noise burn-down campaign', 'Change management, fairness and pager culture']
Origin of the steps of the new proposals, matched by title by the tool (evidence for "taken"; ideas can also travel without a matching title):
Proposal 1: 27 of its 32 steps match its own previous version, 4 are new; step 10 resembles step 10 of proposal 4
Proposal 2: 18 of its 26 steps match its own previous version, 7 are new; step 11 resembles step 10 of proposal 4
Proposal 3: 5 of its 32 steps match its own previous version, 0 are new; steps 2, 3, 7, 8, 10, 11, 13, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 32 resemble steps 2, 3, 6, 7, 9, 10, 12, 17, 18, 19, 20, 22, 23, 24, 25, 28, 26, 27, 29, 30, 32 of proposal 1; step 5 resembles step 4 of proposal 2; steps 4, 6, 9, 12, 15 resemble steps 4, 5, 8, 11, 14 of proposal 4
Proposal 4: 7 of its 26 steps match its own previous version, 10 are new; steps 5, 12, 20, 22, 23, 25 resemble steps 5, 12, 23, 25, 26, 30 of proposal 1; steps 7, 11 resemble steps 6, 12 of proposal 2; step 6 resembles step 5 of proposal 5
Proposal 5: 6 of its 26 steps match its own previous version, 4 are new; steps 12, 16, 22, 26 resemble steps 11, 18, 27, 32 of proposal 1; steps 7, 10, 15, 20 resemble steps 6, 9, 14, 20 of proposal 2; steps 1, 6, 21, 24, 25 resemble steps 1, 5, 23, 25, 26 of proposal 3; steps 4, 17, 19 resemble steps 4, 17, 19 of proposal 4
PROPOSALS OF ROUND 2 (to assess):
--- PROPOSAL 1 (agent opus5_refine_1, anthropic/claude-opus-5) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (32):
1. Charter the program: one owner, one mandate, funded, dated before the audit
Turn the CEO email into a chartered company program with a single accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the **burning-platform narrative** and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the written on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (depends on: 5, 7)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 8, 4)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (depends on: 10, 5)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (depends on: 7, 12, 13)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
18. SLA credit and financial impact workflow (depends on: 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Use the credit delta as the standing business case for on-call pay and reserved reliability capacity.
19. Blameless postmortem standard and Incident Review Board (depends on: 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, reserved capacity and enforcement (depends on: 19, 12)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
21. Runbooks, readiness bar and ledger blast-radius reduction (depends on: 5, 8)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (depends on: 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (depends on: 22, 12, 21)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
24. Publish Incident Management Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 11, 12, 21, 22, 9)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Alert noise burn-down campaign (depends on: 10, 12, 25)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
28. Wave rollout to all 28 teams with readiness gates (depends on: 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
29. Change management, fairness and pager culture (depends on: 4, 9, 25)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (depends on: 24, 26, 28)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 26, 28, 30)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 2 (agent gpt5.6-sol_refine_2, openai/gpt-5.6-sol) ---
Estimated complexity: high
Success metrics: - By day 7, every suspected major incident uses one incident record, one coordination channel, and a named commander.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes; at least 95% receive command within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and technical coverage.
- By week 18, all 180 services have an owner, tier, tested escalation path, and coverage appropriate to their risk.
- No mandatory night rotation starts before compensation, training, access, runbooks, and staffing controls are active.
- Every direct 24x7 technical rotation has at least six qualified responders or a documented, expiring executive exception.
- No responder is routinely assigned primary duty more often than one week in six or simultaneously assigned to two primary rotations.
- At least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of qualifying SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet their required internal and customer update cadence by month 3.
- Monthly human pages fall from 3,400 alert events to no more than 1,500 by day 90 and no more than 500 actionable pages by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, with no decline in Tier 0/1 detection coverage.
- Average after-hours load remains at or below two pages per responder per week; every sustained breach produces a remediation plan.
- All six monitoring sources route human pages through the single paging platform by week 12; direct legacy paging paths are disabled by week 18.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Monthly contracted availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- Quarterly on-call surveys reach at least 75% favorable responses on fairness, ownership boundaries, compensation, and sustainability by month 6.
- The month-7 mock audit finds no unowned high-risk control gap, and at least 95% of sampled incidents contain complete operating evidence.
Steps (26):
1. Establish the mandate, owner, funding, and schedule
Launch incident management as a **company operating program within 48 hours**. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Form a decision group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Approve the non-negotiables: one severity scale, one human-paging platform, one incident record, paid on-call, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, and incident actions.
- Fund tooling, compensation, training, exercises, and reliability work. Use the $1.3M in credits as the minimum financial comparison.
- Set milestones: interim process by day 7, policy and critical-path pilot by week 6, Tier 0/1 coverage by week 10, company rollout by week 18, internal audit in month 6, and mock audit in month 7.
2. Install a seven-day incident-response floor (depends on: 1)
Do not wait for policy design or tooling migration. Put a minimum process into operation immediately and begin retaining evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communication schedule.
- Create one monitored declaration path through chat, telephone, and the current paging environment.
- Use one incident channel, bridge, timeline document, and naming convention for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced on-call engineers and engineering leaders.
- Require a named commander within 10 minutes. The duty engineering director assumes command if nobody else does.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Compensate interim duty retroactively under the permanent policy.
- Hold a 15-minute daily operational review until the permanent process is live.
3. Create the factual, legal, and audit baseline (depends on: 1)
Build one defensible baseline for process design, executive decisions, and SOC 2 testing. Confirm the auditor's expected Type II observation period immediately.
- Reconstruct all 31 incidents from first customer impact through resolution, communications, credits, postmortem, and corrective action.
- Analyze the two incidents with unclear command minute by minute.
- Identify the missing detection signal for every customer-first incident.
- Inventory the six alert sources, 3,400 monthly alert events, noisy rules, duplicates, missing owners, and missing runbooks.
- Record the current rotations, unpaid work, after-hours load, and teams without coverage.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M credits, 85% noise, and 11 of 64 actions closed.
- Map incident-response controls to applicable SOC 2 criteria with Compliance and the auditor.
- Establish retention, confidentiality, legal-hold, and access requirements for incident evidence.
4. Co-design the fairness contract with engineers (depends on: 1)
Treat pager resistance as a valid design constraint. Make the employment and ownership bargain explicit before expanding on-call.
- Interview representatives from all 28 teams and the Support, Customer Success, Security, and Operations groups.
- Distinguish objections involving unpaid work, unfamiliar code, bad alerts, weak runbooks, sleep disruption, or blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for services they own or have formally accepted, supported by trained command, and given capacity to remove recurring defects.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Measure baseline trust, fairness, fatigue, and psychological safety. Repeat the survey at days 60 and 120, then quarterly.
5. Build the service catalog and criticality model (depends on: 3)
Make a machine-readable catalog the source of truth for routing, escalation, customer impact, and audit evidence. Every production service must have one accountable owner.
- Record the owning team, manager, product capability, escalation policy, communication channel, dashboard, runbook, dependencies, regions, and data stores for all 180 services.
- Map payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, and reporting to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing services whose failure can create material contractual impact.
- Classify Tier 2 as internal or deferrable services, and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional mode, recovery method, and contractual obligations for Tier 0 and Tier 1.
- Give orphan services an owner or approved decommission date within 30 days.
- Maintain separate but coordinated ownership for the ledger application and PostgreSQL platform.
6. Adopt one severity scale and incident lifecycle (depends on: 3, 5)
Use four impact-based levels. Classify on actual or credible customer, financial, security, regulatory, and contractual harm rather than organizational seniority.
- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; both regions impaired; or a core payment journey broadly unavailable. Page every role immediately, open a bridge, notify executives, publish customer status within 15 minutes when applicable, begin legal assessment, and require a postmortem.
- **SEV2 — major:** material payment degradation; settlement deadline at risk; a critical customer or material cohort unavailable; regional impairment with reduced resilience; or likely SLA breach. Page command and technical roles, publish customer status within 30 minutes when customer-visible, and require a postmortem.
- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads; page only when immediate action is necessary.
- **SEV4 — operational event:** no current customer impact and no urgent risk. Create a ticket and handle during normal operations.
- Treat percentages as supporting guardrails, never reasons to under-classify integrity, settlement, security, or contractual risk.
- Anyone may declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Automatically use at least SEV2 posture for suspected ledger-integrity events, cross-team incidents with unknown ownership, and materially unknown impact lasting 15 minutes.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed. Resolution requires stability, backlog recovery, and necessary reconciliation.
7. Define roles, authority, and handoffs (depends on: 6)
Separate command, communications, recordkeeping, and technical repair. One named person must hold command throughout every SEV1 and SEV2.
- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not serve as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- **Subject-Matter Responders:** diagnose and mitigate only services for which they have ownership, access, training, or formally accepted support responsibility.
- **Executive Duty Officer:** removes organizational barriers and makes exceptional business decisions without displacing the commander.
- Security, Legal, Compliance, Vendor Management, and Finance join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Require distinct commander, communications, scribe, and primary technical lead for SEV1. Communications and scribe may combine temporarily for bounded SEV2 incidents.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, and reconciliation requirements for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, with the exact transfer time.
8. Create sustainable 24x7 coverage across 28 teams (depends on: 4, 5, 7)
Use central command coverage and risk-based technical coverage rather than creating 28 fragile night rotations. All services receive a response path, but only critical domains maintain direct overnight technical rotations.
- Create a 24x7 Incident Command corps of 24–30 certified people, with primary and secondary coverage at all times.
- Create 18–24 trained Communications Leads and a similarly sized scribe pool using Support, Customer Operations, Engineering Operations, and qualified managers.
- Maintain a surge roster for simultaneous incidents and a 24x7 Executive Duty Officer schedule.
- Group Tier 0 and Tier 1 ownership into approximately 8–12 coherent domains only where responders have training, access, runbooks, and explicit acceptance.
- Staff each critical domain with primary and secondary responders and at least six qualified people. Target no more than one primary week in six.
- Give Tier 2 and Tier 3 services business-hours coverage plus a maintained manager and director escalation path.
- When lower-tier impact becomes SEV1 or SEV2, central command activates the manager escalation path and obtains the necessary owner.
- Route unknown-owner pages to Duty Command and platform triage temporarily. Log each one as a catalog control failure.
- Do not merge small teams merely for schedule convenience. Provide training, reassign services, add staffing, or decommission unsupported services.
9. Approve compensation and fatigue protections (depends on: 4, 8)
End unpaid on-call before expanding mandatory night coverage. HR, Finance, Payroll, and employment counsel should approve the policy within 14 days.
- Use market-validated weekly bands, initially budgeting approximately $800–$1,200 for Tier 0/1 primary duty and $300–$500 for secondary duty.
- Budget approximately $900–$1,300 for Duty Incident Commander weeks and $400–$800 for Communications Lead or scribe duty, adjusted for actual burden.
- Pay holiday premiums and compensate active after-hours work according to exempt or non-exempt status and applicable federal and New York rules.
- Provide a protected paid recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal sprint commitments by approximately 15% during a primary on-call week.
- Prohibit simultaneous primary assignments, consecutive primary weeks, on-call during leave, and invisible schedule swaps.
- Allow responders to declare themselves temporarily unfit after disruptive night work without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Budget roughly $0.9M–$1.2M annually, then refine using actual rotation count, employment classification, and activation data.
10. Establish one paging and incident system of record (depends on: 3, 5, 8)
Monitoring tools may remain specialized, but all human pages must enter one controlled platform. This removes conflicting schedules and creates one evidence trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page.
- Ingest events from the six existing monitoring tools before disabling their direct paging paths.
- Route pages through service-catalog ownership and deduplicate related events.
- Provide a single declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, severity changes, decisions, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status-page publishing, and conference facilities.
- Require role-based access, MFA, access reviews, and immutable or tamper-evident history.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, or an AWS region.
- Retire each legacy human-paging path only after ownership review, end-to-end tests, and two weeks of verified operation.
11. Enforce an alert-quality contract and page budget (depends on: 3, 5, 10)
Treat every page as a production interface with an owner and an expected action. Noise reduction must not create detection gaps.
- Require every paging rule to identify the service, owning team, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Page only when prompt human judgment or intervention can materially reduce customer, financial, security, or contractual risk.
- Route informational, capacity, and non-urgent infrastructure conditions to dashboards or ticket queues.
- Prefer payment-outcome, settlement-risk, queue-age, and error-budget-burn alerts over raw CPU, memory, pod, or log thresholds.
- Run new paging rules in shadow mode for at least seven days unless an emergency exception is approved.
- Review rules with repeated no-action acknowledgements, low actionability, or excessive firing within two business days.
- Set a load budget of no more than two after-hours pages per responder per week, measured over four weeks.
- Require an owner, compensating detection, and recorded approval before suppressing or deleting a rule.
- Burn down the top 50 noisy rules first. Review missed detections and noise together so teams cannot improve metrics by becoming blind.
12. Detect payment and ledger failures before customers (depends on: 5, 11)
Move detection from host health to customer journeys and financial outcomes. Use internal SLOs with enough headroom to protect the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, API access, webhooks, and reporting freshness.
- Run external synthetic transactions through critical payment journeys at least every minute and from paths independent of the production platform.
- Validate each AWS region and expose dependencies that undermine nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed value relative to settlement deadlines, not only queue depth.
- Add tenant and cohort anomaly detection for high-value customers and major payment methods.
- Convert high-priority Support, account-manager, processor, sponsor-bank, and network reports into incident candidates within five minutes.
- Record the detection source for every incident. Treat customer-first detection as a mandatory missed-detection review.
13. Codify detection, escalation, and live execution (depends on: 7, 10)
Create one time-bound path from the first credible signal to named command and mitigation. Notification delivery does not count as human acknowledgement.
- Page the owning critical-domain primary and Duty Incident Commander immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at five minutes, manager at 10, and director at 15.
- Escalate an unclaimed command page to backup commander at five minutes. The Executive Duty Officer assumes command at 10 minutes until a certified handoff occurs.
- Require Support and account managers to use the same declaration path as automated monitoring and engineers.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Begin with a standard statement of severity, known impact, assigned roles, current objective, workstreams, and next update time.
- Freeze unrelated changes during SEV1 and normally during SEV2. Record any exception.
- Prefer reversible mitigation such as rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require formal command handoff for long incidents, shift changes, or fatigue. Do not leave an incident unowned during transfer.
- Reconcile the ledger and safely drain backlogs before resolving payment or ledger incidents.
14. Standardize internal incident communications (depends on: 7, 10, 13)
Give responders one working room and stakeholders one controlled information source. Executives must not interrupt the technical command path.
- Create one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue an initial internal brief within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within 15 minutes for SEV1 and 30 minutes for SEV2.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record material decisions and outbound messages in the incident timeline.
15. Standardize status-page and customer communications (depends on: 6, 10, 14)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate impact before root cause is known.
- Publish customer status within 15 minutes of declaring a customer-visible SEV1 and within 30 minutes for customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice promptly after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map status-page components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, data-integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, known workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Give affected strategic accounts direct account-manager outreach within 30 minutes for SEV1 and 60 minutes for SEV2.
- Require account managers to use the approved briefing and prohibit independent technical explanations.
- Provide a customer-facing incident summary within five business days for SEV1 and qualifying SEV2 events.
- Record any legally necessary delay or restriction of public detail, its approver, and the alternative communication plan.
16. Operationalize legal, regulatory, contractual, and credit decisions (depends on: 3, 6, 15)
Some payments incidents start external notification clocks. Make assessment mandatory without assuming that every operational incident is reportable.
- Build a counsel-validated matrix covering applicable NYDFS rules, state breach laws, GLBA or FTC obligations, PCI requirements, money-transmitter obligations, sponsor-bank and network contracts, cyber insurance, and customer contracts.
- Engage Legal and Compliance immediately for every SEV1, security event, and suspected ledger-integrity event.
- Record an initial reportability assessment within one hour for SEV1 and within two hours for other potentially reportable events.
- Record non-reportable decisions as evidence, including facts considered, approver, timestamp, and reassessment trigger.
- Maintain tested 24x7 contacts for counsel, regulators, sponsor banks, networks, insurers, and critical vendors.
- Encode customer-specific notice deadlines and channels in the customer record used by the Communications Lead.
- Have Legal own regulatory text and submission. Keep technical command with the Incident Commander.
- Have Finance calculate affected minutes, delayed value, likely credits, and contractual exposure within five business days.
- Track credits by incident and recurring cause to support reliability investment decisions.
17. Make postmortems mandatory, consistent, and blameless (depends on: 6, 7, 10)
Use one learning standard with fixed deadlines. Keep postmortems separate from performance, misconduct, and disciplinary processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact over two hours, SLA credits, contractual breach, repeated contributing factors, control failures, and ledger-integrity near misses.
- Produce the factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the Incident Commander responsible for the timeline and the owning engineering director accountable for completion.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit analysis of why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the incident commander.
- Describe decisions in the context and information available at the time. Do not name an individual as the root cause.
- Publish broadly useful findings internally while restricting security, privacy, personnel, or privileged material appropriately.
18. Give corrective actions enforceable ownership (depends on: 10, 17)
Treat incident actions as risk commitments, not suggestions. A ticket is not complete until the expected risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Escalate overdue high-risk actions to the manager after seven days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit overdue actions to block related releases where the unrepaired condition could reproduce severe impact.
- Verify effectiveness through tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days. Complete, re-plan, or formally accept each risk.
19. Set the service-readiness bar and critical playbooks (depends on: 5, 8, 11, 12)
A critical service must be supportable at 3 a.m. before its team is placed on direct overnight coverage. Existing critical detection must remain active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and data-integrity constraints.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional failure, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document controlled failover, split-brain prevention, replay protection, backlog recovery, and post-recovery reconciliation.
- Exercise critical runbooks at least twice per year and after material changes.
- Block new Tier 0/1 releases and new paging rules when readiness requirements are missing.
- Handle existing gaps through a named owner, compensating control, executive-approved expiry date, and remediation plan rather than disabling detection.
20. Publish policy and certify every response role (depends on: 6, 7, 8, 9, 11, 13, 14, 15, 16, 17, 18)
Convert the operating design into concise, signed documents and practical training. Training and exercises occur during paid working time.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, and exception standards.
- Provide one-page severity, role, authority, escalation, and communication cards inside the incident tool.
- Train all employees to recognize and declare incidents.
- Train all engineers in severity, acknowledgement, evidence preservation, handoff, and financial-integrity precautions.
- Certify responders only after they demonstrate access, dashboards, runbooks, rollback, and escalation competence.
- Certify Incident Commanders through formal instruction, simulation, and at least two shadowed incidents or exercises.
- Train Communications Leads in status writing, account segmentation, contractual clocks, and legal escalation.
- Train scribes in timeline quality and fact-versus-hypothesis labeling.
- Require two shadow shifts before independent primary duty and annual recertification.
- Maintain the training and certification register as operational and audit evidence.
21. Pilot on the payment critical path (depends on: 9, 10, 12, 19, 20)
Run a four-to-six-week pilot across the highest-risk journey before expanding. Use real incidents and exercises to correct the model quickly.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, API edge, authentication, settlement, and Support intake.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently assuming command.
- Correct critical process or tool defects within 48 hours and update the standard visibly.
- Exit only after 95% timely command assignment, 95% communication compliance, no unpaid pages, complete required postmortems, and at least 50% lower noise.
22. Roll out by customer journey and risk (depends on: 21)
Expand in controlled waves while the interim process remains active company-wide. Complete rollout early enough to accumulate operating evidence before the audit.
- Roll out remaining Tier 0 domains first, then Tier 1, Tier 2, and Tier 3.
- Use two-to-three-week waves with a named program coach and director sign-off.
- Gate each team on catalog ownership, appropriate coverage, compensation, trained responders, tested escalation, alert quality, runbooks, access, and a passed tabletop.
- Require six or more responders only for direct 24x7 technical rotations. Apply business-hours coverage and manager escalation to lower tiers.
- Reschedule failed gates or approve a time-limited executive exception with a compensating control.
- Disable legacy human-paging paths after verified cutover for each wave.
- Publish an internal adoption dashboard by team, service tier, and control gap.
- Finish critical coverage by week 10 and all 28 teams by week 18.
23. Exercise command, communications, regional recovery, and fallbacks (depends on: 19, 20)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous security and operational events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors where appropriate.
- Conduct at least one unannounced after-hours paging test before the audit.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure time to acknowledgement, command, customer notice, mitigation decision, handoff, and recovery.
- Create tracked actions for every material exercise finding.
24. Measure performance and operate fixed review forums (depends on: 3, 10, 17, 18)
Use a balanced scorecard that exposes weak controls without rewarding hidden incidents or suppressed alerts. Report medians and 90th percentiles rather than averages alone.
- Measure time from first impact to detection, declaration, acknowledgement, command assignment, mitigation, recovery, and resolution.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track pages, actionability, duplicates, after-hours load, missed detections, and page-budget breaches.
- Track postmortem timeliness, action age, closure by due date, verified effectiveness, and recurring contributing factors.
- Track availability by customer journey, error-budget burn, failed or delayed payment value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight activations, recovery days, schedule exceptions, sentiment, and attrition.
- Hold a weekly Incident Review Board for incidents, postmortems, control failures, noisy alerts, and overdue actions.
- Hold a monthly executive reliability review for trends, funding, contractual exposure, and accepted risks.
- Hold a quarterly control and resilience review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile incident records monthly against support cases, status history, credits, and customer complaints to detect under-reporting.
25. Prove SOC 2 operating effectiveness before fieldwork (depends on: 20, 22, 23, 24)
Generate audit evidence through normal operation rather than reconstructing it later. Test both control design and consistent execution.
- Maintain approved and versioned policies, exceptions, catalog records, schedules, compensation activation, access reviews, training, incidents, communications, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, non-reportable legal decisions, missed timelines, and exercises in the population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the populations, evidence requests, and interviews expected from the external auditor.
- Record deviations honestly with owners, remediation dates, and compensating controls. Never rewrite historical records.
- Confirm that evidence retention covers the full auditor-defined observation period.
- Brief commanders, responders, Support, and Compliance on the real process without scripting inaccurate answers.
26. Institutionalize improvement and reduce structural risk (depends on: 22, 24, 25)
Prevent the program from decaying after the audit. Use incident evidence to drive permanent ownership and architectural investment.
- Assign permanent owners for policy, service catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review severity thresholds, staffing, compensation, and communication timings annually and after material process failures.
- Recertify command and communications personnel annually.
- Review recurring failure families quarterly and require executive decisions where corrective work repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fairness, fatigue, and tool friction.
- Report severe incidents, credits, overdue high-risk actions, and resilience investment to the board or risk committee quarterly.
- Treat the shared ledger cluster as a strategic concentration risk. Fund failover assurance, blast-radius reduction, isolation of non-critical readers, and stronger regional independence.
- Evaluate follow-the-sun command or technical coverage using one year of page-load and staffing data.
- Build year-two plans for automated mitigation, deployment safety, graceful degradation, and error-budget release controls.
--- PROPOSAL 3 (agent qwen3.8-max_refine_3, alibaba/qwen3.8-max) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (32):
1. Establish executive mandate, program office, and funding
Turn the CEO email into a **chartered company program** with one accountable owner and authority over all 28 teams.
- Appoint the CTO as executive sponsor and a Head of Reliability as program owner with full-time authority.
- Create a permanent program office: one program lead, one platform engineer, one analyst.
- Form an eight-person steering group: Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours.
- Lock the non-negotiables: one severity scale, one paging platform, one postmortem format, paid on-call, mandatory action tracking, named service ownership.
- Approve budget anchored against the $1.3M in credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs.
- Reserve 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Publish a one-page charter stating incident response is a company operating process, not a per-team choice.
2. Install a seven-day interim command bridge (depends on: 1)
Do not leave the company unprotected while the permanent process is designed. Put a **crude but real** command structure in place within one week.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Direct Support to escalate credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical actions: complete, re-plan, or formally risk-accept, prioritizing ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Build the forensic baseline of incidents, alerts, and money lost (depends on: 1)
Rebuild the facts before designing anything. This becomes both the **design input** and the before picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific missing signal. This list drives the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Freeze baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the on-call fairness deal (depends on: 1)
Engineer pushback is the **largest delivery risk**. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS, and Sales in two weeks. Separate the real objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey covering fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
5. Map compliance, evidence, and audit requirements from day one (depends on: 1)
Design evidence as a **by-product of operations**, not a reconstruction before the auditor arrives. Confirm the SOC 2 observation window immediately.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for communications; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure, training register, drill records, reportability decisions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve all policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has a **named owner**. Build a machine-readable catalog as the single source of truth.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact: Tier 0 (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for each Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
7. Adopt the severity scale, declaration rules, and lifecycle (depends on: 3, 6)
Replace judgment calls with a **lookup table**. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter.
- SEV1 (crisis): money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of all roles and executive; bridge in 5 minutes; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 (critical): material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- SEV3 (major, contained): narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected, Declared, Triaged, Mitigated (customer impact ends), Monitoring, Resolved (backlog processed and ledger reconciled), Reviewed.
- Publish a decision tree with 12 worked examples from the real 31 incidents.
8. Define incident roles, authority, and handover discipline (depends on: 7)
Solve nobody-in-charge-for-an-hour by making command **explicit, single-holder, transferable, and logged**. Separate coordination from debugging.
- Incident Commander: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates for SEV1/SEV2.
- Subject-Matter Responders: engineers of the owning team; they mitigate, they do not run the room.
- Executive Duty Officer (SEV1): removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Design the three-layer 24x7 coverage model (depends on: 6, 8)
Do not create 28 night rotations. **Centralise coordination** in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- Layer A, Incident Command corps: approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of approximately 18 from Support, CS, and engineering management. A scribe pool used as the training entry point.
- Layer B, critical-path domain rotations: consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- Layer C, everyone else: business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing: a US-only paid night rotation now, a follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
10. Approve on-call compensation, labor compliance, and fatigue safeguards (depends on: 9, 4)
Unpaid on-call in New York is both a **retention problem and a legal exposure**. Pay must be live in payroll before any mandatory night rotation starts.
- Indicative scheme: approximately $1,000 per primary 24x7 week, $400 secondary, $250 for business-hours rotations, a separate $1,200 Duty Commander stipend, holiday premiums, approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- HR, Finance, and employment counsel publish amounts, eligibility, tax treatment, and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a 15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
11. Set the alert quality standard, page budget, and noise burn-down (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes. Make alert quality a **condition of being allowed to page a human**.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly.
- Target 3,400 to under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
12. Build payment-outcome detection and ledger assurance (depends on: 11, 6)
Stop customers telling you first. Detection must be driven by **payment outcomes and ledger truth**, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer detected first becomes a named defect class with a mandatory tracked action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Consolidate to one paging platform, one incident record, one status page (depends on: 7, 9, 11)
Collapse six alerting tools into a **single operational system of record** so there is one queue, one timeline, and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is unavailable. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
14. Codify the escalation ladder and five-minute command rule (depends on: 8, 9, 13)
Write one unskippable path from something looks wrong to **someone is in charge**. The default action is never waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately, secondary at 5 minutes, domain manager and Duty Commander at 10, Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension, and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
15. Write the live incident execution doctrine and major-incident playbooks (depends on: 8, 13, 14)
Give responders one short operating procedure from the first minutes through closure. Priority is **limiting customer and financial harm**, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the single largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window, an explicit handback to the owning team, Support, and Customer Success, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
16. Standardize internal communications protocol (depends on: 8, 13)
Standardise the internal picture so executives, Support, and Sales are informed **without interrupting** the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
17. Build customer communications, status page, and account-manager outreach (depends on: 7, 16)
Replace whoever is around with a **timed, owned, pre-approved process**. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook, and RSS. All 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity, or blame.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
18. Create the regulatory, partner, and legal notification playbook (depends on: 7, 17)
In payments, some incidents start a **legal clock at detection**. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded even when the answer is not reportable, with evidence, decision-maker, and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
19. Operationalize SLA credit and financial impact workflow (depends on: 7, 17)
Tie incidents to money so severity, credits, and investment decisions **stay honest**, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one. Use the delta as the standing business case for on-call pay and reliability capacity.
20. Establish the blameless postmortem standard and Incident Review Board (depends on: 7, 8)
Replace some incidents, various formats with **one mandatory format, fixed deadlines, and a forum with teeth**.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly top-five recurring causes analysis.
21. Enforce action ownership, capacity reservation, and tracking (depends on: 20, 13)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of **customer commitments** and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
22. Build the training and certification academy (depends on: 8, 14, 16, 20)
Command is a skill, not a title. **Certify before assigning duty**, and use paid working time for all of it.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days + shadowing): command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Comms Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum; never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Run the exercise programme: tabletops, game days, and unannounced drills (depends on: 22, 13, 15)
The process must meet a **simulated SEV1** before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger. Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (depends on: 7, 8, 9, 10, 11, 14, 16, 17, 18, 20, 21)
Collapse the design into a document people will **actually open mid-outage**, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 12, 13, 15, 22, 10)
Prove the process on the **highest-risk surface** with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment.
- Publish a one-page result to the whole company.
26. Run the alert noise burn-down campaign (depends on: 11, 13, 25)
Run noise reduction as a **visible, quota-driven campaign** in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness. Review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 to 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
27. Establish metrics, dashboards, review cadence, and anti-gaming (depends on: 13, 20, 25)
Instrument the process itself so improvement is visible and the auditor sees **evidence of monitoring and review**. Report median and 90th percentile, never averages alone.
- Response: time to detect, declare, commander, acknowledgement, mitigate, resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
28. Wave rollout to all 28 teams with readiness gates (depends on: 25, 27)
Roll out in four waves of six to eight teams every two to three weeks, ordered by **customer risk**. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. Failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Publish a live adoption scoreboard.
29. Run change management, fairness, and pager culture programme (depends on: 4, 10, 25)
Run this from day one in parallel. Engineers judge the process on **fairness**; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. Build SOC 2 evidence by design, internal testing, and mock audit (depends on: 24, 27, 28)
Design the evidence as a **by-product of doing the work**. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any change as an exception.
31. Maintain the program risk register and contingencies (depends on: 1)
Name the ways this programme fails and **pre-commit the response**. Review it monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 27, 28, 30)
Guard against the classic failure: the process **decays once the audit is signed**. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 4 (agent grok4.6_refine_4, xai/grok-4.6) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (26):
1. Charter the program and start the audit clock
Convert the CEO email into a company operating process with one owner, a budget, and an observation window that starts this week.
Incident response is no longer a per-team choice.
- Name the CTO as sponsor and a Head of Reliability as the single accountable owner, with a three-person program office.
- Form a small decision group: Engineering, SRE, Support, Customer Success, Security, Legal, Finance, and HR. The sponsor decides within 48 hours.
- Lock non-negotiables: one severity scale, one paging path, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, and reserved engineering capacity against the $1.3M in SLA credits.
- Timeline: operating floor in 7 days, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, mock audit in month 7, audit in month 8.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for new tools or the final policy. Put a crude but real process in place this week so the next outage has a named commander.
This week is the first audit evidence.
- Publish a one-page interim severity card and one declaration path: Slack command, phone number, and existing pagers, all reaching the same duty person.
- Staff interim primary and backup Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in the incident channel.
- Tell Support to declare from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline (depends on: 1)
Rebuild the facts before locking design. This is both the design input and the before picture for the CEO and the auditor.
Freeze the numbers so they cannot drift during design.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisy rules and every rule with no owner or runbook.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11 of 64 actions closed, on-call in 12 of 28 teams.
4. Map resistance and publish the fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
Publish the deal in writing before any new mandatory pager is assigned.
- Interview all 28 teams plus Support, Customer Success, and Sales in two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the **fairness contract**: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and money-path tiers (depends on: 3)
You cannot page the right person across 180 services until every service has one owner. A wrong owner recreates the pager objection.
The catalog is the single source of truth for paging, impact, status-page components, and audit.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to services, data stores, regions, and third parties, including sponsor banks and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
6. Lock severity levels and what each one triggers (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by who reported it or how hard the fix looks.
Anyone may declare. Nobody is punished for over-declaring. Only the Incident Commander may downgrade, with the evidence recorded.
- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 major, contained: narrow or single-customer impact with a workaround. Owning team leads; business-hours comms; postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages.
- Attach a financial-integrity or security flag to any severity. The flag forces dual-control, Legal, and the regulatory checkpoint without inventing a fifth level.
- Auto-escalate: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer impact ends) → Monitoring → Resolved (backlog processed and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged.
Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. A human validates for SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1 and SEV2. Every handover is announced verbally and in writing with the exact time.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
8. Staff 24x7 with a command corps, not 28 night rotations (depends on: 5, 7)
Do not create 28 night rotations. That is what engineers are rejecting.
Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified volunteers from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to seven months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate 28 teams into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat Lisbon or APAC follow-the-sun as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
9. Pay on-call, meet New York labor rules, and cap fatigue (depends on: 4, 8)
Unpaid on-call in New York is a retention problem and a legal exposure. Pay must be in payroll before any new mandatory night rotation starts.
Publish the numbers. Then ask people to sign up.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
10. Set the alert-quality bar and a hard page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them.
Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers do (depends on: 5, 10)
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail.
Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One Slack command creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is down. Test weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Write the five-minute path from signal to command (depends on: 7, 8, 12)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting.
If nobody claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Codify live execution, readiness, and ledger playbooks (depends on: 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving a root cause.
A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down. No command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers.
15. Run one communications clock for internals, customers, and regulators (depends on: 6, 7, 12)
Replace whoever is around with one timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is not reportable.
- Obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
16. Tie incidents to SLA credits and true financial cost (depends on: 6, 15)
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
Credit calculation is an output of the incident record, not a negotiation.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
17. Make blameless postmortems mandatory and consistent (depends on: 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth.
The discipline lives in the deadlines and the review, not the template.
- Mandatory for every SEV1 and SEV2, any incident detected by a customer first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed.
18. Track actions as risk commitments with reserved capacity (depends on: 12, 17)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
Closing a ticket without evidence of effectiveness does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Target 90% of high-priority actions closed by due date within two quarters.
19. Train and certify every role before independent duty (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
The certification register is an audit artefact.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
20. Rehearse with tabletops, game days, and night drills (depends on: 12, 14, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
21. Publish Policy v1 and the signed fairness contract (depends on: 6, 7, 8, 9, 10, 13, 15, 17, 18)
Collapse the design into a document people will actually open mid-outage, and make it official.
Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
22. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 19, 21)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse. Publish a one-page result to the whole company.
23. Instrument metrics, reviews, and anti-gaming (depends on: 12, 17, 22)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review.
Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
24. Roll out in risk-ordered waves with readiness gates (depends on: 22, 23)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave, not as a background hope. Give every team its noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
25. Produce SOC 2 evidence by operating, then mock-audit (depends on: 21, 23, 24)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions. Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (depends on: 23, 24, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
Name the ways this programme fails and pre-commit the response.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 5 (agent deepseek-v4-pro_refine_5, deepseek/deepseek-v4-pro) ---
Estimated complexity: high
Success metrics: - Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, no loss of Tier 0/1 detection.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% Tier 0/1 services have named owning team, tier, escalation policy, dashboard and runbook by day 30; all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10-12 domain rotations of at least six trained responders.
- 30 certified incident commanders and 18 certified communications leads active.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% required postmortems drafted in 3 business days, reviewed in 5, published in 10 from month 2.
- Postmortem action closure rises from 17% to 90% closed by due date within two quarters; all 53 legacy actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k annualised within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% SEV1/security SEV2 including not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two cross-company exercises completed before audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% sampled incidents; SOC 2 Type II incident response controls pass zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days; no increase in attrition among responders.
Steps (26):
1. Executive mandate, budget, and governance
Secure a written CTO/CEO charter making incident management a company operating process, not a team option. Name one accountable Head of Reliability and a small steering group with authority to decide within 48 hours.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking, and protected engineering capacity.
- Approve budget against the $1.3M annual credits: tooling $150-250k/yr, on-call compensation, training, and 2-3 program FTEs.
- Set timeline: interim process in week 1, design weeks 1-6, pilot weeks 7-12, rollout weeks 13-24, mock audit month 7, SOC 2 at month 8.
- Freeze current baselines: 31 incidents, 22-min MTTD, 40% customer-first detection, 3h10 MTTM, $1.3M credits, 3,400 alerts/month at 85% noise, 11/64 actions closed.
2. Interim 7-day command (depends on: 1)
Put a crude but real process in place immediately so no incident remains unowned while the permanent design proceeds.
- Publish a one-page interim severity card, one declaration path (phone, Slack command, pager), and one incident channel/bridge/timeline naming convention.
- Staff an interim 24x7 duty commander with primary and backup from engineering managers and senior SREs; compensate retroactively under final policy.
- Require a named incident commander within 10 minutes of any suspected major incident, announced in channel.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions in 30 days, prioritizing ledger integrity, duplicate payment, regional failover, security and detection gaps.
- Hold a daily 15-minute operations review until the permanent process is live.
3. Forensic baseline and alert estate analysis (depends on: 1)
Reconstruct the true before picture from the last 12 months; it drives design and serves as audit baseline.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact start/detect/declare/commander/mitigate/resolve, who led, credits paid, contributing factors.
- For every customer-first detection, name the missing signal; this becomes the detection backlog.
- Reconstruct the two `nobody in charge` incidents minute by minute; use them as the burning-platform narrative and test case.
- Profile the 3,400 monthly alerts by tool, team, rule, outcome; identify top 50 noisy rules and every rule with no owner or runbook.
- Freeze all baseline metrics in one signed document for the executive and the auditor.
4. Listening tour and fairness contract (depends on: 1, 3)
Treat engineer pushback as the main delivery risk and convert it into a written deal that removes the objection to carrying a pager for other teams' code.
- Interview all 28 teams plus Support, CS and Sales in two weeks to separate objections: unpaid work, nights, unfamiliar code, missing runbooks, or fear of blame.
- Harvest practices from the 12 teams already on-call; they supply pilot teams and first commanders.
- Publish the deal: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Recruit 10-15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey and repeat at 60, 120 and 365 days.
5. Service ownership catalog and criticality tiers (depends on: 3)
Create a machine-readable catalog as the single source of truth for paging, routing, status page components, and audit evidence.
- Assign one accountable team, engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list per service.
- Tier by business impact: Tier 0 money movement/ledger/auth/shared PostgreSQL; Tier 1 customer-facing degradable; Tier 2 internal/batch; Tier 3 non-critical.
- Map customer journeys (initiate, authorize, settle, reconcile, report, onboard) to services, data stores, regions, third parties.
- Give orphan services an owner within 30 days or a decommission date approved by the sponsor; Tier 0 without an owner is an executive escalation.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
6. Severity scale, declaration rules, and automatic triggers (depends on: 3, 5)
Adopt one severity scale as a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade with recorded rationale. Start high when uncertain.
- SEV1: money moved wrongly/duplicated/lost, ledger integrity in doubt, confirmed security/data exposure, both regions impaired, payment processing halted. Full role page, bridge in 5 min, status page in 15 min, regulator assessment within 1 hour, mandatory postmortem.
- SEV2: material degradation of a payment journey, settlement window at risk, large/strategic customer fully down, SLA breach likely. Commander and SMEs paged, status page in 30 min, mandatory postmortem.
- SEV3: narrow or single-customer impact with workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any ledger cluster incident, any SEV3 open >2h, any unknown impact after 15 min, any cross-team incident become at least SEV2.
- Publish a decision tree with 12 worked examples from the 31 real incidents.
7. Incident roles, authority, and handoffs (depends on: 6)
Codify roles so coordination never depends on seniority or heroics. The Incident Commander coordinates; responders fix only services they own.
- IC: owns severity, priorities, roles, cadence, closure; pre-authorized to freeze deploys, rollback, disable features, shift traffic, invoke failover, put ledger read-only, commit spend; does not type in terminals; keeps command when VP joins.
- Communications Lead: sole author for status page, account managers, executives, and hand-off to Legal for regulators; speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, state changes; human validates for SEV1/SEV2.
- Subject-Matter Responders: diagnose and mitigate only their owned services.
- Executive Duty Officer (SEV1): removes obstacles, shields IC from exec questions, owns board/regulator escalation; does not take command unless formally transferred.
- Rule: IC claimed within 5 min and announced; distinct people for IC, Comms, technical lead at SEV1/SEV2; every handover announced verbally and in writing with time.
- Financial controls survive: IC coordinates ledger recovery but cannot bypass dual control, reconciliation or privileged access.
8. 24x7 command and responder coverage model (depends on: 5, 7)
Do not create 28 fragile night rotations. Centralize coordination in a trained command corps, keep technical ownership local.
- Layer A: incident command corps of ~30 certified volunteers with primary/secondary 24x7; paired Comms Lead pool ~18 and scribe pool as entry.
- Layer B: consolidate 28 teams into 10-12 critical-path domains (ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes, data/reporting, integrations) each with 24x7 primary+secondary and minimum six trained responders.
- Layer C: all other teams business-hours on-call with written after-hours escalation lists held by managers.
- Platform on-call is safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate follow-the-sun only as a 12-month option, not year-1 dependency.
- Acknowledgement discipline: 5 min at SEV1/SEV2 with automatic failover.
9. Paid on-call and fatigue safeguards (depends on: 8, 4)
Unpaid on-call in New York is a retention and legal risk. Pay must be in payroll before any mandatory night rotation starts.
- Indicative: ~$1,000 per primary 24x7 week, secondary ~$400, business-hours ~$250, duty commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond one hour.
- Mandatory recovery: paid recovery day after >2h overnight work, SEV1, or qualifying SEV2; managers arrange coverage.
- HR/Legal/Finance publish amounts, tax treatment, FLSA/NY wage-hour status within 14 days.
- Load rules: no primary more often than 1 in 6, no consecutive primary/secondary weeks, no primary on two rotations, no on-call during PTO, annual cap.
- If a rotation averages >2 out-of-hours pages per person per week for 4 weeks, trigger staffing or alert-remediation review.
- Count on-call as ~15% delivery load; document exemption path for health/caring without career penalty.
10. Service readiness bar and runbooks (depends on: 5, 8)
A service must earn the right to page a human at 3 a.m. Define a minimum readiness bar and major incident playbooks before any Tier 0/1 service goes live with on-call.
- Require for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switch, escalation contacts, data-loss/latency impact.
- Write playbooks for top failures: shared PostgreSQL failure/corruption, cross-region failover, Kubernetes control-plane loss, processor/sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, duplicate payment.
- Treat ledger cluster as largest structural risk: documented failover, read-only degraded mode, reconciliation after recovery, executive-signed RPO/RTO.
- Start a parallel workstream on blast-radius reduction: tenant/function partitioning, read replicas, isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging at night unless manager accepts a dated exception in writing.
- Runbooks are peer-reviewed, version-controlled, marked stale if not exercised twice a year.
11. Alert quality standard and page budget (depends on: 3, 5)
Make alert quality a condition of paging a human. Burn down noise deliberately rather than by mass silencing.
- Every paging alert must declare owning team, affected service, customer/SLO impact, expected action, dashboard, runbook, dedup key, severity mapping, escalation policy. Anything failing becomes a ticket.
- Page on symptoms of customer harm (SLO burn rate, payment success rate, queue age vs settlement deadlines), not raw CPU/memory.
- Run new alerts in shadow mode 7 days; test firing and recovery.
- Set page budget of 2 out-of-hours pages per person per week; breach blocks new alert creation and triggers tuning sprint.
- Auto-quarantine rules firing >5 times/month without action or >70% no-action acks; return to owner with correction deadline.
- Never silently disable: verify compensating detection, record decision, name owner, review missed detections monthly.
- Target 3,400 to <500 pages/month, actionability >75% within six months, no loss of Tier 0/1 detection.
12. Detection uplift on money path (depends on: 5, 11)
Stop customers telling you first by detecting on payment outcomes and ledger truth.
- Define SLOs and business SLIs per customer journey: initiation success, authorization latency, settlement timeliness, reconciliation break rate, API availability, reporting freshness; set internal targets stricter than 99.95%.
- Run external synthetic end-to-end payments every 60 seconds from both regions, covering all critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate detection, replication lag, failover-readiness, settlement-window countdown.
- Add per-customer anomaly detection for top 100 accounts; five similar support tickets in 10 minutes auto-creates triage incident.
- Route partner and processor notifications into declaration path within 5 minutes.
- Record detection source on every incident; treat `customer detected first` as a named defect class with mandatory action.
13. Consolidate to one paging and incident platform (depends on: 6, 8, 11, 12)
Collapse six alerting tools into one operational system of record: one queue, one timeline, one audit trail, without creating a monitoring gap.
- Select one paging/scheduling platform, one incident record, one hosted status page; time-box selection to 2 weeks.
- Ingest all six sources first, deduplicate/correlate; retire legacy path only after named owners, successful end-to-end test, and 2 weeks verified operation.
- One-command declaration in Slack creates channel/bridge, pages commander, sets severity, opens timeline, starts clock.
- Route every page through service catalog: service label -> owning domain schedule -> escalation policy.
- Capture evidence automatically: declaration, acks, role assignments, severity changes, decisions, comms, mitigated/resolved times, postmortem link; retain 12 months.
- Test out-of-band SMS/phone paging, mobile fallback, offline runbook weekly; ensure works during one AWS region/chat/identity provider failure.
- Set hard date after which pages outside this tool create no on-call obligation.
14. Escalation paths and acknowledgement SLAs (depends on: 7, 8, 13)
Write one unskippable path from signal to named commander in under 5 minutes; default action is never waiting.
- Converge all entry points on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately -> secondary at 5 min -> domain manager + Duty Commander at 10 -> Executive Duty Officer at 15. SEV2 requires command assigned within 15 min.
- If no one claims command in 5 min, platform assigns and announces it; assignee may hand over but not decline.
- Commander may page any domain's on-call directly with 10-min ack obligation; this reciprocity makes single-team ownership viable.
- If ownership unclear after 10 min, commander keeps incident and names temporary owner; missing catalog entry logged as control defect.
- Pre-authorize regional failover, ledger read-only, payment suspension, partner-bank notification so IC never waits for an executive.
15. Standardize internal and customer communications (depends on: 6, 7, 13)
Replace whoever is around with a timed, owned, pre-approved process; the Communications Lead is single author and never writes from scratch under pressure.
- Internal: one working channel/bridge plus one read-only broadcast for executives/Support/Sales; cadence SEV1 15-30 min, SEV2 60 min even if nothing changed; fixed template with impact, action, next update, IC/Comms names.
- Executives ask questions only of Executive Duty Officer; publish this as signed behavior rule.
- Status page: SEV1 within 15 min, SEV2 within 30 min; updates every 30/60 min; resolution notice within 30 min of verified recovery; customer-facing summary within 5 business days for SEV1/qualifying SEV2.
- Pre-approve 12-15 templates with Legal for degradation, settlement delay, API errors, regional failure, data-integrity investigation, security event.
- Top 100 accounts get named account manager call/email within 30 min of SEV1 with briefing pack; all 2,100 subscribed by default.
- Language rules: state impact and next update; never speculate on cause, recovery time, data integrity or blame.
16. Regulatory, partner, and financial impact workflows (depends on: 6, 15)
In payments some incidents start a legal clock at detection. Build obligation assessment into the process and tie incidents to money.
- Legal/Compliance produce obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC, PCI, money-transmitter, sponsor-bank/card-network windows, cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and security-related SEV2 within 1 hour, recorded even when `not reportable`, with decision-maker and evidence.
- Maintain 24x7 contact matrix for regulators, sponsor banks, networks, outside counsel, insurer.
- Agree availability measurement method per contract; compute affected minutes per customer from incident record and journey telemetry; propose credit schedule within 5 business days.
- Attribute credits to root-cause families; track failed payment volume, delayed value, reconciliation breaks; target annual credits from $1.3M to under $400k.
17. Blameless postmortem standard and action tracking (depends on: 6, 7)
Replace various formats with one mandatory format, fixed deadlines, and enforceable actions. Closing a ticket without evidence does not close the action.
- Mandatory for every SEV1/SEV2, customer-first detection, incident >2h, repeat of known cause, credit-generating/contract breach, ledger near-miss.
- Draft within 3 business days, review within 5, publish within 10; IC owns delivery, owning manager accountable.
- One template: summary, customer/financial impact, detection source/gap, timeline, response analysis, contributing conditions, what worked, actions.
- Two mandatory questions: why did a customer see this first? why did mitigation take as long as it did?
- Blameless in writing: systems and context, no individual named as cause, never used in performance reviews; HR handles personnel/misconduct separately.
- Every action gets owner, priority, due date, verification method, ticket auto-created; classes containment 7 days, corrective 30, strategic 90; reserve 15% engineering capacity.
- Escalation: manager at +7 days, director at +14, CTO dashboard at +30; overdue P0 can block related releases.
- Weekly Incident Review Board ratifies severity, challenges quality, monitors actions; target 90% closure on time within two quarters.
18. Train and certify incident roles (depends on: 7, 13, 15)
Command is a skill, not a title. Certify before duty; use paid working time.
- All employees: 30-min module on recognizing impact, declaring, finding channel/status page.
- Responder: half-day on severity, escalation, runbook, financial integrity; mandatory before joining rotation.
- Scribe: 2 hours timeline discipline; entry point.
- Incident Commander: 2 days + shadowing; command presence, delegation, severity calls, running room, handover; certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead: 1 day on status writing, customer tiering, legal boundaries, regulator triggers.
- Domain responders demonstrate dashboard, runbook, rollback, failover, access before primary; two shadow shifts; no primary in first 90 days.
- Certification valid 12 months, renewed by simulation; register is audit artifact.
19. Simulations and game days (depends on: 10, 14, 15, 18)
Rehearse process before real SEV1; exercise tooling failure and ledger scenarios.
- Monthly tabletop per group reusing an incident from baseline, rotating commander.
- Quarterly game day in staging or tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes degradation.
- Twice-yearly unannounced paging drill to measure real overnight ack times.
- Annually one combined operational/security exercise and one regulator-notification exercise with Legal/CISO.
- Exercise status page down, chat/paging provider unavailable.
- Never inject uncontrolled change into production ledger; validate backups/restore/RPO/RTO on replicas.
- Every exercise produces tracked actions; publish time-to-commander, time-to-first-update, time-to-mitigation.
20. Pilot on critical path (depends on: 9, 11, 12, 13, 14, 15, 17, 18, 19)
Prove full process on highest-risk surface with willing teams for six weeks before full rollout.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, Support intake, including two of the 12 already on-call teams.
- Activate severity scale, command corps, single paging tool, page budget, status page policy, mandatory postmortems, paid rotations.
- Parallel old paths for one week then cut over; real incidents use new process only.
- Program lead attends every SEV2+ as coach, never shadow commander; review every pilot page within one business day.
- Exit gate: commander named within 5 min in 95% cases, first status update on time, MTTD under 10 min for pilot services, pages halved, no unpaid page, all required postmortems on time, positive sentiment.
- Publish one-page result company-wide.
21. Metrics, dashboards, and review cadence (depends on: 13, 17)
Instrument the process so improvement is visible and audit sees monitoring/review evidence. Report median and p90, never averages alone.
- Response: time to detect/declare/commander/ack/mitigate/resolve split by severity, tier, journey, region, detection source.
- Quality: customer-first rate, status page timeliness, update cadence, missed escalations, role conflicts, alert actionability, out-of-hours pages per person.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat factors.
- Business: availability per journey, error budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, frequency, recovery days, sentiment, attrition.
- Weekly Incident Review Board; monthly reliability review per team; monthly executive review to CEO; quarterly control review with Security/Compliance/Risk; quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, status page history; team scorecards direct help, never individual penalties.
22. Wave rollout to all 28 teams with readiness gates (depends on: 20, 21)
Roll out in four waves by customer risk, each passing an explicit gate rather than a date. Complete all teams by week 24 to leave ~3 months of evidence before audit fieldwork.
- Wave 1 remaining Tier 0, Wave 2 Tier 1, Wave 3 Tier 2, Wave 4 Tier 3/internal.
- Per-team onboarding kit: catalog entry complete, alerts migrated within budget, runbooks at readiness bar, rotation staffed with 6 trained responders or time-limited exception, one IC candidate nominated, one tabletop passed, payroll active.
- Named coach per wave for three weeks; director signs gate.
- Legacy paging paths disabled per wave, not kept as comfort fallback.
- Publish live adoption scoreboard.
- Any team unable to staff fair rotation gets headcount, service reassignment, or explicit executive risk acceptance.
23. Culture, fairness, and continuous feedback (depends on: 4, 9)
Run this in parallel from day one. Engineers judge fairness; executives judge results.
- Repeat the deal in every forum: paid on-call, paged only for owned services, trained commander, real sprint capacity.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
- Weekly office hours first 8 weeks, Slack support channel 4-hour SLA, per-team champion, weekly newsletter with metrics and bad news.
- Incident response contribution in promotion criteria; quarterly awards for best postmortem and biggest noise reduction; public thanks after SEV1.
- Pulse survey at 60, 120, 365 days; if fairness or load red, pause expansion until fixed.
24. SOC 2 evidence design and internal dry-run (depends on: 13, 17, 21, 22)
Make evidence a by-product of operations; test before external auditor.
- Map with Compliance to Trust Services Criteria CC7.2-CC7.5, CC2.2/CC2.3, A1.2; confirm observation window early.
- Evidence set automated and indexed: versioned policies/exception, rotation schedules, compensation activation, incident records with timestamps/roles, paging/ack logs, status history, reportability decisions, postmortems, action closure with verification, training/certification register, drill records, access reviews.
- Internal Audit tests process in months 4 and 6, sampling incidents end-to-end.
- Run formal mock audit month 7 using same evidence populations; include interviews with IC, engineer, Support, Compliance.
- Correct failures through tracked actions, never by editing history.
- Freeze process wording after month 7; any change logged as exception.
25. Risk register and contingency planning (depends on: 1)
Name likely failure modes and pre-commit responses; review monthly with sponsor.
- Too few commander volunteers: make command rostered duty for managers and staff engineers until pool reaches 30.
- Compensation not approved in time: fallback to time-off-in-lieu plus phased stipend, never launch mandatory night on-call uncompensated.
- Tool migration slips: cut scope on incident record layer, never on paging consolidation.
- Noise pruning hides real failure: demote to ticket first, observe 30 days, keep recovery path, monthly missed-detection review.
- Burnout in experienced teams: weekly load monitoring with page caps.
- Major SEV1 mid-rollout: program lead becomes full-time responder, wave schedule slips one wave, sponsor told same day.
- Shared ledger concentration: if blast-radius workstream slips, escalate to board as accepted risk with dated plan.
26. Inspect and adapt; year-two sustainability (depends on: 22, 24)
Prevent decay after audit by revising on data and assigning permanent owners.
- At 90 days live, revise policy using measurements: severity calibration, Layer B/C membership from page data, uncovered shifts, commander burn, missed updates, action closure, survey results.
- Assign permanent owners for policy, paging platform, status page, service catalog, training, metrics.
- Re-baseline targets every six months; shift from lagging to leading metrics: error budget burn, near-miss rate, drill performance.
- Year-two candidates: follow-the-sun coverage, automated mitigation for top recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, completion of ledger blast-radius reduction.
- Keep annual policy review, certification renewal, exercise calendar, board reporting as permanent commitments.
Your answer has these parts:
- "round_summary": one or two sentences on how the round went as a whole.
- "converging": true if the proposals of this round are more similar to each other than those of the previous round, false otherwise.
- "differences": a short list (four items at most) of what still separates the proposals of this round, naming the proposals and the steps concerned.
- "influences": a short list (five items at most) of the flows of ideas of the round: who took what from whom, and what nobody took, naming proposals and steps ("P1 and P3 took P2's CDC sync (steps 2 and 13)").
- "proposals": one entry per proposal of round 2, each with:
- "proposal": its number,
- "assessment": "improved", "worsened", "mixed" or "unchanged" with respect to its previous version ("no_previous_version" if that agent produced nothing in the previous round),
- "what_changed": how it improved or worsened and why: two or three short sentences, or a list,
- "improvements": a list of concrete gains (specific steps, metrics, structure),
- "regressions": a list of concrete losses (dropped steps, vaguer metrics, broken dependencies...),
- "taken": the ideas this proposal visibly adopted from the OTHER proposals of round 1 (not from its own previous version): one entry per idea with "from_proposal" (the number of the proposal it came from), "steps" (the numbers of the steps of that proposal where the idea lives, as listed above; empty if it is not tied to specific steps), "what" (the idea, one sentence) and "why" (how it was used or adapted, one sentence),
- "rejected": the ideas of the OTHER proposals of round 1 that this proposal visibly declined: an explicit contradiction, or a prominent idea it saw and left out while taking the opposite approach. Same fields; "why" gives the evidence (what the proposal does instead). Do not list mere omissions without evidence; an empty list is a valid answer.
[ROUND 3]
[SYSTEM]
You are an expert reviewer of multi-agent planning processes.
Several LLM agents drafted plans for a task, refined them over a number of rounds while seeing each other's proposals, and finally voted for the best one.
Be exhaustive but precise: name concrete steps, ideas and metrics, never generalities. Judge plans by their fitness for the task as stated, their realism, their completeness, the soundness of their order and dependencies, how measurable their success is and how they handle things going wrong.
You are an impartial evaluator, not a chronicler: assess the proposals and the process on their merits, never rationalise what happened or assume that the outcome was right.
After your analysis, answer in the requested structure.
Every text field you write will be read by a busy person who skims. Make it easy to skim: short sentences and short paragraphs; when you name several things, prefer a list to a paragraph, with sub-items when an item has parts, but keep a single fact as a sentence; lead with the point and then the evidence; name proposals and steps by number (P2, step 4); no preamble, no repetition of the question, no closing summary; bold at most one key phrase per item or paragraph. Text fields accept Markdown: a blank line between paragraphs, "- " for lists, **bold**.
[HUMAN]
Task given to the agents: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
This is round 3, a refinement round: every agent received ALL the proposals of round 2 and wrote a new plan, improving on them or taking a different approach. By convention, the previous version of proposal N is proposal N of round 2, written by the same model.
PROPOSALS OF ROUND 2 (the previous versions):
--- PROPOSAL 1 (agent opus5_refine_1, anthropic/claude-opus-5) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (32):
1. Charter the program: one owner, one mandate, funded, dated before the audit
Turn the CEO email into a chartered company program with a single accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the **burning-platform narrative** and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Listening tour, resistance map and the written on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (depends on: 5, 7)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 8, 4)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (depends on: 10, 5)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (depends on: 7, 12, 13)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
18. SLA credit and financial impact workflow (depends on: 6, 16)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Use the credit delta as the standing business case for on-call pay and reserved reliability capacity.
19. Blameless postmortem standard and Incident Review Board (depends on: 6, 7)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
20. Action ownership, reserved capacity and enforcement (depends on: 19, 12)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
21. Runbooks, readiness bar and ledger blast-radius reduction (depends on: 5, 8)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks, and bad runbooks are a real part of the pager resistance. A service must earn the right to page a human.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
22. Training and certification academy (depends on: 7, 13, 15, 19)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Exercise programme: tabletops, game days and unannounced drills (depends on: 22, 12, 21)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
24. Publish Incident Management Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 19, 20)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 11, 12, 21, 22, 9)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
26. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 19, 25)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
27. Alert noise burn-down campaign (depends on: 10, 12, 25)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
28. Wave rollout to all 28 teams with readiness gates (depends on: 25, 26)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
29. Change management, fairness and pager culture (depends on: 4, 9, 25)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. SOC 2 evidence by design, internal testing and mock audit (depends on: 24, 26, 28)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
31. Risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 26, 28, 30)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 2 (agent gpt5.6-sol_refine_2, openai/gpt-5.6-sol) ---
Estimated complexity: high
Success metrics: - By day 7, every suspected major incident uses one incident record, one coordination channel, and a named commander.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes; at least 95% receive command within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and technical coverage.
- By week 18, all 180 services have an owner, tier, tested escalation path, and coverage appropriate to their risk.
- No mandatory night rotation starts before compensation, training, access, runbooks, and staffing controls are active.
- Every direct 24x7 technical rotation has at least six qualified responders or a documented, expiring executive exception.
- No responder is routinely assigned primary duty more often than one week in six or simultaneously assigned to two primary rotations.
- At least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of qualifying SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet their required internal and customer update cadence by month 3.
- Monthly human pages fall from 3,400 alert events to no more than 1,500 by day 90 and no more than 500 actionable pages by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, with no decline in Tier 0/1 detection coverage.
- Average after-hours load remains at or below two pages per responder per week; every sustained breach produces a remediation plan.
- All six monitoring sources route human pages through the single paging platform by week 12; direct legacy paging paths are disabled by week 18.
- 100% of required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 by month 3.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due date with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Monthly contracted availability meets or exceeds 99.95% by month 6, subject to the contractually defined measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- Quarterly on-call surveys reach at least 75% favorable responses on fairness, ownership boundaries, compensation, and sustainability by month 6.
- The month-7 mock audit finds no unowned high-risk control gap, and at least 95% of sampled incidents contain complete operating evidence.
Steps (26):
1. Establish the mandate, owner, funding, and schedule
Launch incident management as a **company operating program within 48 hours**. Give one leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Form a decision group with Engineering, Support, Customer Success, Security, Legal, Compliance, HR, Finance, and Internal Audit.
- Approve the non-negotiables: one severity scale, one human-paging platform, one incident record, paid on-call, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, and incident actions.
- Fund tooling, compensation, training, exercises, and reliability work. Use the $1.3M in credits as the minimum financial comparison.
- Set milestones: interim process by day 7, policy and critical-path pilot by week 6, Tier 0/1 coverage by week 10, company rollout by week 18, internal audit in month 6, and mock audit in month 7.
2. Install a seven-day incident-response floor (depends on: 1)
Do not wait for policy design or tooling migration. Put a minimum process into operation immediately and begin retaining evidence.
- Publish a one-page provisional severity guide, declaration procedure, role card, and communication schedule.
- Create one monitored declaration path through chat, telephone, and the current paging environment.
- Use one incident channel, bridge, timeline document, and naming convention for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced on-call engineers and engineering leaders.
- Require a named commander within 10 minutes. The duty engineering director assumes command if nobody else does.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Compensate interim duty retroactively under the permanent policy.
- Hold a 15-minute daily operational review until the permanent process is live.
3. Create the factual, legal, and audit baseline (depends on: 1)
Build one defensible baseline for process design, executive decisions, and SOC 2 testing. Confirm the auditor's expected Type II observation period immediately.
- Reconstruct all 31 incidents from first customer impact through resolution, communications, credits, postmortem, and corrective action.
- Analyze the two incidents with unclear command minute by minute.
- Identify the missing detection signal for every customer-first incident.
- Inventory the six alert sources, 3,400 monthly alert events, noisy rules, duplicates, missing owners, and missing runbooks.
- Record the current rotations, unpaid work, after-hours load, and teams without coverage.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M credits, 85% noise, and 11 of 64 actions closed.
- Map incident-response controls to applicable SOC 2 criteria with Compliance and the auditor.
- Establish retention, confidentiality, legal-hold, and access requirements for incident evidence.
4. Co-design the fairness contract with engineers (depends on: 1)
Treat pager resistance as a valid design constraint. Make the employment and ownership bargain explicit before expanding on-call.
- Interview representatives from all 28 teams and the Support, Customer Success, Security, and Operations groups.
- Distinguish objections involving unpaid work, unfamiliar code, bad alerts, weak runbooks, sleep disruption, or blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for services they own or have formally accepted, supported by trained command, and given capacity to remove recurring defects.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Measure baseline trust, fairness, fatigue, and psychological safety. Repeat the survey at days 60 and 120, then quarterly.
5. Build the service catalog and criticality model (depends on: 3)
Make a machine-readable catalog the source of truth for routing, escalation, customer impact, and audit evidence. Every production service must have one accountable owner.
- Record the owning team, manager, product capability, escalation policy, communication channel, dashboard, runbook, dependencies, regions, and data stores for all 180 services.
- Map payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, and reporting to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing services whose failure can create material contractual impact.
- Classify Tier 2 as internal or deferrable services, and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional mode, recovery method, and contractual obligations for Tier 0 and Tier 1.
- Give orphan services an owner or approved decommission date within 30 days.
- Maintain separate but coordinated ownership for the ledger application and PostgreSQL platform.
6. Adopt one severity scale and incident lifecycle (depends on: 3, 5)
Use four impact-based levels. Classify on actual or credible customer, financial, security, regulatory, and contractual harm rather than organizational seniority.
- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; both regions impaired; or a core payment journey broadly unavailable. Page every role immediately, open a bridge, notify executives, publish customer status within 15 minutes when applicable, begin legal assessment, and require a postmortem.
- **SEV2 — major:** material payment degradation; settlement deadline at risk; a critical customer or material cohort unavailable; regional impairment with reduced resilience; or likely SLA breach. Page command and technical roles, publish customer status within 30 minutes when customer-visible, and require a postmortem.
- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads; page only when immediate action is necessary.
- **SEV4 — operational event:** no current customer impact and no urgent risk. Create a ticket and handle during normal operations.
- Treat percentages as supporting guardrails, never reasons to under-classify integrity, settlement, security, or contractual risk.
- Anyone may declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Automatically use at least SEV2 posture for suspected ledger-integrity events, cross-team incidents with unknown ownership, and materially unknown impact lasting 15 minutes.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed. Resolution requires stability, backlog recovery, and necessary reconciliation.
7. Define roles, authority, and handoffs (depends on: 6)
Separate command, communications, recordkeeping, and technical repair. One named person must hold command throughout every SEV1 and SEV2.
- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not serve as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, and approved customer language.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- **Subject-Matter Responders:** diagnose and mitigate only services for which they have ownership, access, training, or formally accepted support responsibility.
- **Executive Duty Officer:** removes organizational barriers and makes exceptional business decisions without displacing the commander.
- Security, Legal, Compliance, Vendor Management, and Finance join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Require distinct commander, communications, scribe, and primary technical lead for SEV1. Communications and scribe may combine temporarily for bounded SEV2 incidents.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, and reconciliation requirements for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, with the exact transfer time.
8. Create sustainable 24x7 coverage across 28 teams (depends on: 4, 5, 7)
Use central command coverage and risk-based technical coverage rather than creating 28 fragile night rotations. All services receive a response path, but only critical domains maintain direct overnight technical rotations.
- Create a 24x7 Incident Command corps of 24–30 certified people, with primary and secondary coverage at all times.
- Create 18–24 trained Communications Leads and a similarly sized scribe pool using Support, Customer Operations, Engineering Operations, and qualified managers.
- Maintain a surge roster for simultaneous incidents and a 24x7 Executive Duty Officer schedule.
- Group Tier 0 and Tier 1 ownership into approximately 8–12 coherent domains only where responders have training, access, runbooks, and explicit acceptance.
- Staff each critical domain with primary and secondary responders and at least six qualified people. Target no more than one primary week in six.
- Give Tier 2 and Tier 3 services business-hours coverage plus a maintained manager and director escalation path.
- When lower-tier impact becomes SEV1 or SEV2, central command activates the manager escalation path and obtains the necessary owner.
- Route unknown-owner pages to Duty Command and platform triage temporarily. Log each one as a catalog control failure.
- Do not merge small teams merely for schedule convenience. Provide training, reassign services, add staffing, or decommission unsupported services.
9. Approve compensation and fatigue protections (depends on: 4, 8)
End unpaid on-call before expanding mandatory night coverage. HR, Finance, Payroll, and employment counsel should approve the policy within 14 days.
- Use market-validated weekly bands, initially budgeting approximately $800–$1,200 for Tier 0/1 primary duty and $300–$500 for secondary duty.
- Budget approximately $900–$1,300 for Duty Incident Commander weeks and $400–$800 for Communications Lead or scribe duty, adjusted for actual burden.
- Pay holiday premiums and compensate active after-hours work according to exempt or non-exempt status and applicable federal and New York rules.
- Provide a protected paid recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal sprint commitments by approximately 15% during a primary on-call week.
- Prohibit simultaneous primary assignments, consecutive primary weeks, on-call during leave, and invisible schedule swaps.
- Allow responders to declare themselves temporarily unfit after disruptive night work without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Budget roughly $0.9M–$1.2M annually, then refine using actual rotation count, employment classification, and activation data.
10. Establish one paging and incident system of record (depends on: 3, 5, 8)
Monitoring tools may remain specialized, but all human pages must enter one controlled platform. This removes conflicting schedules and creates one evidence trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page.
- Ingest events from the six existing monitoring tools before disabling their direct paging paths.
- Route pages through service-catalog ownership and deduplicate related events.
- Provide a single declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, severity changes, decisions, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status-page publishing, and conference facilities.
- Require role-based access, MFA, access reviews, and immutable or tamper-evident history.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, or an AWS region.
- Retire each legacy human-paging path only after ownership review, end-to-end tests, and two weeks of verified operation.
11. Enforce an alert-quality contract and page budget (depends on: 3, 5, 10)
Treat every page as a production interface with an owner and an expected action. Noise reduction must not create detection gaps.
- Require every paging rule to identify the service, owning team, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Page only when prompt human judgment or intervention can materially reduce customer, financial, security, or contractual risk.
- Route informational, capacity, and non-urgent infrastructure conditions to dashboards or ticket queues.
- Prefer payment-outcome, settlement-risk, queue-age, and error-budget-burn alerts over raw CPU, memory, pod, or log thresholds.
- Run new paging rules in shadow mode for at least seven days unless an emergency exception is approved.
- Review rules with repeated no-action acknowledgements, low actionability, or excessive firing within two business days.
- Set a load budget of no more than two after-hours pages per responder per week, measured over four weeks.
- Require an owner, compensating detection, and recorded approval before suppressing or deleting a rule.
- Burn down the top 50 noisy rules first. Review missed detections and noise together so teams cannot improve metrics by becoming blind.
12. Detect payment and ledger failures before customers (depends on: 5, 11)
Move detection from host health to customer journeys and financial outcomes. Use internal SLOs with enough headroom to protect the contractual 99.95% SLA.
- Define SLIs and SLOs for payment initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, API access, webhooks, and reporting freshness.
- Run external synthetic transactions through critical payment journeys at least every minute and from paths independent of the production platform.
- Validate each AWS region and expose dependencies that undermine nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed value relative to settlement deadlines, not only queue depth.
- Add tenant and cohort anomaly detection for high-value customers and major payment methods.
- Convert high-priority Support, account-manager, processor, sponsor-bank, and network reports into incident candidates within five minutes.
- Record the detection source for every incident. Treat customer-first detection as a mandatory missed-detection review.
13. Codify detection, escalation, and live execution (depends on: 7, 10)
Create one time-bound path from the first credible signal to named command and mitigation. Notification delivery does not count as human acknowledgement.
- Page the owning critical-domain primary and Duty Incident Commander immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at five minutes, manager at 10, and director at 15.
- Escalate an unclaimed command page to backup commander at five minutes. The Executive Duty Officer assumes command at 10 minutes until a certified handoff occurs.
- Require Support and account managers to use the same declaration path as automated monitoring and engineers.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Begin with a standard statement of severity, known impact, assigned roles, current objective, workstreams, and next update time.
- Freeze unrelated changes during SEV1 and normally during SEV2. Record any exception.
- Prefer reversible mitigation such as rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, or partner rerouting.
- Separate mitigation and diagnosis workstreams when staffing permits.
- Require formal command handoff for long incidents, shift changes, or fatigue. Do not leave an incident unowned during transfer.
- Reconcile the ledger and safely drain backlogs before resolving payment or ledger incidents.
14. Standardize internal incident communications (depends on: 7, 10, 13)
Give responders one working room and stakeholders one controlled information source. Executives must not interrupt the technical command path.
- Create one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue an initial internal brief within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within 15 minutes for SEV1 and 30 minutes for SEV2.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record material decisions and outbound messages in the incident timeline.
15. Standardize status-page and customer communications (depends on: 6, 10, 14)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate impact before root cause is known.
- Publish customer status within 15 minutes of declaring a customer-visible SEV1 and within 30 minutes for customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice promptly after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map status-page components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, data-integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, known workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Give affected strategic accounts direct account-manager outreach within 30 minutes for SEV1 and 60 minutes for SEV2.
- Require account managers to use the approved briefing and prohibit independent technical explanations.
- Provide a customer-facing incident summary within five business days for SEV1 and qualifying SEV2 events.
- Record any legally necessary delay or restriction of public detail, its approver, and the alternative communication plan.
16. Operationalize legal, regulatory, contractual, and credit decisions (depends on: 3, 6, 15)
Some payments incidents start external notification clocks. Make assessment mandatory without assuming that every operational incident is reportable.
- Build a counsel-validated matrix covering applicable NYDFS rules, state breach laws, GLBA or FTC obligations, PCI requirements, money-transmitter obligations, sponsor-bank and network contracts, cyber insurance, and customer contracts.
- Engage Legal and Compliance immediately for every SEV1, security event, and suspected ledger-integrity event.
- Record an initial reportability assessment within one hour for SEV1 and within two hours for other potentially reportable events.
- Record non-reportable decisions as evidence, including facts considered, approver, timestamp, and reassessment trigger.
- Maintain tested 24x7 contacts for counsel, regulators, sponsor banks, networks, insurers, and critical vendors.
- Encode customer-specific notice deadlines and channels in the customer record used by the Communications Lead.
- Have Legal own regulatory text and submission. Keep technical command with the Incident Commander.
- Have Finance calculate affected minutes, delayed value, likely credits, and contractual exposure within five business days.
- Track credits by incident and recurring cause to support reliability investment decisions.
17. Make postmortems mandatory, consistent, and blameless (depends on: 6, 7, 10)
Use one learning standard with fixed deadlines. Keep postmortems separate from performance, misconduct, and disciplinary processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact over two hours, SLA credits, contractual breach, repeated contributing factors, control failures, and ledger-integrity near misses.
- Produce the factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the Incident Commander responsible for the timeline and the owning engineering director accountable for completion.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit analysis of why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the incident commander.
- Describe decisions in the context and information available at the time. Do not name an individual as the root cause.
- Publish broadly useful findings internally while restricting security, privacy, personnel, or privileged material appropriately.
18. Give corrective actions enforceable ownership (depends on: 10, 17)
Treat incident actions as risk commitments, not suggestions. A ticket is not complete until the expected risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within seven days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of team capacity for approved reliability and incident work.
- Escalate overdue high-risk actions to the manager after seven days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit overdue actions to block related releases where the unrepaired condition could reproduce severe impact.
- Verify effectiveness through tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days. Complete, re-plan, or formally accept each risk.
19. Set the service-readiness bar and critical playbooks (depends on: 5, 8, 11, 12)
A critical service must be supportable at 3 a.m. before its team is placed on direct overnight coverage. Existing critical detection must remain active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and data-integrity constraints.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional failure, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document controlled failover, split-brain prevention, replay protection, backlog recovery, and post-recovery reconciliation.
- Exercise critical runbooks at least twice per year and after material changes.
- Block new Tier 0/1 releases and new paging rules when readiness requirements are missing.
- Handle existing gaps through a named owner, compensating control, executive-approved expiry date, and remediation plan rather than disabling detection.
20. Publish policy and certify every response role (depends on: 6, 7, 8, 9, 11, 13, 14, 15, 16, 17, 18)
Convert the operating design into concise, signed documents and practical training. Training and exercises occur during paid working time.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, and exception standards.
- Provide one-page severity, role, authority, escalation, and communication cards inside the incident tool.
- Train all employees to recognize and declare incidents.
- Train all engineers in severity, acknowledgement, evidence preservation, handoff, and financial-integrity precautions.
- Certify responders only after they demonstrate access, dashboards, runbooks, rollback, and escalation competence.
- Certify Incident Commanders through formal instruction, simulation, and at least two shadowed incidents or exercises.
- Train Communications Leads in status writing, account segmentation, contractual clocks, and legal escalation.
- Train scribes in timeline quality and fact-versus-hypothesis labeling.
- Require two shadow shifts before independent primary duty and annual recertification.
- Maintain the training and certification register as operational and audit evidence.
21. Pilot on the payment critical path (depends on: 9, 10, 12, 19, 20)
Run a four-to-six-week pilot across the highest-risk journey before expanding. Use real incidents and exercises to correct the model quickly.
- Include payment orchestration, ledger application, PostgreSQL platform, Kubernetes or cloud platform, API edge, authentication, settlement, and Support intake.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking.
- Run old paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently assuming command.
- Correct critical process or tool defects within 48 hours and update the standard visibly.
- Exit only after 95% timely command assignment, 95% communication compliance, no unpaid pages, complete required postmortems, and at least 50% lower noise.
22. Roll out by customer journey and risk (depends on: 21)
Expand in controlled waves while the interim process remains active company-wide. Complete rollout early enough to accumulate operating evidence before the audit.
- Roll out remaining Tier 0 domains first, then Tier 1, Tier 2, and Tier 3.
- Use two-to-three-week waves with a named program coach and director sign-off.
- Gate each team on catalog ownership, appropriate coverage, compensation, trained responders, tested escalation, alert quality, runbooks, access, and a passed tabletop.
- Require six or more responders only for direct 24x7 technical rotations. Apply business-hours coverage and manager escalation to lower tiers.
- Reschedule failed gates or approve a time-limited executive exception with a compensating control.
- Disable legacy human-paging paths after verified cutover for each wave.
- Publish an internal adoption dashboard by team, service tier, and control gap.
- Finish critical coverage by week 10 and all 28 teams by week 18.
23. Exercise command, communications, regional recovery, and fallbacks (depends on: 19, 20)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first company-wide command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous security and operational events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors where appropriate.
- Conduct at least one unannounced after-hours paging test before the audit.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure time to acknowledgement, command, customer notice, mitigation decision, handoff, and recovery.
- Create tracked actions for every material exercise finding.
24. Measure performance and operate fixed review forums (depends on: 3, 10, 17, 18)
Use a balanced scorecard that exposes weak controls without rewarding hidden incidents or suppressed alerts. Report medians and 90th percentiles rather than averages alone.
- Measure time from first impact to detection, declaration, acknowledgement, command assignment, mitigation, recovery, and resolution.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track pages, actionability, duplicates, after-hours load, missed detections, and page-budget breaches.
- Track postmortem timeliness, action age, closure by due date, verified effectiveness, and recurring contributing factors.
- Track availability by customer journey, error-budget burn, failed or delayed payment value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight activations, recovery days, schedule exceptions, sentiment, and attrition.
- Hold a weekly Incident Review Board for incidents, postmortems, control failures, noisy alerts, and overdue actions.
- Hold a monthly executive reliability review for trends, funding, contractual exposure, and accepted risks.
- Hold a quarterly control and resilience review with Product, Security, Compliance, Risk, and Internal Audit.
- Reconcile incident records monthly against support cases, status history, credits, and customer complaints to detect under-reporting.
25. Prove SOC 2 operating effectiveness before fieldwork (depends on: 20, 22, 23, 24)
Generate audit evidence through normal operation rather than reconstructing it later. Test both control design and consistent execution.
- Maintain approved and versioned policies, exceptions, catalog records, schedules, compensation activation, access reviews, training, incidents, communications, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, non-reportable legal decisions, missed timelines, and exercises in the population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the populations, evidence requests, and interviews expected from the external auditor.
- Record deviations honestly with owners, remediation dates, and compensating controls. Never rewrite historical records.
- Confirm that evidence retention covers the full auditor-defined observation period.
- Brief commanders, responders, Support, and Compliance on the real process without scripting inaccurate answers.
26. Institutionalize improvement and reduce structural risk (depends on: 22, 24, 25)
Prevent the program from decaying after the audit. Use incident evidence to drive permanent ownership and architectural investment.
- Assign permanent owners for policy, service catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review severity thresholds, staffing, compensation, and communication timings annually and after material process failures.
- Recertify command and communications personnel annually.
- Review recurring failure families quarterly and require executive decisions where corrective work repeatedly loses priority.
- Maintain quarterly responder surveys and publish actions addressing fairness, fatigue, and tool friction.
- Report severe incidents, credits, overdue high-risk actions, and resilience investment to the board or risk committee quarterly.
- Treat the shared ledger cluster as a strategic concentration risk. Fund failover assurance, blast-radius reduction, isolation of non-critical readers, and stronger regional independence.
- Evaluate follow-the-sun command or technical coverage using one year of page-load and staffing data.
- Build year-two plans for automated mitigation, deployment safety, graceful degradation, and error-budget release controls.
--- PROPOSAL 3 (agent qwen3.8-max_refine_3, alibaba/qwen3.8-max) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (32):
1. Establish executive mandate, program office, and funding
Turn the CEO email into a **chartered company program** with one accountable owner and authority over all 28 teams.
- Appoint the CTO as executive sponsor and a Head of Reliability as program owner with full-time authority.
- Create a permanent program office: one program lead, one platform engineer, one analyst.
- Form an eight-person steering group: Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, HR. It proposes; the sponsor decides within 48 hours.
- Lock the non-negotiables: one severity scale, one paging platform, one postmortem format, paid on-call, mandatory action tracking, named service ownership.
- Approve budget anchored against the $1.3M in credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs.
- Reserve 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Publish a one-page charter stating incident response is a company operating process, not a per-team choice.
2. Install a seven-day interim command bridge (depends on: 1)
Do not leave the company unprotected while the permanent process is designed. Put a **crude but real** command structure in place within one week.
- Publish a one-page interim severity card and one declaration path: a phone number, a Slack command, and the paging tool, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Direct Support to escalate credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical actions: complete, re-plan, or formally risk-accept, prioritizing ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Build the forensic baseline of incidents, alerts, and money lost (depends on: 1)
Rebuild the facts before designing anything. This becomes both the **design input** and the before picture for the executive and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the specific missing signal. This list drives the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Freeze baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the on-call fairness deal (depends on: 1)
Engineer pushback is the **largest delivery risk**. Treat it as a design constraint, not an attitude problem, and answer it with a written deal.
- Interview all 28 teams plus Support, CS, and Sales in two weeks. Separate the real objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call; they supply the pilot teams and the first commanders.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the incident, nights are paid, noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey covering fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
5. Map compliance, evidence, and audit requirements from day one (depends on: 1)
Design evidence as a **by-product of operations**, not a reconstruction before the auditor arrives. Confirm the SOC 2 observation window immediately.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for communications; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure, training register, drill records, reportability decisions.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records.
- Version and approve all policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has a **named owner**. Build a machine-readable catalog as the single source of truth.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact: Tier 0 (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for each Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation.
- Missing ownership or runbooks is a release blocker for Tier 0 and Tier 1.
7. Adopt the severity scale, declaration rules, and lifecycle (depends on: 3, 6)
Replace judgment calls with a **lookup table**. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter.
- SEV1 (crisis): money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of all roles and executive; bridge in 5 minutes; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 (critical): material degradation of a payment journey, settlement window at risk, a large or strategic customer fully down, SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, mandatory postmortem.
- SEV3 (major, contained): narrow or single-customer impact with a workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, any incident with unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Anyone may declare and nobody is penalised for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected, Declared, Triaged, Mitigated (customer impact ends), Monitoring, Resolved (backlog processed and ledger reconciled), Reviewed.
- Publish a decision tree with 12 worked examples from the real 31 incidents.
8. Define incident roles, authority, and handover discipline (depends on: 7)
Solve nobody-in-charge-for-an-hour by making command **explicit, single-holder, transferable, and logged**. Separate coordination from debugging.
- Incident Commander: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates for SEV1/SEV2.
- Subject-Matter Responders: engineers of the owning team; they mitigate, they do not run the room.
- Executive Duty Officer (SEV1): removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Design the three-layer 24x7 coverage model (depends on: 6, 8)
Do not create 28 night rotations. **Centralise coordination** in a trained corps and keep technical ownership local. This is the structural answer to the pager objection.
- Layer A, Incident Command corps: approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. A paired Comms Lead pool of approximately 18 from Support, CS, and engineering management. A scribe pool used as the training entry point.
- Layer B, critical-path domain rotations: consolidate the 28 teams into 10–12 coherent response domains (ledger, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, PostgreSQL platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- Layer C, everyone else: business-hours on-call with a written after-hours escalation list held by the engineering manager. No night pager.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate in writing: a US-only paid night rotation now, a follow-the-sun cell as a 12-month option, and a first-line triage desk for overnight detection.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2 with automatic failover to secondary, then manager, then director. Delivery to a device does not count as acknowledgement.
10. Approve on-call compensation, labor compliance, and fatigue safeguards (depends on: 9, 4)
Unpaid on-call in New York is both a **retention problem and a legal exposure**. Pay must be live in payroll before any mandatory night rotation starts.
- Indicative scheme: approximately $1,000 per primary 24x7 week, $400 secondary, $250 for business-hours rotations, a separate $1,200 Duty Commander stipend, holiday premiums, approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- HR, Finance, and employment counsel publish amounts, eligibility, tax treatment, and payroll mechanics within 14 days, with explicit FLSA exempt/non-exempt and New York wage-hour treatment documented.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Non-cash elements: on-call counted as delivery load with a 15% sprint reduction, incident leadership in promotion criteria, a documented exemption path for caring or health reasons, and a public quarterly report of on-call load by team.
11. Set the alert quality standard, page budget, and noise burn-down (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes. Make alert quality a **condition of being allowed to page a human**.
- Every paging alert must declare: owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days; test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint for the owning team.
- Auto-quarantine any rule firing more than five times a month without action or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly.
- Target 3,400 to under 500 pages a month with actionability above 75% within six months, with no loss of Tier 0/1 coverage.
12. Build payment-outcome detection and ledger assurance (depends on: 11, 6)
Stop customers telling you first. Detection must be driven by **payment outcomes and ledger truth**, not host metrics.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner and processor notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer detected first becomes a named defect class with a mandatory tracked action.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
13. Consolidate to one paging platform, one incident record, one status page (depends on: 7, 9, 11)
Collapse six alerting tools into a **single operational system of record** so there is one queue, one timeline, and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One-command declaration in Slack that creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is unavailable. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
14. Codify the escalation ladder and five-minute command rule (depends on: 8, 9, 13)
Write one unskippable path from something looks wrong to **someone is in charge**. The default action is never waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately, secondary at 5 minutes, domain manager and Duty Commander at 10, Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension, and partner-bank notification, so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
15. Write the live incident execution doctrine and major-incident playbooks (depends on: 8, 13, 14)
Give responders one short operating procedure from the first minutes through closure. Priority is **limiting customer and financial harm**, not proving a root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down; no command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the single largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Long incidents: fatigue rule and formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window, an explicit handback to the owning team, Support, and Customer Success, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
16. Standardize internal communications protocol (depends on: 8, 13)
Standardise the internal picture so executives, Support, and Sales are informed **without interrupting** the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a behavioural commitment signed by the executive team.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
17. Build customer communications, status page, and account-manager outreach (depends on: 7, 16)
Replace whoever is around with a **timed, owned, pre-approved process**. The Comms Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with subscriptions, email, webhook, and RSS. All 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email.
- Language rules: state impact and the next update time; never speculate on cause, recovery time, data integrity, or blame.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
18. Create the regulatory, partner, and legal notification playbook (depends on: 7, 17)
In payments, some incidents start a **legal clock at detection**. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and recorded even when the answer is not reportable, with evidence, decision-maker, and timestamp.
- Maintain a 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters held under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook once a quarter as part of the exercise programme.
19. Operationalize SLA credit and financial impact workflow (depends on: 7, 17)
Tie incidents to money so severity, credits, and investment decisions **stay honest**, and Finance stops being surprised.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k in year one. Use the delta as the standing business case for on-call pay and reliability capacity.
20. Establish the blameless postmortem standard and Incident Review Board (depends on: 7, 8)
Replace some incidents, various formats with **one mandatory format, fixed deadlines, and a forum with teeth**.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed. Publish a searchable library plus a quarterly top-five recurring causes analysis.
21. Enforce action ownership, capacity reservation, and tracking (depends on: 20, 13)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of **customer commitments** and give them real capacity.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance; they can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives. Target: 90% closed by due date within two quarters.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
22. Build the training and certification academy (depends on: 8, 14, 16, 20)
Command is a skill, not a title. **Certify before assigning duty**, and use paid working time for all of it.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days + shadowing): command presence, delegation, decision-making under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Comms Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum; never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
23. Run the exercise programme: tabletops, game days, and unannounced drills (depends on: 22, 13, 15)
The process must meet a **simulated SEV1** before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Never inject uncontrolled change into the production ledger. Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-correct-mitigation.
24. Publish Incident Management Policy v1 (depends on: 7, 8, 9, 10, 11, 14, 16, 17, 18, 20, 21)
Collapse the design into a document people will **actually open mid-outage**, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
25. Pilot on the payments critical path (depends on: 24, 12, 13, 15, 22, 10)
Prove the process on the **highest-risk surface** with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, positive on-call sentiment.
- Publish a one-page result to the whole company.
26. Run the alert noise burn-down campaign (depends on: 11, 13, 25)
Run noise reduction as a **visible, quota-driven campaign** in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so noise reduction never becomes blindness. Review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 to 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
27. Establish metrics, dashboards, review cadence, and anti-gaming (depends on: 13, 20, 25)
Instrument the process itself so improvement is visible and the auditor sees **evidence of monitoring and review**. Report median and 90th percentile, never averages alone.
- Response: time to detect, declare, commander, acknowledgement, mitigate, resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
28. Wave rollout to all 28 teams with readiness gates (depends on: 25, 27)
Roll out in four waves of six to eight teams every two to three weeks, ordered by **customer risk**. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. Failing teams are rescheduled, not waived. Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Publish a live adoption scoreboard.
29. Run change management, fairness, and pager culture programme (depends on: 4, 10, 25)
Run this from day one in parallel. Engineers judge the process on **fairness**; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a Slack support channel with a four-hour answer SLA, a per-team champion, and a short weekly rollout newsletter with metrics and bad news included.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
30. Build SOC 2 evidence by design, internal testing, and mock audit (depends on: 24, 27, 28)
Design the evidence as a **by-product of doing the work**. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. Auditors respond better to a documented exception log than to a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any change as an exception.
31. Maintain the program risk register and contingencies (depends on: 1)
Name the ways this programme fails and **pre-commit the response**. Review it monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as an accepted risk with a dated plan.
32. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 27, 28, 30)
Guard against the classic failure: the process **decays once the audit is signed**. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 4 (agent grok4.6_refine_4, xai/grok-4.6) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days and reviewed in 5, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
Steps (26):
1. Charter the program and start the audit clock
Convert the CEO email into a company operating process with one owner, a budget, and an observation window that starts this week.
Incident response is no longer a per-team choice.
- Name the CTO as sponsor and a Head of Reliability as the single accountable owner, with a three-person program office.
- Form a small decision group: Engineering, SRE, Support, Customer Success, Security, Legal, Finance, and HR. The sponsor decides within 48 hours.
- Lock non-negotiables: one severity scale, one paging path, one incident record, one postmortem format, mandatory action tracking, and **paid on-call**.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, and reserved engineering capacity against the $1.3M in SLA credits.
- Timeline: operating floor in 7 days, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, mock audit in month 7, audit in month 8.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for new tools or the final policy. Put a crude but real process in place this week so the next outage has a named commander.
This week is the first audit evidence.
- Publish a one-page interim severity card and one declaration path: Slack command, phone number, and existing pagers, all reaching the same duty person.
- Staff interim primary and backup Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in the incident channel.
- Tell Support to declare from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline (depends on: 1)
Rebuild the facts before locking design. This is both the design input and the before picture for the CEO and the auditor.
Freeze the numbers so they cannot drift during design.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the specific signal that was missing. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisy rules and every rule with no owner or runbook.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11 of 64 actions closed, on-call in 12 of 28 teams.
4. Map resistance and publish the fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
Publish the deal in writing before any new mandatory pager is assigned.
- Interview all 28 teams plus Support, Customer Success, and Sales in two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the **fairness contract**: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and money-path tiers (depends on: 3)
You cannot page the right person across 180 services until every service has one owner. A wrong owner recreates the pager objection.
The catalog is the single source of truth for paging, impact, status-page components, and audit.
- Assign one accountable team per service, plus engineering manager, Slack channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to services, data stores, regions, and third parties, including sponsor banks and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform.
- Orphan services get an owner in 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
6. Lock severity levels and what each one triggers (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by who reported it or how hard the fix looks.
Anyone may declare. Nobody is punished for over-declaring. Only the Incident Commander may downgrade, with the evidence recorded.
- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 major, contained: narrow or single-customer impact with a workaround. Owning team leads; business-hours comms; postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages.
- Attach a financial-integrity or security flag to any severity. The flag forces dual-control, Legal, and the regulatory checkpoint without inventing a fifth level.
- Auto-escalate: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident all become at least SEV2.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer impact ends) → Monitoring → Resolved (backlog processed and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged.
Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit spend. Retains command when a VP joins.
- Communications Lead: single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. A human validates for SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers.
- Command claimed within 5 minutes, stated in channel. Distinct people for command, comms, and technical lead at SEV1 and SEV2. Every handover is announced verbally and in writing with the exact time.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
8. Staff 24x7 with a command corps, not 28 night rotations (depends on: 5, 7)
Do not create 28 night rotations. That is what engineers are rejecting.
Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified volunteers from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to seven months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate 28 teams into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat Lisbon or APAC follow-the-sun as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
9. Pay on-call, meet New York labor rules, and cap fatigue (depends on: 4, 8)
Unpaid on-call in New York is a retention problem and a legal exposure. Pay must be in payroll before any new mandatory night rotation starts.
Publish the numbers. Then ask people to sign up.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
10. Set the alert-quality bar and a hard page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them.
Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers do (depends on: 5, 10)
The goal is simple: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail.
Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One Slack command creates the channel and bridge, pages the commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, Slack, or the identity provider is down. Test weekly.
- Set a hard date after which pages outside this tool create no on-call obligation.
13. Write the five-minute path from signal to command (depends on: 7, 8, 12)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting.
If nobody claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Codify live execution, readiness, and ledger playbooks (depends on: 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving a root cause.
A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis.
- Separate diagnosis and mitigation workstreams once there are enough responders. Decisions are stated aloud and written down. No command by direct message.
- Payments-specific guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers.
15. Run one communications clock for internals, customers, and regulators (depends on: 6, 7, 12)
Replace whoever is around with one timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is not reportable.
- Obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
16. Tie incidents to SLA credits and true financial cost (depends on: 6, 15)
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
Credit calculation is an output of the incident record, not a negotiation.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
17. Make blameless postmortems mandatory and consistent (depends on: 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth.
The discipline lives in the deadlines and the review, not the template.
- Mandatory for every SEV1 and SEV2, any incident detected by a customer first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. Personnel and misconduct processes stay separate and are invoked only by HR.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items.
- Facilitators are trained and are never the commander of the incident being reviewed.
18. Track actions as risk commitments with reserved capacity (depends on: 12, 17)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
Closing a ticket without evidence of effectiveness does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Target 90% of high-priority actions closed by due date within two quarters.
19. Train and certify every role before independent duty (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
The certification register is an audit artefact.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
20. Rehearse with tabletops, game days, and night drills (depends on: 12, 14, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
21. Publish Policy v1 and the signed fairness contract (depends on: 6, 7, 8, 9, 10, 13, 15, 17, 18)
Collapse the design into a document people will actually open mid-outage, and make it official.
Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and comms timings.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in Slack.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
22. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 19, 21)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse. Publish a one-page result to the whole company.
23. Instrument metrics, reviews, and anti-gaming (depends on: 12, 17, 22)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review.
Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
24. Roll out in risk-ordered waves with readiness gates (depends on: 22, 23)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave, not as a background hope. Give every team its noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
25. Produce SOC 2 evidence by operating, then mock-audit (depends on: 21, 23, 24)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
Documented exceptions beat a claim of perfection. Never edit historical records to look clean.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training and certification register, drill records, access reviews of the incident and paging tools.
- Internal Audit or an independent control owner tests the process in month 4 and month 6, sampling incidents end to end from first signal to verified action.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions. Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (depends on: 23, 24, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
Name the ways this programme fails and pre-commit the response.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 5 (agent deepseek-v4-pro_refine_5, deepseek/deepseek-v4-pro) ---
Estimated complexity: high
Success metrics: - Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90 and under 10% by month 6.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, no loss of Tier 0/1 detection.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% Tier 0/1 services have named owning team, tier, escalation policy, dashboard and runbook by day 30; all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10-12 domain rotations of at least six trained responders.
- 30 certified incident commanders and 18 certified communications leads active.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication.
- 100% required postmortems drafted in 3 business days, reviewed in 5, published in 10 from month 2.
- Postmortem action closure rises from 17% to 90% closed by due date within two quarters; all 53 legacy actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k annualised within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Regulatory reportability assessed and recorded within 1 hour for 100% SEV1/security SEV2 including not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two cross-company exercises completed before audit, each producing tracked actions.
- Month-7 mock audit finds no unowned high-risk gap, with complete evidence for at least 95% sampled incidents; SOC 2 Type II incident response controls pass zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days; no increase in attrition among responders.
Steps (26):
1. Executive mandate, budget, and governance
Secure a written CTO/CEO charter making incident management a company operating process, not a team option. Name one accountable Head of Reliability and a small steering group with authority to decide within 48 hours.
- Lock non-negotiables: one severity scale, one paging tool, one postmortem format, paid on-call, mandatory action tracking, and protected engineering capacity.
- Approve budget against the $1.3M annual credits: tooling $150-250k/yr, on-call compensation, training, and 2-3 program FTEs.
- Set timeline: interim process in week 1, design weeks 1-6, pilot weeks 7-12, rollout weeks 13-24, mock audit month 7, SOC 2 at month 8.
- Freeze current baselines: 31 incidents, 22-min MTTD, 40% customer-first detection, 3h10 MTTM, $1.3M credits, 3,400 alerts/month at 85% noise, 11/64 actions closed.
2. Interim 7-day command (depends on: 1)
Put a crude but real process in place immediately so no incident remains unowned while the permanent design proceeds.
- Publish a one-page interim severity card, one declaration path (phone, Slack command, pager), and one incident channel/bridge/timeline naming convention.
- Staff an interim 24x7 duty commander with primary and backup from engineering managers and senior SREs; compensate retroactively under final policy.
- Require a named incident commander within 10 minutes of any suspected major incident, announced in channel.
- Let Support and Customer Success declare incidents from credible customer reports without waiting for engineering confirmation.
- Triage the 53 open historical actions in 30 days, prioritizing ledger integrity, duplicate payment, regional failover, security and detection gaps.
- Hold a daily 15-minute operations review until the permanent process is live.
3. Forensic baseline and alert estate analysis (depends on: 1)
Reconstruct the true before picture from the last 12 months; it drives design and serves as audit baseline.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact start/detect/declare/commander/mitigate/resolve, who led, credits paid, contributing factors.
- For every customer-first detection, name the missing signal; this becomes the detection backlog.
- Reconstruct the two `nobody in charge` incidents minute by minute; use them as the burning-platform narrative and test case.
- Profile the 3,400 monthly alerts by tool, team, rule, outcome; identify top 50 noisy rules and every rule with no owner or runbook.
- Freeze all baseline metrics in one signed document for the executive and the auditor.
4. Listening tour and fairness contract (depends on: 1, 3)
Treat engineer pushback as the main delivery risk and convert it into a written deal that removes the objection to carrying a pager for other teams' code.
- Interview all 28 teams plus Support, CS and Sales in two weeks to separate objections: unpaid work, nights, unfamiliar code, missing runbooks, or fear of blame.
- Harvest practices from the 12 teams already on-call; they supply pilot teams and first commanders.
- Publish the deal: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- Recruit 10-15 credible engineers as a design working group so the process is co-authored.
- Run a baseline sentiment survey and repeat at 60, 120 and 365 days.
5. Service ownership catalog and criticality tiers (depends on: 3)
Create a machine-readable catalog as the single source of truth for paging, routing, status page components, and audit evidence.
- Assign one accountable team, engineering manager, Slack channel, escalation policy, dashboards, runbook link and dependency list per service.
- Tier by business impact: Tier 0 money movement/ledger/auth/shared PostgreSQL; Tier 1 customer-facing degradable; Tier 2 internal/batch; Tier 3 non-critical.
- Map customer journeys (initiate, authorize, settle, reconcile, report, onboard) to services, data stores, regions, third parties.
- Give orphan services an owner within 30 days or a decommission date approved by the sponsor; Tier 0 without an owner is an executive escalation.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform.
6. Severity scale, declaration rules, and automatic triggers (depends on: 3, 5)
Adopt one severity scale as a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade with recorded rationale. Start high when uncertain.
- SEV1: money moved wrongly/duplicated/lost, ledger integrity in doubt, confirmed security/data exposure, both regions impaired, payment processing halted. Full role page, bridge in 5 min, status page in 15 min, regulator assessment within 1 hour, mandatory postmortem.
- SEV2: material degradation of a payment journey, settlement window at risk, large/strategic customer fully down, SLA breach likely. Commander and SMEs paged, status page in 30 min, mandatory postmortem.
- SEV3: narrow or single-customer impact with workaround. Owning team leads, business-hours comms, postmortem on request.
- SEV4: no customer impact; ticket only, never pages.
- Auto-escalations: any ledger cluster incident, any SEV3 open >2h, any unknown impact after 15 min, any cross-team incident become at least SEV2.
- Publish a decision tree with 12 worked examples from the 31 real incidents.
7. Incident roles, authority, and handoffs (depends on: 6)
Codify roles so coordination never depends on seniority or heroics. The Incident Commander coordinates; responders fix only services they own.
- IC: owns severity, priorities, roles, cadence, closure; pre-authorized to freeze deploys, rollback, disable features, shift traffic, invoke failover, put ledger read-only, commit spend; does not type in terminals; keeps command when VP joins.
- Communications Lead: sole author for status page, account managers, executives, and hand-off to Legal for regulators; speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, state changes; human validates for SEV1/SEV2.
- Subject-Matter Responders: diagnose and mitigate only their owned services.
- Executive Duty Officer (SEV1): removes obstacles, shields IC from exec questions, owns board/regulator escalation; does not take command unless formally transferred.
- Rule: IC claimed within 5 min and announced; distinct people for IC, Comms, technical lead at SEV1/SEV2; every handover announced verbally and in writing with time.
- Financial controls survive: IC coordinates ledger recovery but cannot bypass dual control, reconciliation or privileged access.
8. 24x7 command and responder coverage model (depends on: 5, 7)
Do not create 28 fragile night rotations. Centralize coordination in a trained command corps, keep technical ownership local.
- Layer A: incident command corps of ~30 certified volunteers with primary/secondary 24x7; paired Comms Lead pool ~18 and scribe pool as entry.
- Layer B: consolidate 28 teams into 10-12 critical-path domains (ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes, data/reporting, integrations) each with 24x7 primary+secondary and minimum six trained responders.
- Layer C: all other teams business-hours on-call with written after-hours escalation lists held by managers.
- Platform on-call is safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate follow-the-sun only as a 12-month option, not year-1 dependency.
- Acknowledgement discipline: 5 min at SEV1/SEV2 with automatic failover.
9. Paid on-call and fatigue safeguards (depends on: 8, 4)
Unpaid on-call in New York is a retention and legal risk. Pay must be in payroll before any mandatory night rotation starts.
- Indicative: ~$1,000 per primary 24x7 week, secondary ~$400, business-hours ~$250, duty commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond one hour.
- Mandatory recovery: paid recovery day after >2h overnight work, SEV1, or qualifying SEV2; managers arrange coverage.
- HR/Legal/Finance publish amounts, tax treatment, FLSA/NY wage-hour status within 14 days.
- Load rules: no primary more often than 1 in 6, no consecutive primary/secondary weeks, no primary on two rotations, no on-call during PTO, annual cap.
- If a rotation averages >2 out-of-hours pages per person per week for 4 weeks, trigger staffing or alert-remediation review.
- Count on-call as ~15% delivery load; document exemption path for health/caring without career penalty.
10. Service readiness bar and runbooks (depends on: 5, 8)
A service must earn the right to page a human at 3 a.m. Define a minimum readiness bar and major incident playbooks before any Tier 0/1 service goes live with on-call.
- Require for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switch, escalation contacts, data-loss/latency impact.
- Write playbooks for top failures: shared PostgreSQL failure/corruption, cross-region failover, Kubernetes control-plane loss, processor/sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, duplicate payment.
- Treat ledger cluster as largest structural risk: documented failover, read-only degraded mode, reconciliation after recovery, executive-signed RPO/RTO.
- Start a parallel workstream on blast-radius reduction: tenant/function partitioning, read replicas, isolation of non-critical readers.
- Enforcement: no readiness sign-off, no paging at night unless manager accepts a dated exception in writing.
- Runbooks are peer-reviewed, version-controlled, marked stale if not exercised twice a year.
11. Alert quality standard and page budget (depends on: 3, 5)
Make alert quality a condition of paging a human. Burn down noise deliberately rather than by mass silencing.
- Every paging alert must declare owning team, affected service, customer/SLO impact, expected action, dashboard, runbook, dedup key, severity mapping, escalation policy. Anything failing becomes a ticket.
- Page on symptoms of customer harm (SLO burn rate, payment success rate, queue age vs settlement deadlines), not raw CPU/memory.
- Run new alerts in shadow mode 7 days; test firing and recovery.
- Set page budget of 2 out-of-hours pages per person per week; breach blocks new alert creation and triggers tuning sprint.
- Auto-quarantine rules firing >5 times/month without action or >70% no-action acks; return to owner with correction deadline.
- Never silently disable: verify compensating detection, record decision, name owner, review missed detections monthly.
- Target 3,400 to <500 pages/month, actionability >75% within six months, no loss of Tier 0/1 detection.
12. Detection uplift on money path (depends on: 5, 11)
Stop customers telling you first by detecting on payment outcomes and ledger truth.
- Define SLOs and business SLIs per customer journey: initiation success, authorization latency, settlement timeliness, reconciliation break rate, API availability, reporting freshness; set internal targets stricter than 99.95%.
- Run external synthetic end-to-end payments every 60 seconds from both regions, covering all critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate detection, replication lag, failover-readiness, settlement-window countdown.
- Add per-customer anomaly detection for top 100 accounts; five similar support tickets in 10 minutes auto-creates triage incident.
- Route partner and processor notifications into declaration path within 5 minutes.
- Record detection source on every incident; treat `customer detected first` as a named defect class with mandatory action.
13. Consolidate to one paging and incident platform (depends on: 6, 8, 11, 12)
Collapse six alerting tools into one operational system of record: one queue, one timeline, one audit trail, without creating a monitoring gap.
- Select one paging/scheduling platform, one incident record, one hosted status page; time-box selection to 2 weeks.
- Ingest all six sources first, deduplicate/correlate; retire legacy path only after named owners, successful end-to-end test, and 2 weeks verified operation.
- One-command declaration in Slack creates channel/bridge, pages commander, sets severity, opens timeline, starts clock.
- Route every page through service catalog: service label -> owning domain schedule -> escalation policy.
- Capture evidence automatically: declaration, acks, role assignments, severity changes, decisions, comms, mitigated/resolved times, postmortem link; retain 12 months.
- Test out-of-band SMS/phone paging, mobile fallback, offline runbook weekly; ensure works during one AWS region/chat/identity provider failure.
- Set hard date after which pages outside this tool create no on-call obligation.
14. Escalation paths and acknowledgement SLAs (depends on: 7, 8, 13)
Write one unskippable path from signal to named commander in under 5 minutes; default action is never waiting.
- Converge all entry points on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline.
- SEV1/SEV2 ladder: owning primary immediately -> secondary at 5 min -> domain manager + Duty Commander at 10 -> Executive Duty Officer at 15. SEV2 requires command assigned within 15 min.
- If no one claims command in 5 min, platform assigns and announces it; assignee may hand over but not decline.
- Commander may page any domain's on-call directly with 10-min ack obligation; this reciprocity makes single-team ownership viable.
- If ownership unclear after 10 min, commander keeps incident and names temporary owner; missing catalog entry logged as control defect.
- Pre-authorize regional failover, ledger read-only, payment suspension, partner-bank notification so IC never waits for an executive.
15. Standardize internal and customer communications (depends on: 6, 7, 13)
Replace whoever is around with a timed, owned, pre-approved process; the Communications Lead is single author and never writes from scratch under pressure.
- Internal: one working channel/bridge plus one read-only broadcast for executives/Support/Sales; cadence SEV1 15-30 min, SEV2 60 min even if nothing changed; fixed template with impact, action, next update, IC/Comms names.
- Executives ask questions only of Executive Duty Officer; publish this as signed behavior rule.
- Status page: SEV1 within 15 min, SEV2 within 30 min; updates every 30/60 min; resolution notice within 30 min of verified recovery; customer-facing summary within 5 business days for SEV1/qualifying SEV2.
- Pre-approve 12-15 templates with Legal for degradation, settlement delay, API errors, regional failure, data-integrity investigation, security event.
- Top 100 accounts get named account manager call/email within 30 min of SEV1 with briefing pack; all 2,100 subscribed by default.
- Language rules: state impact and next update; never speculate on cause, recovery time, data integrity or blame.
16. Regulatory, partner, and financial impact workflows (depends on: 6, 15)
In payments some incidents start a legal clock at detection. Build obligation assessment into the process and tie incidents to money.
- Legal/Compliance produce obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC, PCI, money-transmitter, sponsor-bank/card-network windows, cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and security-related SEV2 within 1 hour, recorded even when `not reportable`, with decision-maker and evidence.
- Maintain 24x7 contact matrix for regulators, sponsor banks, networks, outside counsel, insurer.
- Agree availability measurement method per contract; compute affected minutes per customer from incident record and journey telemetry; propose credit schedule within 5 business days.
- Attribute credits to root-cause families; track failed payment volume, delayed value, reconciliation breaks; target annual credits from $1.3M to under $400k.
17. Blameless postmortem standard and action tracking (depends on: 6, 7)
Replace various formats with one mandatory format, fixed deadlines, and enforceable actions. Closing a ticket without evidence does not close the action.
- Mandatory for every SEV1/SEV2, customer-first detection, incident >2h, repeat of known cause, credit-generating/contract breach, ledger near-miss.
- Draft within 3 business days, review within 5, publish within 10; IC owns delivery, owning manager accountable.
- One template: summary, customer/financial impact, detection source/gap, timeline, response analysis, contributing conditions, what worked, actions.
- Two mandatory questions: why did a customer see this first? why did mitigation take as long as it did?
- Blameless in writing: systems and context, no individual named as cause, never used in performance reviews; HR handles personnel/misconduct separately.
- Every action gets owner, priority, due date, verification method, ticket auto-created; classes containment 7 days, corrective 30, strategic 90; reserve 15% engineering capacity.
- Escalation: manager at +7 days, director at +14, CTO dashboard at +30; overdue P0 can block related releases.
- Weekly Incident Review Board ratifies severity, challenges quality, monitors actions; target 90% closure on time within two quarters.
18. Train and certify incident roles (depends on: 7, 13, 15)
Command is a skill, not a title. Certify before duty; use paid working time.
- All employees: 30-min module on recognizing impact, declaring, finding channel/status page.
- Responder: half-day on severity, escalation, runbook, financial integrity; mandatory before joining rotation.
- Scribe: 2 hours timeline discipline; entry point.
- Incident Commander: 2 days + shadowing; command presence, delegation, severity calls, running room, handover; certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead: 1 day on status writing, customer tiering, legal boundaries, regulator triggers.
- Domain responders demonstrate dashboard, runbook, rollback, failover, access before primary; two shadow shifts; no primary in first 90 days.
- Certification valid 12 months, renewed by simulation; register is audit artifact.
19. Simulations and game days (depends on: 10, 14, 15, 18)
Rehearse process before real SEV1; exercise tooling failure and ledger scenarios.
- Monthly tabletop per group reusing an incident from baseline, rotating commander.
- Quarterly game day in staging or tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes degradation.
- Twice-yearly unannounced paging drill to measure real overnight ack times.
- Annually one combined operational/security exercise and one regulator-notification exercise with Legal/CISO.
- Exercise status page down, chat/paging provider unavailable.
- Never inject uncontrolled change into production ledger; validate backups/restore/RPO/RTO on replicas.
- Every exercise produces tracked actions; publish time-to-commander, time-to-first-update, time-to-mitigation.
20. Pilot on critical path (depends on: 9, 11, 12, 13, 14, 15, 17, 18, 19)
Prove full process on highest-risk surface with willing teams for six weeks before full rollout.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, Support intake, including two of the 12 already on-call teams.
- Activate severity scale, command corps, single paging tool, page budget, status page policy, mandatory postmortems, paid rotations.
- Parallel old paths for one week then cut over; real incidents use new process only.
- Program lead attends every SEV2+ as coach, never shadow commander; review every pilot page within one business day.
- Exit gate: commander named within 5 min in 95% cases, first status update on time, MTTD under 10 min for pilot services, pages halved, no unpaid page, all required postmortems on time, positive sentiment.
- Publish one-page result company-wide.
21. Metrics, dashboards, and review cadence (depends on: 13, 17)
Instrument the process so improvement is visible and audit sees monitoring/review evidence. Report median and p90, never averages alone.
- Response: time to detect/declare/commander/ack/mitigate/resolve split by severity, tier, journey, region, detection source.
- Quality: customer-first rate, status page timeliness, update cadence, missed escalations, role conflicts, alert actionability, out-of-hours pages per person.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat factors.
- Business: availability per journey, error budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, frequency, recovery days, sentiment, attrition.
- Weekly Incident Review Board; monthly reliability review per team; monthly executive review to CEO; quarterly control review with Security/Compliance/Risk; quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, status page history; team scorecards direct help, never individual penalties.
22. Wave rollout to all 28 teams with readiness gates (depends on: 20, 21)
Roll out in four waves by customer risk, each passing an explicit gate rather than a date. Complete all teams by week 24 to leave ~3 months of evidence before audit fieldwork.
- Wave 1 remaining Tier 0, Wave 2 Tier 1, Wave 3 Tier 2, Wave 4 Tier 3/internal.
- Per-team onboarding kit: catalog entry complete, alerts migrated within budget, runbooks at readiness bar, rotation staffed with 6 trained responders or time-limited exception, one IC candidate nominated, one tabletop passed, payroll active.
- Named coach per wave for three weeks; director signs gate.
- Legacy paging paths disabled per wave, not kept as comfort fallback.
- Publish live adoption scoreboard.
- Any team unable to staff fair rotation gets headcount, service reassignment, or explicit executive risk acceptance.
23. Culture, fairness, and continuous feedback (depends on: 4, 9)
Run this in parallel from day one. Engineers judge fairness; executives judge results.
- Repeat the deal in every forum: paid on-call, paged only for owned services, trained commander, real sprint capacity.
- Publicly close the two historical nobody-in-charge incidents with what would be different now.
- Weekly office hours first 8 weeks, Slack support channel 4-hour SLA, per-team champion, weekly newsletter with metrics and bad news.
- Incident response contribution in promotion criteria; quarterly awards for best postmortem and biggest noise reduction; public thanks after SEV1.
- Pulse survey at 60, 120, 365 days; if fairness or load red, pause expansion until fixed.
24. SOC 2 evidence design and internal dry-run (depends on: 13, 17, 21, 22)
Make evidence a by-product of operations; test before external auditor.
- Map with Compliance to Trust Services Criteria CC7.2-CC7.5, CC2.2/CC2.3, A1.2; confirm observation window early.
- Evidence set automated and indexed: versioned policies/exception, rotation schedules, compensation activation, incident records with timestamps/roles, paging/ack logs, status history, reportability decisions, postmortems, action closure with verification, training/certification register, drill records, access reviews.
- Internal Audit tests process in months 4 and 6, sampling incidents end-to-end.
- Run formal mock audit month 7 using same evidence populations; include interviews with IC, engineer, Support, Compliance.
- Correct failures through tracked actions, never by editing history.
- Freeze process wording after month 7; any change logged as exception.
25. Risk register and contingency planning (depends on: 1)
Name likely failure modes and pre-commit responses; review monthly with sponsor.
- Too few commander volunteers: make command rostered duty for managers and staff engineers until pool reaches 30.
- Compensation not approved in time: fallback to time-off-in-lieu plus phased stipend, never launch mandatory night on-call uncompensated.
- Tool migration slips: cut scope on incident record layer, never on paging consolidation.
- Noise pruning hides real failure: demote to ticket first, observe 30 days, keep recovery path, monthly missed-detection review.
- Burnout in experienced teams: weekly load monitoring with page caps.
- Major SEV1 mid-rollout: program lead becomes full-time responder, wave schedule slips one wave, sponsor told same day.
- Shared ledger concentration: if blast-radius workstream slips, escalate to board as accepted risk with dated plan.
26. Inspect and adapt; year-two sustainability (depends on: 22, 24)
Prevent decay after audit by revising on data and assigning permanent owners.
- At 90 days live, revise policy using measurements: severity calibration, Layer B/C membership from page data, uncovered shifts, commander burn, missed updates, action closure, survey results.
- Assign permanent owners for policy, paging platform, status page, service catalog, training, metrics.
- Re-baseline targets every six months; shift from lagging to leading metrics: error budget burn, near-miss rate, drill performance.
- Year-two candidates: follow-the-sun coverage, automated mitigation for top recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, completion of ledger blast-radius reduction.
- Keep annual policy review, certification renewal, exercise calendar, board reporting as permanent commitments.
Step-level differences computed by the tool:
Proposal 1 vs the previous-round proposal it resembles most (opus5_refine_1): 31 steps kept, added ['Audit, legal and evidence scoping in week one', 'Change intelligence and deployment safety on the critical path', 'Support and account management as a detection and intake tier', 'SLA credit workflow, true cost model and incentive guardrails'], removed ['SLA credit and financial impact workflow']
Proposal 2 vs the previous-round proposal it resembles most (gpt5.6-sol_refine_2): 18 steps kept, added ['Charter the program and fund immediate action', 'Publish the on-call fairness contract', 'Codify the live incident lifecycle', 'Implement the detection and escalation ladder', 'Tie incidents to SLA credits and financial exposure', 'Establish mandatory blameless postmortems', 'Publish the signed policy and control set', 'Instrument the scorecard and review forums', 'Roll out by risk with readiness gates', 'Inspect, adapt, and institutionalize ownership'], removed ['Establish the mandate, owner, funding, and schedule', 'Co-design the fairness contract with engineers', 'Codify detection, escalation, and live execution', 'Operationalize legal, regulatory, contractual, and credit decisions', 'Make postmortems mandatory, consistent, and blameless', 'Roll out by customer journey and risk', 'Measure performance and operate fixed review forums', 'Institutionalize improvement and reduce structural risk']
Proposal 3 vs the previous-round proposal it resembles most (grok4.6_refine_4): 21 steps kept, added ['Map compliance, evidence, and audit requirements from day one', 'Adopt the severity scale, declaration rights, and incident lifecycle', 'Codify the five-minute escalation path and live execution doctrine', 'Set the service-readiness bar and write major-incident playbooks', 'Enforce action ownership, reserved capacity, and tracking', 'Publish Incident Management Policy v1 and the exception register', 'Run change management, fairness, and pager culture from day one', 'Maintain the program risk register and pre-committed contingencies'], removed ['Lock severity levels and what each one triggers', 'Write the five-minute path from signal to command', 'Codify live execution, readiness, and ledger playbooks', 'Track actions as risk commitments with reserved capacity', 'Publish Policy v1 and the signed fairness contract']
Proposal 4 vs the previous-round proposal it resembles most (grok4.6_refine_4): 18 steps kept, added ['Lock severity levels, integrity flags, and the incident lifecycle', 'Staff 24x7 with command, domains, and a Duty Triage Desk', 'Codify escalation, the five-minute command rule, and vendor incidents', 'Operationalize regulatory notice, partner clocks, and SLA credits', 'Publish Policy v1 and the exception register', 'Burn down alert noise and roll out by risk with gates', 'Operate the fairness and culture program in parallel', 'Prove SOC 2 operating effectiveness before fieldwork'], removed ['Map resistance and publish the fairness contract', 'Lock severity levels and what each one triggers', 'Staff 24x7 with a command corps, not 28 night rotations', 'Write the five-minute path from signal to command', 'Tie incidents to SLA credits and true financial cost', 'Track actions as risk commitments with reserved capacity', 'Roll out in risk-ordered waves with readiness gates', 'Produce SOC 2 evidence by operating, then mock-audit']
Proposal 5 vs the previous-round proposal it resembles most (opus5_refine_1): 29 steps kept, added ['Charter the program and start the audit clock'], removed ['Charter the program: one owner, one mandate, funded, dated before the audit', 'SLA credit and financial impact workflow', 'Runbooks, readiness bar and ledger blast-radius reduction']
Origin of the steps of the new proposals, matched by title by the tool (evidence for "taken"; ideas can also travel without a matching title):
Proposal 1: 30 of its 35 steps match its own previous version, 4 are new; step 19 resembles step 17 of proposal 3
Proposal 2: 18 of its 28 steps match its own previous version, 2 are new; step 18 resembles step 20 of proposal 3; steps 1, 2, 4, 17, 23, 26, 28 resemble steps 1, 2, 21, 16, 21, 24, 26 of proposal 4
Proposal 3: 14 of its 29 steps match its own previous version, 0 are new; steps 10, 22 resemble steps 9, 24 of proposal 1; steps 14, 15 resemble steps 13, 19 of proposal 2; steps 1, 9, 11, 12, 16, 17, 18, 20, 25, 27, 29 resemble steps 1, 8, 10, 11, 15, 16, 17, 19, 24, 25, 26 of proposal 4
Proposal 4: 17 of its 26 steps match its own previous version, 2 are new; step 19 resembles step 24 of proposal 1; steps 6, 10, 16, 25 resemble steps 6, 11, 16, 25 of proposal 2; steps 13, 24 resemble steps 14, 29 of proposal 3
Proposal 5: 1 of its 30 steps match its own previous version, 0 are new; steps 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 27, 28, 29, 30 resemble steps 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 20, 22, 23, 24, 25, 26, 27, 29, 30, 31, 32 of proposal 1; step 1 resembles step 1 of proposal 4
PROPOSALS OF ROUND 3 (to assess):
--- PROPOSAL 1 (agent opus5_refine_1, anthropic/claude-opus-5) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90 a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95% internally and externally.
- Top-100 account outreach completed within 30 minutes of SEV1 in 95% of qualifying cases, using the approved briefing pack.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- Change-correlated incidents are identified within 10 minutes of declaration in 90% of cases, and the change-correlated share of incidents declines quarter over quarter.
- SLA credits fall to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, one concurrency drill and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and a tested read-only mode by month 6, with milestones reported to the board quarterly.
Steps (35):
1. Charter the programme: one owner, one mandate, funded, dated against the audit
Turn the CEO email into a chartered company programme with a single accountable owner and authority over all 28 teams. Incident response stops being a per-team preference and becomes a **company operating process**.
- Name the CTO as executive sponsor and a full-time Head of Reliability & Incident Management as accountable owner, supported by a programme office of three: programme lead, incident-platform engineer, reliability analyst.
- Form a decision group (Engineering, Platform/SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables on day one: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, named service ownership, and **paid on-call**.
- Publish the timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal control test month 5, mock audit month 7, SOC 2 fieldwork month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.3M/yr, training and exercises, plus 15% of engineering capacity reserved for reliability work.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence immediately.
- Publish a one-page interim severity card and a single declaration path: one chat command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim 24x7 Duty Incident Commander roster (primary plus backup) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document and one naming convention for every major incident, starting now.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is the design input, the executive narrative and the frozen "before" picture for the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the **detection backlog**.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Correlate incidents with deployments, config changes and feature flags to quantify how many started with a change we made.
- Quantify true cost beyond credits: failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost, churn risk on named accounts.
- Freeze the baselines in a signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Audit, legal and evidence scoping in week one (depends on: 1)
Design evidence as a by-product of operating, never as a reconstruction before fieldwork. Settle the scope and the legal handling of incident records now, not in month seven.
- Confirm with the external auditor the Type II observation window, the incident population definition, and the evidence they will sample. Everything from the day-7 floor onwards must count.
- Map incident response to the Trust Services Criteria with Compliance: CC7.2–CC7.5 (monitoring, identification, response, recovery), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), A1.2 (availability).
- Define the evidence set and where it is produced automatically: incident records, paging and acknowledgement logs, role assignments, status-page history, reportability decisions, postmortems, action closure with verification, training register, drill records, access reviews.
- Agree retention, confidentiality, legal hold and access rules. Decide with counsel which postmortem content is privileged and how privileged material is segregated **without making the ordinary postmortem secret**.
- Start the obligation matrix with Legal: NYDFS 23 NYCRR 500, state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network contractual windows, cyber-insurer notice.
- Begin monthly evidence sampling from month 1 so control operation is visible long before the audit.
5. Listening tour, resistance map and the written on-call deal (depends on: 1, 3)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written deal signed by the sponsor.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
- State the hard gate publicly: no new mandatory night rotation starts before compensation, training, runbooks and staffing rules are live.
6. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- Record for every service: one accountable team, engineering manager, chat channel, escalation policy, dashboards, runbook link, dependency list, regions and data stores.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record SLO, RTO, RPO, active-active or region-bound, failover method, and the dependencies that make nominal two-region redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
7. Severity standard, declaration rights and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning responders and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and responders paged, comms and scribe if customer-visible, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, credit-generating or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Attach a financial-integrity flag or a security flag to any severity. The flag forces dual control, Legal engagement and the reportability checkpoint without inventing a fifth level.
- Anchor on payments reality alongside error rates: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including a customer report.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
8. Roles, authority, concurrency and handover discipline (depends on: 7)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice to the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1 and SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, Finance and Vendor Management join on defined triggers rather than by invitation.
- Rules: command claimed within 5 minutes and stated in channel; distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- **Concurrency doctrine:** two simultaneous SEV1/SEV2 incidents activate the secondary commander and a surge roster; a designated Multi-Incident Coordinator arbitrates shared resources such as the ledger, the database platform and the deploy freeze.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
9. Three-layer 24x7 coverage: command corps, domain rotations, overnight triage desk (depends on: 5, 6, 8)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving roughly one primary week per person every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement and reconciliation, auth, API edge, Kubernetes platform, data and reporting, partner integrations). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Publish the staffing arithmetic: 10–12 domain rotations of six to eight people plus a 30-person command corps means roughly 100 of 260 engineers carry a night obligation, at about one week in six. Twenty-eight independent rotations would be unstaffable and is therefore rejected on the numbers.
- Teams too small for a fair rotation get headcount, service reassignment, or a time-limited executive exception. **Never a two-person 24x7 rotation.**
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: a human acknowledgement within 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
10. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 5, 9)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour, and a premium for the overnight triage desk.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per responder per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty since day 7 is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons without career penalty, and on-call load is published quarterly by team.
11. Alert quality contract and page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression can never hide an incident.
12. Detection uplift on the money path, validated by incident replay (depends on: 6, 11)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action and a missed-detection review.
13. Change intelligence and deployment safety on the critical path (depends on: 6, 12)
Most of these incidents start with something we changed. Make change the first hypothesis the tooling answers, and make changes safer to reverse.
- Stream every deployment, configuration change, feature-flag flip, schema migration and infrastructure change into the incident timeline with service labels and owners.
- Give the commander an automatic "what changed in the last 60 minutes on the affected journey" panel at declaration time.
- Require Tier 0/1 changes to be progressively delivered with a documented rollback that is tested, time-bounded and executable by the on-call responder without the author.
- Treat ledger schema migrations and settlement-affecting changes as a separate class: dual approval, rehearsed rollback, no deploys inside the settlement window.
- Enforce a change freeze during SEV1 and SEV2, lifted only by the commander and logged.
- Report change-correlated incidents monthly; a rising ratio is a signal to strengthen release safety, not to blame a team.
14. One pager, one incident record, one status page — migrated without a detection gap (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy paging path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access, MFA and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page delivered outside this platform creates no on-call obligation.
15. Escalation ladder and the five-minute command rule (depends on: 8, 9, 14)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts and support entitlements: AWS and database premium support, processors, sponsor banks, card networks.
16. Live execution doctrine and payments safety rules (depends on: 8, 14, 15)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed in advance rather than improvised.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and automatic reopening if impact recurs.
17. Service readiness bar, major-incident playbooks and ledger blast-radius reduction (depends on: 6, 9, 16)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human, and the shared ledger is the largest structural risk in the estate.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as a concentration risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers, stronger regional independence — with board-visible quarterly milestones.
- Enforcement: no readiness sign-off, no night paging alerts unless the manager accepts the gap in writing with an expiry date and a compensating control. Never respond to a gap by turning detection off.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
18. Internal communications protocol (depends on: 8, 14)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, Sales and Security.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change. First internal brief within 10 minutes of SEV1 and 15 of SEV2.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
19. Customer communications, status page and account-manager outreach (depends on: 7, 18)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; monitoring notice after mitigation; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, monitoring, resolution.
- Component-level status page mapped to customer journeys rather than internal service names, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1 and 60 of SEV2, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state observed impact, affected capabilities, any workaround and the next update time; **never speculate on cause, recovery time, data integrity or blame**.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
- Record any legally required restriction or delay of public detail, its approver, and the alternative stakeholder plan.
20. Support and account management as a detection and intake tier (depends on: 7, 12, 19)
Customers detected 40% of incidents first, which means the front line already holds the signal. Turn Support and account managers into an instrumented detection channel rather than a bystander.
- Give Support explicit declaration rights, a one-page trigger card, and a macro that opens an incident candidate directly in the incident platform.
- Automate the clustering rule: five similar tickets or calls in ten minutes auto-creates a triage incident assigned to the duty commander.
- Route credible partner, processor and sponsor-bank notifications into the same declaration path within five minutes.
- Generate the affected-customer list automatically from journey telemetry and the incident record, and push it to Support, CS and the account-manager briefing.
- Train Support and account managers on approved language and prohibit independent technical explanations to customers.
- Measure and publish "signal was in Support before it was in monitoring" as a detection defect, and feed each instance into the detection backlog.
21. Regulatory, partner and legal notification playbook (depends on: 4, 7, 19)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Complete the obligation matrix started in scoping and have counsel validate the triggers, deadlines, channels and submitting authority for each obligation.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with facts considered, decision-maker, timestamp and reassessment trigger.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, card networks, outside counsel, cyber-insurer, critical vendors.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the operational facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly inside the exercise programme, including one full regulator-notification simulation per year.
22. SLA credit workflow, true cost model and incentive guardrails (depends on: 7, 19)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component, including how partial degradation counts.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Guardrail against the perverse incentive: better detection will surface incidents that previously went unbilled, so credits may rise before they fall. Publish this expectation to the executive team in advance, and make it a written rule that no finance or commercial pressure may influence severity or declaration.
23. Blameless postmortem standard and Incident Review Board (depends on: 4, 7, 8)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger; any control failure.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning engineering director is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions where security, privacy or privileged content requires it.
24. Action ownership, reserved capacity and enforcement (depends on: 14, 23)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, accountable manager, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness through tests, telemetry, exercises or production evidence before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
25. Training and certification academy (depends on: 8, 15, 18, 23)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, the 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, contractual clocks, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days of joining a domain.
- Certification is valid 12 months and renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
26. Exercise programme: tabletops, game days, night drills and vendor rehearsals (depends on: 14, 17, 25)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander. First company-wide tabletop within 30 days.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times, plus a concurrency drill with two simultaneous incidents.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down. Rehearse joint escalation with AWS, a processor and a sponsor bank.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
27. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 15, 18, 19, 21, 23, 24)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, escalation and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix, Readiness Bar.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
28. Pilot on the payments critical path (depends on: 10, 12, 13, 14, 17, 25, 27)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement and reconciliation, PostgreSQL platform, Kubernetes platform, API edge, auth and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, action tracking, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The programme lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time in 95% of cases, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
29. Alert noise burn-down campaign (depends on: 11, 14, 28)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → under 1,500 pages a month by day 90, under 500 with noise below 15% by month 6.
30. Metrics, dashboards, review cadence and anti-gaming (depends on: 14, 23, 28)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time from first impact to detect, declare, commander, acknowledge, mitigate, resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors, change-correlated incident ratio.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
31. Wave rollout to all 28 teams with readiness gates (depends on: 28, 30)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the programme office for three weeks; the director signs the gate.
- **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard by team, tier and control gap.
32. Change management, fairness and pager culture (depends on: 5, 10, 28)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
33. SOC 2 evidence by design, internal control test and mock audit (depends on: 4, 27, 30, 31)
The real process is the evidence. Never build a parallel audit process, and never reconstruct records after the fact.
- Maintain the evidence set defined in scoping, produced automatically and indexed: versioned policies and exceptions, catalog records, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Sample incidents monthly from first signal through verified corrective action, deliberately including customer-reported events, downgraded incidents and missed timelines.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling end to end and reporting gaps to the sponsor.
- Run a formal mock audit in month 7 using the populations, evidence requests and interviews the auditor will use: a commander, a random engineer, Support, Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
34. Programme risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps and immediate staffing intervention.
- **A major SEV1 mid-rollout:** the programme lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
35. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 30, 31, 33)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, exercise calendar and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and quarterly board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 2 (agent gpt5.6-sol_refine_2, openai/gpt-5.6-sol) ---
Estimated complexity: high
Success metrics: - By day 7, every suspected major incident uses one record, one coordination channel, and a named Incident Commander within 10 minutes.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes.
- By month 3, at least 95% of SEV1 and SEV2 incidents have a named commander within 5 minutes.
- By month 3, at least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and appropriate technical coverage.
- By week 18, all 180 services have an owner, tier, escalation path, and tested coverage model.
- No new mandatory night rotation begins before compensation, training, access, runbooks, and minimum staffing are active.
- Every direct 24x7 technical rotation has at least six qualified responders or an approved, expiring executive exception.
- No responder is routinely primary more often than one week in six or assigned to two simultaneous primary rotations.
- Median impact-to-detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- By day 90, incident replay identifies a current detector for at least 90% of the 31 historical incidents, including its expected detection minute.
- Median time to mitigation falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of applicable SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet the required internal and customer update cadence by month 3.
- Monthly human notification episodes fall from the current 3,400 alert events to no more than 1,500 by day 90 and 500 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without reducing Tier 0 or Tier 1 replay coverage.
- Average after-hours load remains at or below two notification episodes per responder per week; every sustained breach receives a dated remediation plan.
- All six monitoring sources route human pages through the controlled paging platform by week 12, with direct legacy routes retired by week 18.
- All required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 from month 3 onward.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due dates with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Reportability is assessed and recorded within 1 hour for 100% of SEV1 and qualifying SEV2 incidents, including not-reportable decisions.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Approved ledger RTO and RPO, controlled failover, read-only mode, and post-recovery reconciliation are exercised before month 6.
- Quarterly responder surveys reach at least 75% favorable responses for fairness, ownership boundaries, compensation, and sustainability by month 6.
- Monthly contracted availability meets or exceeds 99.95% by month 6 using the contractually authoritative measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- The month-7 mock audit finds no unowned high-risk control gap and at least 95% of sampled incidents contain complete operating evidence.
- The SOC 2 Type II incident-response controls complete external testing without an unresolved material exception.
Steps (28):
1. Charter the program and fund immediate action
Make incident management a **company operating process** within 48 hours. Give one accountable leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Include Engineering, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit in a decision group. The sponsor resolves blocked decisions within 48 hours.
- Approve non-negotiables: one severity model, one human-paging path, one incident record, paid on-call, named service ownership, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, resilience, and incident actions.
- Fund tooling, compensation, training, exercises, and program staff. Compare the cost with the existing $1.3M in annual SLA credits.
- Set the schedule: operating floor by day 7, design during weeks 1–4, pilot during weeks 5–10, rollout during weeks 11–18, control tests in months 4 and 6, and mock audit in month 7.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for the final policy or new tooling. Put a minimum viable process into operation immediately and retain evidence from the first day.
- Publish a one-page interim severity guide, declaration procedure, role card, and communications clock.
- Provide one monitored declaration route through chat, telephone, and the current paging environment.
- Create one channel, bridge, timeline, and incident identifier for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced managers and engineers.
- Require a named commander within 10 minutes. The duty engineering director assumes command if the command page is unclaimed.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Stop uncompensated mandatory after-hours expansion. Pay interim duty under a temporary stipend, retroactive to program launch.
- Hold a 15-minute daily operational review until the permanent process is active.
3. Create the factual, contractual, and control baseline (depends on: 1)
Build one defensible baseline for process design, investment decisions, and SOC 2 testing. Preserve the original data so progress cannot be created by changing definitions.
- Reconstruct all 31 incidents from impact start through detection, declaration, command, mitigation, recovery, communications, credits, and corrective actions.
- Reconstruct the two incidents with unclear command minute by minute.
- Identify the missing signal for every customer-first detection.
- Inventory all six alert sources, 3,400 monthly alert events, duplicates, noisy rules, missing owners, and missing runbooks.
- Record current rotations, unpaid duty, overnight activations, schedule size, and uncovered services.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, 85% noise, $1.3M in credits, and 11 of 64 actions closed.
- Inventory customer-specific availability definitions, notice periods, credit terms, sponsor-bank obligations, and other incident-related contracts.
- Confirm the required SOC 2 Type II observation period and evidence expectations with the auditor during week 1.
4. Publish the on-call fairness contract (depends on: 1)
Treat pager resistance as a legitimate design constraint. Adoption depends on a written agreement that separates command from technical ownership.
- Interview representatives from all 28 teams and from Support, Customer Success, Security, and Operations.
- Separate concerns about unpaid work, sleep loss, unfamiliar systems, noisy alerts, inadequate runbooks, and blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for systems they own or have formally accepted and trained to support, and assisted by a separate Incident Commander.
- State that Platform may temporarily triage unknown ownership but does not inherit another team's service.
- Provide confidential accommodations for health, disability, pregnancy, or caregiving constraints without career penalty.
- Measure baseline trust, fairness, fatigue, and alert confidence. Repeat at days 60 and 120, then quarterly.
5. Build the service catalog and customer-journey map (depends on: 3)
Make a machine-readable catalog the source of truth for routing, impact analysis, status-page components, and control evidence. Every production service must have one accountable owner.
- Record the owning team, manager, business capability, repository, channel, dashboard, runbook, escalation policy, dependencies, regions, and data stores for all 180 services.
- Map initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, reporting, and onboarding to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing systems that can create material contractual impact.
- Classify Tier 2 as deferrable internal or batch systems and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional recovery mode, contractual commitments, and critical vendors for Tier 0 and Tier 1.
- Keep separate but coordinated ownership for the ledger application and PostgreSQL platform.
- Give orphan services an owner or approved decommission date within 30 days. Treat unowned Tier 0 services as release-blocking executive risks.
6. Adopt the severity model and incident modifiers (depends on: 3, 5)
Classify incidents by credible customer, financial, security, regulatory, and contractual harm. Start at the higher plausible severity while scope or integrity remains unknown.
- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; broad loss of a core payment journey; both regions impaired; or processing must be stopped. Immediately page all response roles and the Executive Duty Officer. Open the bridge within 5 minutes, freeze unrelated changes, issue internal notice within 10 minutes, publish applicable customer status within 15 minutes, start legal assessment within 1 hour, and require a postmortem.
- **SEV2 — major:** material payment degradation; settlement deadline at risk; regional impairment with reduced resilience; a critical customer or material cohort unavailable; or an SLA breach is likely. Page command and technical roles immediately. Issue internal notice within 15 minutes, applicable customer status within 30 minutes, and require a postmortem.
- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads. Page only when immediate action can reduce harm.
- **SEV4 — operational event:** no current customer impact and no credible imminent harm. Create a ticket and handle during normal hours.
- Add FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, and VENDOR modifiers. These invoke specialist controls without distorting customer-impact severity.
- Use at least SEV2 posture for unknown impact lasting 15 minutes, credible ledger-integrity risk, or a cross-domain incident without clear ownership.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Permit anyone to declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Publish a decision tree and examples based on the 31 historical incidents.
7. Define roles, authority, and handoffs (depends on: 6)
Separate coordination, communications, recordkeeping, and technical repair. One named person holds command continuously throughout every SEV1 and SEV2.
- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not act as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, executive updates, and coordination with Legal.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- **Subject-Matter Responders:** diagnose and mitigate only systems for which they have ownership, access, training, or a formally accepted support agreement.
- **Executive Duty Officer:** removes organizational obstacles and makes exceptional business decisions without displacing the commander.
- Add Security, Legal, Compliance, Finance, and Vendor Management on modifier-specific triggers.
- Require distinct commander, Communications Lead, scribe, and technical lead for SEV1. Communications and scribe may combine for the first 10 minutes of a bounded SEV2.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, reconciliation, and change evidence for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, including the exact transfer time and unresolved risks.
8. Create sustainable 24x7 coverage (depends on: 4, 5, 7)
Use central command coverage and risk-based technical coverage instead of creating 28 fragile night rotations. Responders do not carry pagers for unfamiliar code.
- Create a 24x7 Incident Command corps of approximately 30 certified people with primary and backup schedules. Two schedules create about 104 weekly assignments a year, or roughly three to four weeks per person annually.
- Create a 24x7 Communications pool of 18–24 trained people from Support, Customer Operations, Engineering Operations, and management.
- Create a similarly sized scribe pool. The backup commander temporarily records the first minutes if a scribe has not joined.
- Maintain a 24x7 Executive Duty Officer schedule and specialist contact paths for Security and Legal.
- Group Tier 0 and Tier 1 systems into roughly 8–12 coherent response domains only where responders share training, access, runbooks, and explicit support acceptance.
- Staff every critical domain with primary and secondary responders and at least six qualified people. Target eight where overnight activation is frequent.
- Give Tier 2 and Tier 3 services business-hours ownership plus tested manager and director escalation.
- Reclassify any lower-tier service capable of causing severe overnight harm rather than hiding the risk behind manager callback.
- Route unknown-owner incidents to Duty Command and Platform temporarily. Record each occurrence as a catalog control defect.
- Prohibit simultaneous primary assignments and two-person 24x7 rotations.
9. Implement compensation and fatigue protections (depends on: 4, 8)
End unpaid on-call before expanding mandatory coverage. Use fixed duty compensation so responders are not rewarded for alert volume.
- Use planning bands of $900–$1,200 per Tier 0/1 primary week and $300–$500 per secondary week.
- Use planning bands of $1,000–$1,300 per Duty Commander week and $400–$700 for Communications or scribe primary duty.
- Pay holiday premiums. Compensate all legally compensable active and waiting time for non-exempt staff, including overtime where required.
- Have HR, Finance, Payroll, and employment counsel approve final bands, tax handling, FLSA classification, New York wage-hour treatment, and schedule constraints within 14 days.
- Provide a protected recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal delivery commitments by about 15% during a primary week.
- Prohibit consecutive primary weeks, on-call during leave, hidden schedule swaps, and primary duty more often than one week in six.
- Allow responders to declare temporary fatigue-related unfitness without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Make active payroll setup, training, access, and readiness hard gates before any new mandatory night rotation starts.
10. Codify the live incident lifecycle (depends on: 6, 7)
Give every incident the same operational flow from first signal through verified recovery. The first objective is limiting customer and financial harm, not proving root cause.
- Use the states Detected, Declared, Triaged, Mitigating, Mitigated, Monitoring, Resolved, and Reviewed.
- Record impact start separately from detection. Use the earliest defensible evidence and revise it transparently when later facts emerge.
- Open with a standard command message: severity, known impact, assigned roles, immediate objective, workstreams, and next update time.
- Freeze unrelated production changes during SEV1 and normally during SEV2. Record every exception.
- Prefer reversible mitigation: rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, partner rerouting, or controlled processing suspension.
- Separate mitigation and diagnosis workstreams when staffing allows.
- Keep decisions in the shared incident record rather than direct messages.
- Require formal command and technical handoffs for shift changes, fatigue, or incidents exceeding four hours.
- For payment incidents, verify backlog handling, duplicate protection, customer state, settlement exposure, and ledger reconciliation before resolution.
- Require a severity-specific stability period and explicit handback to the owning team, Support, and Customer Success.
11. Establish one paging and incident system of record (depends on: 5, 7, 8)
Monitoring tools may remain specialized, but every human page and major-incident record must enter one controlled platform. This provides consistent routing and an audit trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page within two weeks.
- Ingest events from all six monitoring tools before disabling their direct human-paging routes.
- Route pages using the service catalog and deduplicate events belonging to the same symptom.
- Provide one declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, decisions, severity changes, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status publishing, and conference facilities.
- Apply MFA, role-based access, periodic access reviews, and tamper-evident history.
- Keep privileged legal or security material in restricted linked records rather than exposing it in the general timeline.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, the status provider, or an AWS region.
- Retire each legacy paging route only after ownership review, end-to-end testing, and two weeks of verified operation.
12. Enforce the alert-quality contract (depends on: 3, 5)
Treat every human page as a production interface with an owner and a required action. Measure human notification episodes rather than raw monitoring events.
- Require every paging rule to identify the service, owner, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Define an actionable page as one that causes or materially informs a timely intervention or risk decision.
- Define noise as duplicate, non-urgent, unactionable, stale, test-generated, or incorrectly routed notification.
- Page on payment outcomes, error-budget burn, queue age against deadlines, and financial-integrity risk rather than raw CPU, memory, pod, or log thresholds.
- Run new rules in shadow mode for seven days unless a documented emergency exception applies.
- Review repeated no-action pages within two business days.
- Set a page budget of no more than two after-hours notification episodes per responder per week, measured over four weeks.
- Make a sustained budget breach trigger a tuning sprint and block additional non-emergency paging rules.
- Require compensating detection and central approval before suppressing a Tier 0 or Tier 1 rule.
- Never disable an existing critical detector solely because its metadata or runbook is incomplete. Track the gap with a dated remediation owner.
13. Detect payment failures before customers (depends on: 5, 12)
Move detection from infrastructure health to customer journeys and ledger truth. Validate coverage against actual historical failures.
- Define SLIs and internal SLOs for initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, APIs, webhooks, and reporting freshness.
- Set internal objectives with enough headroom to protect the contractual 99.95% availability commitment.
- Run external synthetic transactions through critical journeys at least once per minute from paths independent of the production platform.
- Test each region and expose dependencies that defeat nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed payment value relative to settlement deadlines.
- Add tenant and cohort anomaly detection for high-value customers and critical payment methods.
- Convert credible Support, account-manager, processor, bank, and network reports into incident candidates within five minutes.
- Replay all 31 historical incidents. Record which current detector would fire and at what minute.
- Treat customer-first detection as a mandatory missed-detection review with a tracked action.
14. Implement the detection and escalation ladder (depends on: 7, 8, 11)
Create one time-bound path from the first credible signal to named command and the correct technical owner. Device delivery does not count as acknowledgement.
- Converge automated alerts, engineer observations, support cases, account-manager reports, partner notices, and customer calls on the same declaration path.
- Page the Duty Incident Commander and owning critical-domain primary immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at 5 minutes, manager at 10 minutes, and director at 15 minutes.
- Escalate unclaimed command to the backup commander at 5 minutes. The Executive Duty Officer assumes temporary command at 10 minutes until a certified transfer occurs.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Keep command with the current commander when service ownership remains unclear. Assign a temporary technical lead and record the ownership gap.
- Maintain tested external escalation routes for AWS, database support, processors, sponsor banks, networks, and critical vendors.
- Test the full declaration, acknowledgement, fallback, conference, and status-publishing path weekly.
- Treat missed acknowledgement, failed routing, and unowned incidents as control failures requiring review.
15. Standardize internal communications (depends on: 7, 10, 11)
Give responders one working room and stakeholders one controlled source of truth. Executives must not interrupt the technical command path.
- Maintain one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue the initial internal notice within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within the same initial-notice window.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record every material decision and outbound message in the incident timeline.
- Require the executive team to sign the communication behavior rules.
16. Standardize customer, account, and regulatory communications (depends on: 3, 6, 15)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate observed impact before root cause is known.
- Publish customer status within 15 minutes of a customer-visible SEV1 and within 30 minutes of a customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map public status components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, available workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Have account managers contact affected strategic accounts within 30 minutes for SEV1 and 60 minutes for SEV2 using the approved briefing.
- Offer status subscriptions to all customers. Auto-enroll only where contracts, consent, and applicable communication rules permit.
- Encode customer-specific notice deadlines and channels in the customer record.
- Have Legal and Compliance maintain a counsel-validated matrix covering applicable NYDFS, breach, GLBA or FTC, PCI, money-transmitter, sponsor-bank, network, insurance, and customer obligations.
- Complete and record a reportability assessment within one hour for every SEV1 and every security, privacy, or integrity-related SEV2, including decisions of not reportable.
- Let Legal own regulatory text and submission while the Incident Commander owns operational facts. Record any legally required restriction of public detail and its alternative stakeholder plan.
17. Tie incidents to SLA credits and financial exposure (depends on: 3, 16)
Link the incident record to contractual and financial outcomes. Finance should not discover outages through later credit claims.
- Define the authoritative availability calculation for each contract and customer journey with Legal and Finance.
- Calculate affected customers and minutes from incident scope and journey telemetry.
- Produce a preliminary credit and contractual-exposure estimate within five business days of resolution.
- Record failed payment count, failed or delayed value, settlement exposure, reconciliation breaks, support effort, and engineering effort.
- Establish a documented approval path for proactive credits and claims-based credits.
- Attribute credits and financial harm to recurring failure families.
- Use the quarterly credit analysis to prioritize detection, resilience, and architectural investment.
18. Establish mandatory blameless postmortems (depends on: 6, 7)
Use one learning standard with fixed deadlines. Keep learning separate from disciplinary and misconduct processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact lasting more than two hours, SLA credits, contractual breach, repeated contributing factors, major control failures, and ledger-integrity near misses.
- Produce a factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the owning engineering director accountable for completion. The commander owns response analysis, and the scribe supplies the timeline.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit answers to why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the commander or primary technical responder.
- Describe decisions using the context and information available at the time. Do not name an individual as the root cause.
- Keep HR, misconduct, and personnel matters in separate processes.
- Publish broadly useful findings internally while restricting security, privacy, personnel, and privileged content appropriately.
19. Make corrective actions enforceable risk commitments (depends on: 18)
An action is not complete when its ticket is closed. It is complete when the intended risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within 7 days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Prefer actions that remove hazards or reduce blast radius over vague actions such as retraining or adding monitoring.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of engineering capacity for approved reliability work.
- Escalate overdue high-risk actions to the manager after 7 days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit unrepaired severe conditions to block related releases.
- Verify effectiveness using tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days as complete, re-planned, superseded with evidence, or formally risk-accepted.
20. Set the service-readiness bar and payments playbooks (depends on: 5, 10, 12, 13)
A critical service must be supportable at 3 a.m. before it enters direct overnight coverage. Existing critical detection remains active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and integrity constraints.
- Require responders to demonstrate access and safe execution before independent primary duty.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional impairment, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document split-brain prevention, replay protection, failover, controlled backlog recovery, and post-recovery reconciliation.
- Require executive-approved RTO and RPO for the shared ledger cluster.
- Exercise critical runbooks at least twice a year and after material changes.
- Block new Tier 0 or Tier 1 releases and paging rules when readiness requirements are missing.
- Handle existing gaps using named owners, compensating controls, executive-approved expiry dates, and remediation plans.
- Run a parallel architecture workstream to reduce ledger blast radius, isolate non-critical readers, and strengthen regional independence.
21. Train and certify every response role (depends on: 7, 14, 15, 18, 20)
Command and communication are learned skills. Use paid working time and certify people before independent duty.
- Give all employees a 30-minute module on recognizing impact, declaring incidents, and locating the status page.
- Give responders a half-day module on severity, acknowledgement, escalation, evidence, runbooks, and financial-integrity precautions.
- Train scribes for two hours on timeline quality and fact-versus-hypothesis labeling.
- Train Incident Commanders for two days on delegation, uncertainty, severity, mitigation strategy, fatigue, handoffs, and executive management.
- Require commander candidates to complete a simulated SEV1 and two shadowed incidents or exercises.
- Train Communications Leads for one day on status writing, customer segmentation, legal boundaries, and contractual clocks.
- Require domain responders to demonstrate dashboards, access, rollback, failover, escalation, and relevant playbooks.
- Require two shadow shifts before independent primary duty.
- Renew certification annually through simulation.
- Maintain the training, assessment, and certification register as operational and audit evidence.
- Nominate an incident-management champion in each of the 28 teams.
22. Exercise command, recovery, and tool failure (depends on: 11, 20, 21)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first cross-company command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous operational and security events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors when relevant.
- Conduct at least one unannounced after-hours paging test before the audit and two annually thereafter.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure acknowledgement, command, customer notice, mitigation decision, handoff, and recovery times.
- Create tracked corrective actions for every material exercise finding.
23. Publish the signed policy and control set (depends on: 6, 7, 8, 9, 10, 12, 14, 15, 16, 18, 19)
Convert the design into concise documents that people can use during an incident. The actual operating process must also be the documented and audited process.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, evidence, and exception standards.
- Include one-page cards for severity, roles, authority, escalation, and communication timings inside the incident tool.
- State explicitly that responders support only owned or formally accepted and trained service portfolios.
- Include the compensation structure, fatigue rules, declaration rights, and non-retaliation commitment.
- Obtain approval from the CTO, HR, Legal, Security, Compliance, and Internal Audit.
- Announce the policy at an all-hands and through team briefings.
- Create an exception register with owner, rationale, compensating control, approver, review date, and expiry.
- Version every policy change. Do not rewrite historical records when the process changes.
24. Instrument the scorecard and review forums (depends on: 3, 11, 18, 19)
Measure the process before the pilot so failures become visible immediately. Report medians and 90th percentiles rather than averages alone.
- Measure impact-to-detection, detection-to-declaration, declaration-to-command, acknowledgement, mitigation, recovery, and resolution.
- Split results by severity, service tier, customer journey, region, detection source, and business-hours status.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track notification episodes, actionability, duplicates, after-hours load, missed detection, routing errors, and page-budget breaches.
- Track postmortem timeliness, action age, due-date performance, verified effectiveness, and repeated contributing factors.
- Track journey availability, error-budget burn, failed or delayed value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight work, recovery days, exceptions, sentiment, and responder attrition.
- Hold a weekly Incident Review Board chaired by the Head of Reliability with relevant directors.
- Hold a monthly executive reliability review and a quarterly control and resilience review with Internal Audit.
- Reconcile incident records monthly against support cases, customer complaints, status history, credits, and major operational anomalies to detect under-reporting.
- Use team-level scorecards to direct help and investment. Never penalize an individual for good-faith declaration.
25. Pilot the complete process on the payment path (depends on: 9, 11, 13, 20, 21, 23, 24)
Run a four-to-six-week pilot across the highest-risk journey before expanding. The interim operating floor remains active for the rest of the company.
- Include payment orchestration, ledger application, PostgreSQL platform, API edge, authentication, settlement, reconciliation, Kubernetes platform, and Support intake.
- Include teams with existing on-call experience and teams new to the model.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking together.
- Run legacy and new paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently taking command.
- Correct critical process or tooling defects within 48 hours.
- Exit only after 95% timely command assignment, 95% communications compliance, no unpaid pages, complete required postmortems, tested fallbacks, and at least 50% lower pilot noise.
- Publish the pilot results, defects, and policy changes company-wide.
26. Roll out by risk with readiness gates (depends on: 25)
Expand in controlled waves and finish early enough to accumulate operating evidence before the audit. A calendar date does not override a failed readiness gate.
- Roll out remaining Tier 0 domains first, followed by Tier 1, Tier 2, and Tier 3.
- Use four waves of six to eight teams, each lasting two to three weeks.
- Gate each service on catalog ownership, appropriate coverage, active compensation, trained responders, access, tested escalation, alert quality, runbooks, and a passed tabletop.
- Require at least six responders only for direct 24x7 technical rotations. Use business-hours coverage for lower tiers.
- Give each wave a named coach and director sign-off.
- Reschedule failed gates or use a time-limited executive exception with compensating controls. Do not create silent waivers.
- Disable legacy human-paging routes after verified cutover for each wave.
- Run a quota-based noise reduction sprint in every wave, starting with the highest-volume rules.
- Pair every suppression with a compensating-detection check.
- Publish an internal adoption dashboard by team, service tier, coverage, and control gap.
- Complete critical coverage by approximately week 12 and all 28 teams by week 18.
27. Prove SOC 2 operating effectiveness (depends on: 22, 23, 24, 26)
Generate evidence through normal operation rather than reconstructing it before fieldwork. Test both control design and consistent execution.
- Map controls to the applicable Trust Services Criteria with Compliance and the auditor, including monitoring, incident identification, response, recovery, communications, and availability.
- Retain approved policies, exceptions, service ownership, schedules, compensation activation, access reviews, training, incidents, communications, reportability decisions, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, missed timelines, non-reportable decisions, and exercises in the testing population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the evidence populations and interviews expected from the external auditor.
- Correct deviations through tracked actions with owners and dates. Never edit history to create apparent compliance.
- Verify that evidence retention covers the full auditor-defined observation period.
- Brief commanders, engineers, Support, and Compliance on the actual process without scripting inaccurate answers.
28. Inspect, adapt, and institutionalize ownership (depends on: 26, 27)
Prevent the process from decaying after rollout or the audit. Change it using measured operating evidence rather than opinion.
- Review the policy after 90 days of live operation using severity calibration, page load, missed detection, communication compliance, action closure, fatigue, and survey results.
- Remove steps that create work without reducing risk. Add controls only where incidents, exercises, or evidence show a gap.
- Reassess Tier 0 and Tier 1 classification and domain boundaries every six months.
- Assign permanent owners for policy, catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review compensation bands, rotation burden, accommodations, and staffing annually.
- Report severe incidents, credits, overdue high-risk actions, and ledger concentration risk to the board or risk committee quarterly.
- Maintain the ledger blast-radius program as an executive risk until failover, degraded mode, reconciliation, and regional independence meet approved objectives.
- Evaluate follow-the-sun coverage using one year of actual activation and staffing data.
- Build year-two plans for automated mitigation, safer deployments, graceful degradation, and error-budget release controls.
--- PROPOSAL 3 (agent qwen3.8-max_refine_3, alibaba/qwen3.8-max) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of 'not reportable'.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (29):
1. Charter the program, fund it, and start the audit clock
Convert the CEO email into a company operating process with one accountable owner, a budget, and a dated timeline that starts this week.
- Name the CTO as executive sponsor and appoint a Head of Reliability & Incident Management as the single accountable owner with full-time authority over all 28 teams.
- Stand up a three-person program office: program lead, platform engineer, reliability analyst.
- Form an eight-person steering group spanning Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, and HR. It proposes; the sponsor decides within 48 hours. Never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one human-paging path, one incident record, one postmortem format, paid on-call, mandatory action tracking, and named service ownership.
- Approve budget anchored against the $1.3M in SLA credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs, and 15% reserved engineering capacity.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Publish a one-page charter company-wide on day 2. Incident response is a company process, not a per-team preference.
- Timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day interim command floor (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place this week so the next outage has a named commander and the evidence clock starts immediately.
- Publish a one-page interim severity card and one declaration path: a Slack command, a phone number, and the existing pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs from the 12 teams already on-call. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical action items into complete, re-plan, or formally risk-accept within 30 days, prioritising ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline of incidents, alerts, and money lost (depends on: 1)
Rebuild the facts before locking any design. This is both the design input and the frozen 'before' picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines in one signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the written fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written, signed deal.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- State explicitly that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Map compliance, evidence, and audit requirements from day one (depends on: 1)
Design evidence as a by-product of operations, not a reconstruction before the auditor arrives. The interim process in week 1 is already evidence.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for internal and external communication; CC5 for control activities; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure with verification, training register, drill records, reportability decisions including 'not reportable'.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records. Require role-based access, MFA, and periodic access review.
- Version and approve all policy documents from day one: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard, Alert Quality Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has a named owner. A wrong owner recreates the pager objection. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Assign one accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. **Tier 1**: customer-facing but degradable. **Tier 2**: internal or batch. **Tier 3**: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for every Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform. They fail differently and need different hands.
- Orphan services get an owner within 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Missing ownership or runbooks blocks Tier 0/1 releases.
7. Adopt the severity scale, declaration rights, and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix. Anyone may declare. Nobody is penalised for over-declaring.
- **SEV1 (crisis):** money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; deploy freeze; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads; business-hours comms; postmortem if customer-detected, over two hours, a repeat, or credit-generating.
- **SEV4:** no customer impact. Ticket only. Never pages.
- Payments-specific anchors: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident become at least SEV2.
- Only the Incident Commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
8. Define incident roles, authority, dual-control, and handover discipline (depends on: 7)
Solve 'nobody in charge for an hour' by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Command claimed within 5 minutes and stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Staff 24x7 with a three-layer coverage model, not 28 night rotations (depends on: 6, 8)
Do not create 28 night rotations. That is precisely what engineers are rejecting. Centralise coordination in a trained command corps and keep technical ownership local.
- **Layer A — Incident Command corps:** approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. Paired with a Communications Lead pool of approximately 18 from Support, CS, and engineering management, and a scribe pool used as the training entry point.
- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat a Lisbon or APAC follow-the-sun cell as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
10. Approve paid on-call, New York labor compliance, and fatigue safeguards (depends on: 4, 9)
Unpaid on-call in New York is both a retention problem and a wage-hour exposure. Pay must be live in payroll before any new mandatory night rotation starts. Publish the actual numbers, then ask people to sign up.
- Indicative scheme locked by HR, Finance, and employment counsel within 14 days: approximately $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as approximately 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
- Publish on-call load by team quarterly.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
11. Set the alert-quality standard and a hard page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target: 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, noise below 15%, with no loss of Tier 0/1 detection coverage.
12. Detect payment and ledger failures before customers do (depends on: 6, 11)
The goal is blunt: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
- Run the replay test: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish replay coverage as a leading metric.
13. Consolidate to one paging platform, one incident record, one status page (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
14. Codify the five-minute escalation path and live execution doctrine (depends on: 8, 9, 13)
Write one unskippable path from 'something looks wrong' to 'someone is in charge'. The default action is never waiting. If nobody claims command within 5 minutes, the platform assigns it and announces it.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute.
- Separate diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Closure requires a stability observation window and explicit handback.
15. Set the service-readiness bar and write major-incident playbooks (depends on: 6, 9, 12)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human at night. The shared ledger cluster is the single largest structural risk.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction: partitioning, read replicas, isolation of non-critical readers, with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
16. Run one communications clock for internals, customers, and regulators (depends on: 7, 8, 13)
Replace 'whoever is around' with one timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure. State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is 'not reportable'.
- Obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
17. Tie incidents to SLA credits and true financial cost (depends on: 7, 16)
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
18. Make blameless postmortems mandatory with one format and fixed deadlines (depends on: 7, 8)
Replace 'some incidents, various formats' with one mandatory format, fixed deadlines, and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did?
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and are never the commander of the incident under review.
- Publish a searchable library and a quarterly top-five recurring causes analysis, with restricted versions for security or privileged content.
19. Enforce action ownership, reserved capacity, and tracking (depends on: 13, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity. Closing a ticket without evidence does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
- Target: 90% of high-priority actions closed by due date within two quarters.
20. Train and certify every role before independent duty (depends on: 8, 14, 16, 18)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it. The certification register is an audit artefact.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days plus shadowing): command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
21. Rehearse with tabletops, game days, and unannounced drills (depends on: 13, 15, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-mitigation.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
22. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 14, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 10, 12, 13, 15, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse.
- Publish a one-page result to the whole company.
24. Instrument metrics, dashboards, review cadence, and anti-gaming (depends on: 13, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
25. Roll out in risk-ordered waves with readiness gates (depends on: 23, 24)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, never waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave. Give every team its ranked list of noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Review missed detections monthly with the same seriousness as noise.
- Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
26. Run change management, fairness, and pager culture from day one (depends on: 4, 10)
Run this in parallel from day one. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
27. Produce SOC 2 evidence by operating, test internally, and mock-audit (depends on: 22, 24, 25)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence. Documented exceptions beat a claim of perfection.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5, CC2.2/CC2.3, CC5, and A1.2. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including 'not reportable', postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in months 4 and 6, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
28. Maintain the program risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and pre-commit the response. Review monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
29. Inspect at 90 days, lock year-two ownership, and prevent decay (depends on: 24, 25, 27)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 4 (agent grok4.6_refine_4, xai/grok-4.6) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- By day 7, every suspected major incident uses one record, one channel, and a named commander; Support can declare without engineering confirmation.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of customer-visible SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (26):
1. Charter the program, fund it, and start the audit clock
Turn the CEO email into a chartered company program within 48 hours. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability as full-time owner, with a three-person office: program lead, platform engineer, and analyst.
- Form a small decision group of Engineering, SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit. It proposes. The sponsor decides within 48 hours.
- Lock the non-negotiables now: one severity scale, one human-paging platform, one incident record, one postmortem format, mandatory action tracking, paid on-call, and named service ownership.
- Confirm the SOC 2 Type II observation window with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, exercises, and reserved engineering capacity against the $1.3M in SLA credits.
- Reserve 15% of engineering capacity for detection, runbooks, and incident actions, protected by the sponsor.
- Publish a one-page charter on day 2. Clock: floor by day 7, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for policy, tooling, or the audit. Put a crude but real process in place this week so the next outage already has an owner.
- Publish a one-page interim severity card and one declaration path: chat command, phone number, and current pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline, and one naming convention for every suspected major incident.
- Triage the 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
- Replay one of the two nobody-in-charge incidents as a tabletop within 14 days using this floor process.
3. Rebuild the forensic baseline and incident-replay catalog (depends on: 1)
Rebuild the facts before locking design. This is the design input, the frozen before-picture for the CEO, and the test set for detection work.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact missing signal. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Name the top 50 noisy rules and every rule with no owner or runbook.
- Quantify true cost: credits, failed payment volume, delayed value, reconciliation breaks, and engineering hours lost.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11 of 64 actions closed, on-call in 12 of 28 teams.
- Build a **replay catalog**: for each historical incident, the detector that should now fire, at which minute, and the owner of the gap.
4. Publish the on-call fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders. Do not punish them by making them the company's permanent night watch.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the **fairness contract** in writing: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive exemption path for health, disability, or caregiving.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and journey tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has one owner. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Record owning team, manager, chat channel, escalation policy, dashboards, runbook, dependencies, regions, and data stores for all 180 services.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, refund, report, onboard — to services, data stores, regions, sponsor banks, and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform. They fail differently.
- Orphan services get an owner in 30 days or an approved decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Coverage follows the journey, not the org chart. Small teams that own Tier 0 pieces get headcount, service reassignment, or membership in a domain rotation. Never a two-person 24x7 rota.
6. Lock severity levels, integrity flags, and the incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Classify on actual or credible customer, financial, security, and contractual harm, never on who reported it or how hard the fix looks.
- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 contained: narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages a human.
- Attach an **integrity, security, settlement, or regulatory flag** to any severity. The flag forces dual-control, Legal, and the reportability checkpoint without inventing a fifth level. A one-customer ledger corruption is still a flagged crisis.
- Auto-escalate to at least SEV2: any ledger-cluster event, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident.
- Anyone may declare. Nobody is punished for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- Communications Lead: single voice for the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists. A human validates at SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory text. The commander owns the facts.
- Distinct people for command, comms, and technical lead at SEV1. Comms and scribe may combine only for bounded SEV2.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
- Every role assignment and handover is announced verbally and in writing with the exact time and open risks.
8. Staff 24x7 with command, domains, and a Duty Triage Desk (depends on: 4, 5, 7)
Do not create 28 night rotations. That is what engineers are rejecting. Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified people from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to eight months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate ownership into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- **Duty Triage Desk**: a small paid overnight first-line rotation that owns the first ten minutes of ambiguous, unowned, or low-confidence pages. It verifies, enriches, and applies only the runbook's safe steps, then wakes the owning domain. It never sits on a suspected ledger, payment-halt, or security event. Those page commander and likely domains immediately, in parallel.
- Overnight comms: SEV1 pages a Communications Lead 24x7. Customer-visible SEV2 lets the commander publish the first status from a template; Comms is paged if the incident is still open at 30 minutes or a top-100 account is affected.
- Staffing math: 260 engineers can sustain 10–12 domain rotations, one command corps, and one triage desk. They cannot sustain 28 night rotas. Role exclusivity: nobody is primary on two rotations in the same week. Commanders may also be domain responders in different weeks.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
- Evaluate follow-the-sun as a 12-month option, not a year-one dependency. Seed Layer B from the 12 teams already on-call.
9. Pay on-call, meet New York labor rules, and cap fatigue (depends on: 4, 8)
Unpaid on-call in New York is a retention problem and a wage-hour exposure. Pay must be in payroll before any new mandatory night rotation starts.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, $400–$800 for Comms or triage-desk duty, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- **Hard gate: no mandatory night rotation starts before compensation is live in payroll.** Interim duty is paid retroactively. Budget roughly $0.9M–$1.2M a year, then refine with actual rotation count.
10. Enforce an alert-quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers, then replay the past (depends on: 5, 10)
Stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. A detector is not done until it would have caught the last 12 months.
- Define SLIs and SLOs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner, processor, and sponsor-bank notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a named defect class with a mandatory tracked action.
- Run the **replay test** against the S3 catalog. For each of the 31 incidents, name the detector that would now fire and at which minute. Close gaps the replay exposes before calling detection improved.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One chat command creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy. Unowned pages go to the Duty Triage Desk and Duty Command, and log a catalog defect.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Codify escalation, the five-minute command rule, and vendor incidents (depends on: 7, 8, 12)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting. Payments also fail at processors and banks, which you cannot patch.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Vendor-incident class: processor, sponsor-bank, card-network, or cloud-control-plane failure. Command is still required. Success is time-to-customer-notice, time-to-failover-decision, and queue management, not root cause at the vendor.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Write live execution doctrine, readiness bars, and ledger playbooks (depends on: 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving root cause. A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis. Split diagnosis and mitigation workstreams once staffing allows.
- Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Drain backlogs under control. Complete reconciliation before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date. Existing detection stays on while gaps are repaired.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers, with board-visible milestones and executive-signed RPO/RTO.
15. Run one communications clock for internals, customers, and account managers (depends on: 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for customer-visible SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys, not internal service names. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer record.
- Language rules: state impact and the next update time. **Never speculate** on cause, recovery time, data integrity, or blame.
16. Operationalize regulatory notice, partner clocks, and SLA credits (depends on: 6, 15)
In payments, some incidents start a legal clock at detection. Tie incidents to money so Finance does not learn about outages from invoices.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when not reportable**, with facts, decision-maker, and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families. Track failed payment volume, delayed value, and reconciliation breaks as the true incident cost.
17. Make postmortems mandatory and actions enforceable (depends on: 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth. Eleven of 64 closed is a process nobody enforces.
- Mandatory for every SEV1 and SEV2, any incident a customer detected first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and gap, timeline, response analysis, contributing conditions, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items. Facilitators are trained and are never the commander of the incident under review.
- Every action gets one named owner, priority, due date, expected risk reduction, verification method, and a ticket created automatically. If it is not on the board, it does not exist.
- Classes: containment in 7 days, corrective in 30, strategic in 90. SEV1 recurrence-prevention enters the next sprint ahead of roadmap work. Reserve 15% capacity.
- Overdue ladder: manager at +7 days, director at +14, CTO at +30. Overdue high-risk items need written residual-risk acceptance and can block related releases. Verify effectiveness before closing.
18. Train and certify every response role before independent duty (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Duty Triage Desk, half day: enrichment, safe-step limits, when to wake a domain immediately, when not to delay.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
19. Publish Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17)
Collapse the design into a document people will actually open mid-outage, and make it official. Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
20. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 18, 19)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, Duty Triage Desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
21. Instrument metrics, reviews, and anti-gaming (depends on: 12, 17, 20)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- Response: time from first impact to detect, declare, commander, acknowledge, mitigate, and resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
22. Rehearse with tabletops, game days, and night drills (depends on: 12, 14, 18, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced overnight paging drill to measure real acknowledgement times.
- One combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO, before the audit.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
23. Burn down alert noise and roll out by risk with gates (depends on: 20, 21)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Noise reduction is a quota inside each wave, not a background hope.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy paging paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Give every team its noisiest rules from the baseline. Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed. Pair every suppression with a compensating-detection check.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
- If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
24. Operate the fairness and culture program in parallel (depends on: 4, 9)
Run this from the moment the deal is published. Engineers judge the process on fairness. Executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
25. Prove SOC 2 operating effectiveness before fieldwork (depends on: 19, 21, 22, 23)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (depends on: 21, 23, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, metrics, and the ledger blast-radius workstream, independent of the audit cycle.
- Pre-committed contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 5 (agent deepseek-v4-pro_refine_5, deepseek/deepseek-v4-pro) ---
Estimated complexity: high
Success metrics: - Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (30):
1. Charter the program and start the audit clock
Turn the CEO email into a chartered company program with one accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
- Confirm the SOC 2 Type II observation window with the auditor; map the future response controls to applicable Trust Services Criteria; define evidence set and retention requirements.
4. Listening tour, resistance map and the written on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (depends on: 5, 7)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 4, 8)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (depends on: 5, 10)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (depends on: 7, 13)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 12, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
- Agree with Legal and Finance the availability measurement method per contract and per component; compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide credit posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain; attribute credits to root-cause families for investment decisions.
18. Blameless postmortem standard and Incident Review Board (depends on: 6, 7, 12)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
19. Action ownership, reserved capacity and enforcement (depends on: 12, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
20. Training and certification academy (depends on: 7, 13, 15, 18)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
21. Exercise programme: tabletops, game days and unannounced drills (depends on: 12, 14, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
22. Publish Incident Management Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
24. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
25. Alert noise burn-down campaign (depends on: 10, 12, 23)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
26. Wave rollout to all 28 teams with readiness gates (depends on: 23, 24, 25)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
27. Change management, fairness and pager culture (depends on: 4, 9, 23)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
28. SOC 2 evidence by design, internal testing and mock audit (depends on: 22, 24, 26)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
29. Risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
30. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 24, 26, 28)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
Your answer has these parts:
- "round_summary": one or two sentences on how the round went as a whole.
- "converging": true if the proposals of this round are more similar to each other than those of the previous round, false otherwise.
- "differences": a short list (four items at most) of what still separates the proposals of this round, naming the proposals and the steps concerned.
- "influences": a short list (five items at most) of the flows of ideas of the round: who took what from whom, and what nobody took, naming proposals and steps ("P1 and P3 took P2's CDC sync (steps 2 and 13)").
- "proposals": one entry per proposal of round 3, each with:
- "proposal": its number,
- "assessment": "improved", "worsened", "mixed" or "unchanged" with respect to its previous version ("no_previous_version" if that agent produced nothing in the previous round),
- "what_changed": how it improved or worsened and why: two or three short sentences, or a list,
- "improvements": a list of concrete gains (specific steps, metrics, structure),
- "regressions": a list of concrete losses (dropped steps, vaguer metrics, broken dependencies...),
- "taken": the ideas this proposal visibly adopted from the OTHER proposals of round 2 (not from its own previous version): one entry per idea with "from_proposal" (the number of the proposal it came from), "steps" (the numbers of the steps of that proposal where the idea lives, as listed above; empty if it is not tied to specific steps), "what" (the idea, one sentence) and "why" (how it was used or adapted, one sentence),
- "rejected": the ideas of the OTHER proposals of round 2 that this proposal visibly declined: an explicit contradiction, or a prominent idea it saw and left out while taking the opposite approach. Same fields; "why" gives the evidence (what the proposal does instead). Do not list mere omissions without evidence; an empty list is a valid answer.
[FINAL]
[SYSTEM]
You are an expert reviewer of multi-agent planning processes.
Several LLM agents drafted plans for a task, refined them over a number of rounds while seeing each other's proposals, and finally voted for the best one.
Be exhaustive but precise: name concrete steps, ideas and metrics, never generalities. Judge plans by their fitness for the task as stated, their realism, their completeness, the soundness of their order and dependencies, how measurable their success is and how they handle things going wrong.
You are an impartial evaluator, not a chronicler: assess the proposals and the process on their merits, never rationalise what happened or assume that the outcome was right.
After your analysis, answer in the requested structure.
Every text field you write will be read by a busy person who skims. Make it easy to skim: short sentences and short paragraphs; when you name several things, prefer a list to a paragraph, with sub-items when an item has parts, but keep a single fact as a sentence; lead with the point and then the evidence; name proposals and steps by number (P2, step 4); no preamble, no repetition of the question, no closing summary; bold at most one key phrase per item or paragraph. Text fields accept Markdown: a blank line between paragraphs, "- " for lists, **bold**.
[HUMAN]
Task given to the agents: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
THE INITIAL PROPOSALS (round 0):
--- PROPOSAL 1 (agent opus5_initial_1, anthropic/claude-opus-5) ---
Estimated complexity: high
Success metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 9 months.
- Customer-first detection falls from 40% of incidents to under 10% within 6 months and under 5% within 12.
- Median time to mitigate falls from 3 h 10 min to under 60 minutes within 12 months.
- An Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes.
- Status page updated within 15 minutes of SEV1 declaration and 30 minutes of SEV2 in 95% of cases.
- Monthly alert volume falls from 3,400 to under 500 pages, with actionability above 75%; out-of-hours pages under 2 per person per week.
- All six legacy alerting tools consolidated into one paging platform, legacy paging paths disabled, by week 16.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by their due date.
- SLA credits fall from $1.3M to under $400k in the first 12 months.
- Customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar.
- All 28 teams onboarded by week 24, with 24x7 rotations of 6+ certified responders for every Tier 0/1 team.
- 30+ certified Incident Commanders and 20+ certified Communications Leads, giving 24x7 primary and secondary command cover.
- Paid on-call policy approved by HR, Legal and Finance and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among engineers on rotation.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked action items.
Steps (29):
1. Program charter, executive mandate and funding
Convert the CEO's frustration into a named program with one accountable owner, a budget and a deadline that is earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, the 28 teams will opt out.
- Appoint a single Incident Management Program Lead (Head of Reliability/SRE) with direct exec sponsorship from CTO and CEO.
- Form a steering group: CTO, VP Eng, Head of Support/CS, CISO/Compliance, Legal, Finance (SLA credits), HR (on-call pay).
- Set the non-negotiables: one severity scale, one paging tool, one postmortem format, mandatory action tracking, paid on-call.
- Fix the timeline: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready by week 28 (four weeks of buffer before the audit).
- Approve budget lines: tooling (~$150–250k/yr), on-call compensation (~$600k–1M/yr), 2–3 dedicated program FTEs. Anchor it against $1.3M of credits plus incident cost.
2. Forensic baseline of the 31 incidents and the alert estate (depends on: 1)
Before designing anything, rebuild the facts. Re-open all 31 incidents and profile the 3,400 monthly alerts so every later design decision is evidence-based.
This also creates the "before" picture the exec and the auditor will compare against.
- Re-code each incident: trigger, service, detection source (customer vs monitor), timestamps for detect/acknowledge/declare/mitigate/resolve, who led, credits paid, root cause family.
- Quantify the 40% customer-first detections: which signal was missing in each case.
- Classify the two "nobody in charge" incidents minute by minute; use them as the burning-platform story.
- Audit the six alerting tools: volume per tool, per team, per alert rule; identify the top 50 rules that produce most of the 85% noise; find rules with no owner and no runbook.
- Baseline the numbers formally: MTTD 22 min, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed. Freeze them as the reference line.
3. Stakeholder listening tour and resistance map (depends on: 1)
Engineer pushback against "carrying a pager for other teams' code" is the main delivery risk. Treat it as a design input, not an attitude problem.
Run structured interviews across all 28 teams, plus Support, CS and Sales, in two weeks.
- Test the real objection: is it unpaid work, night sleep, unfamiliar code, poor runbooks, or fear of blame? Each has a different fix.
- Collect current informal practices — the 12 teams already on-call are the pilot candidates and the source of veterans.
- Document the promise that answers the objection: **you are only paged for services your team owns**, plus a trained commander who runs the incident and pulls in others.
- Map influencers and blockers by team; recruit 10–15 credible engineers as a design working group so the process is co-authored, not imposed.
- Survey baseline sentiment (trust in alerts, willingness to be on-call, burnout) to re-measure at 6 and 12 months.
4. Service ownership catalog and criticality tiering (depends on: 2)
You cannot page the right person across 180 services until each service has a named owning team. This is the foundation of both on-call fairness and severity mapping.
Build a machine-readable catalog (Backstage or equivalent) that is the single source of truth for routing.
- One owning team per service, a named engineering manager, a Slack channel, a paging escalation policy, a dependency list.
- Tier services by business impact: Tier 0 (money movement, ledger, auth, shared PostgreSQL cluster), Tier 1 (customer-facing but degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map each Tier 0/1 service to the customer-visible capability it supports (payment initiation, settlement, reporting, onboarding).
- Flag orphan services and cross-team shared components; force an ownership decision for each within 30 days, or schedule decommissioning.
- Publish coverage gaps to the steering group: any Tier 0 service without an owner is an executive escalation.
5. Severity scale and declaration criteria (depends on: 2, 4)
Define a five-level scale with objective, payments-specific triggers so declaration is a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade.
Each level triggers a fixed bundle of response, comms and postmortem obligations.
- **SEV1**: money movement stopped or incorrect, ledger integrity in doubt, data breach, full region loss, >10% of customers impacted. Triggers: immediate 24x7 page of IC + comms + exec, bridge within 5 min, status page within 15 min, mandatory postmortem, regulator assessment.
- **SEV2**: severe degradation, settlement at risk of missing a window, single large/strategic customer fully down, SLA breach likely. Triggers: IC paged, status page within 30 min, mandatory postmortem.
- **SEV3**: partial or workaround-available degradation, no credit exposure. Team-led, business-hours comms, postmortem optional but encouraged.
- **SEV4/5**: minor or internal-only; ticket-tracked, no paging.
- Add auto-escalation rules: any SEV3 open >2 h, or any incident touching the shared ledger cluster, becomes SEV2 automatically. Include a severity decision tree and 12 worked examples drawn from the 31 real incidents.
6. Incident roles, decision authority and handover rules (depends on: 5)
Solve the "nobody in charge for an hour" failure by making command explicit, transferable and logged.
Define five roles with written responsibilities, entry criteria and explicit authority.
- **Incident Commander**: owns the incident, not the fix. Authority to declare severity, pull any engineer, approve customer-impacting mitigations, invoke failover and authorise spend. The IC never types in the terminal.
- **Communications Lead**: owns status page, internal updates, account-manager briefings and the exec summary. Single voice to customers.
- **Scribe**: maintains the timeline, decisions and open questions; feeds the postmortem and the audit evidence trail.
- **Subject-Matter Responders**: engineers from owning teams; they investigate and remediate, and report to the IC.
- **Executive Liaison** (SEV1 only): shields the IC from exec questions and owns regulator/board escalation.
- Rules: the IC role is assumed within 5 minutes of declaration, stated explicitly in the channel ("I am IC"), and any handover is announced and logged. Roles may be combined below SEV2; never at SEV1.
7. 24x7 incident command coverage model (depends on: 6, 4)
Command is staffed by a small, trained, cross-team pool — not by 28 teams individually. This is what makes 24x7 realistic in one New York time zone.
Design a central rotation that scales with a growing certified pool.
- Create a **Duty Incident Commander** rotation of 25–35 certified volunteers (target ~1 per team, plus managers and senior engineers), giving each person roughly one week per 6–8 months.
- Pair with a Duty Comms Lead rotation (Support/CS leads plus engineering managers, ~15–20 people) and a Scribe pool (rotating, lowest barrier, used as the training entry point).
- Coverage: primary + secondary IC at all times; hard 5-minute acknowledgement SLA with automatic failover to secondary, then to the on-call engineering director.
- Night coverage options to evaluate in writing: US-only rotation with paid night stipend now; a Lisbon/Dublin or APAC follow-the-sun cell as a 12-month option; a 24x7 NOC-style triage desk for first-line detection.
- Eligibility: certification required (S21); commanders are volunteers with manager approval and can step out with 30 days' notice.
8. Team on-call structure, rotations and routing rules (depends on: 4, 6)
Rebuild team on-call around the principle that answers the pushback: **you are only paged for code your team owns**.
Apply a tiered obligation so 28 teams are not treated identically.
- Tier 0/1 owning teams (expected 14–18 teams): 24x7 primary + secondary, minimum 6 people per rotation, one-week shifts, handover Wednesday mornings.
- Tier 2/3 teams: business-hours on-call with a best-effort out-of-hours escalation path, no night paging.
- Platform/Infrastructure and Database teams: 24x7, since they own the shared PostgreSQL ledger cluster and the Kubernetes/regional layer.
- Rotations under 6 people are merged across teams or backfilled by hiring; no rotation of fewer than 4 is approved.
- Routing: every page resolves through the service catalog to the owning team's escalation policy; cross-team pages are made by the IC, never by an alert.
- Guardrails: maximum one week in four, no on-call in the week after a SEV1 you led, protected recovery time after any night page, and a per-person page budget (see S10).
9. On-call compensation, labour compliance and fairness policy (depends on: 8, 3)
Unpaid on-call is both a retention risk and a legal exposure in New York. Paying for it is the fastest way to convert resistance into participation.
Design the scheme with HR, Legal, Finance and Payroll, and publish it before asking anyone to sign up.
- Base stipend per week on rotation, differentiated by tier: e.g. $800–1,200 for 24x7 Tier 0/1, $300–500 for business-hours rotations, with premiums for holidays and weekends.
- Per-incident payment for out-of-hours activation (e.g. $150 per night page plus hourly beyond one hour) and guaranteed time-off-in-lieu after night work.
- Separate Duty IC stipend, since command is a distinct and heavier burden.
- Verify FLSA exempt/non-exempt treatment, NY State wage rules and overtime exposure for non-exempt staff; document the legal review.
- Budget and model the annual cost; get board/CFO approval as a line item, benchmarked against $1.3M of credits.
- Add non-cash elements: on-call time counted as delivery load (teams reduce sprint commitment by ~15%), incident leadership recognised in promotion criteria, and a public quarterly report of on-call load per team.
10. Alert quality standard and page budget (depends on: 2, 4)
3,400 alerts a month at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a contractual condition of paging someone.
Publish a standard, then enforce it mechanically.
- Every paging alert must have: a named owning team, a documented customer impact, a runbook link, a tested threshold, and a severity mapping. Alerts failing this are demoted to ticket or deleted.
- Page only on symptoms that affect customers (SLO burn rate, error budget, queue depth against settlement deadlines); cause-based CPU/memory alerts become dashboards or tickets.
- Set a **page budget**: maximum 2 out-of-hours pages per person per week. Breach triggers a mandatory alert-tuning sprint for the owning team and blocks new alert creation.
- Auto-quarantine: any alert that fires more than 5 times a month without action, or has >70% no-action acknowledgements, is silenced automatically and returned to its owner.
- Monthly alert review per team: kill, tune, or keep, with the numbers on screen.
- Target: 3,400 → under 500 pages a month, with actionability above 75% within six months.
11. Detection uplift: SLOs, synthetic journeys and ledger assurance (depends on: 10, 4)
The goal is to stop customers telling you first. Detection must be driven by customer-visible outcomes, not host metrics.
Instrument the money path end to end and alert on it.
- Define SLOs for each Tier 0/1 customer capability: payment initiation success rate, authorisation latency, settlement file timeliness, API availability and reporting freshness. Tie them to the 99.95% contractual SLA with a stricter internal target.
- Deploy synthetic transactions from outside the platform, in both regions, every 60 seconds, covering the full payment lifecycle including a small real-value canary flow where feasible.
- Add ledger assurance checks: continuous double-entry balance reconciliation, replication lag and failover-readiness alarms on the shared PostgreSQL cluster, and settlement-window countdown alerts.
- Build per-customer anomaly detection for the top 100 accounts (volume drop, error spike) so a single-tenant outage is detected before the account manager calls.
- Create an "inbound signal" bridge: any support ticket or account-manager report matching impact keywords auto-creates a triage incident within 5 minutes.
- Track every incident's detection source; make "customer detected first" a reviewed defect with its own follow-up action.
12. Tool consolidation and incident platform implementation (depends on: 5, 7, 8, 10)
Collapse six alerting tools into one paging and incident platform so there is a single queue, a single timeline and a single audit record.
Run a short, time-boxed selection and migrate within the pilot window.
- Select an integrated stack: paging/on-call scheduling plus an incident management layer (e.g. PagerDuty + incident.io/FireHydrant, or a single vendor) and a hosted status page.
- Implement one-command declaration in Slack (`/incident declare`) that creates the channel and bridge, pages the Duty IC, sets severity, opens the timeline and starts the clock.
- Migrate all monitoring sources to route into the one platform; decommission direct paging from the legacy six and block new integrations that bypass it.
- Automate the evidence trail: timestamps, role assignments, severity changes, comms sent, and postmortem linkage exported for SOC 2.
- Integrate with the service catalog for routing, Jira for actions, Salesforce/CS tooling for affected-customer lists, and Zoom/Slack Huddle for the bridge.
- Hard requirement: the platform must work when AWS in one region is down — verify out-of-band paging (SMS/phone) and a printed/offline fallback runbook.
13. Detection-to-escalation path and the five-minute command rule (depends on: 6, 7, 12)
Write the single path from "something looks wrong" to "someone is in charge", and make it impossible to skip.
The design target is detection to commander in under five minutes, any hour.
- Entry points: automated alert, engineer observation, support ticket, account manager, partner bank, customer-facing SEV hotline. All converge on the same declaration command.
- Anyone in the company may declare up to SEV2; nobody is punished for over-declaring. Publish that rule in writing and repeat it.
- Auto-page ladder: Duty IC (5 min) → secondary IC (5 min) → on-call Director (10 min) → CTO. Same ladder for the owning team's responder.
- Cross-team pull: the IC can page any team's on-call directly, with a 10-minute acknowledgement obligation. This is the reciprocal commitment that makes single-team ownership viable.
- Explicit takeover protocol: if no one claims IC within 5 minutes, the platform assigns it and announces it; the assignee cannot decline, only hand over.
- Define standing severity triggers for immediate regional failover, ledger read-only mode and partner-bank notification, with pre-authorised decision rights so the IC does not wait for an executive.
14. Internal communications protocol (depends on: 6, 12)
Standardise the internal channel so responders, executives and support see the same picture without interrupting the IC.
Separate the working channel from the audience channel.
- One incident channel per incident (auto-created), one bridge, and a read-only broadcast channel for executives, Support and Sales.
- Update cadence by severity: SEV1 every 30 minutes even if nothing has changed; SEV2 every 60 minutes; SEV3 at state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, ETA or next update time, current IC and Comms Lead.
- Exec briefing rule: executives ask questions only to the Executive Liaison; the IC is not interrupted. Publish this as a behavioural expectation signed by the exec team.
- Support/CS enablement: a live affected-customer list and a holding statement within 15 minutes of SEV1/SEV2 so the front line is never guessing.
- Handover protocol for incidents beyond 4 hours: formal IC handover checklist, fatigue rule, and staffing of a second shift.
15. Customer communications and status page policy (depends on: 5, 14)
Customers currently learn of outages from their own monitoring and hear from whoever happens to be around. Replace that with a timed, owned, pre-approved process.
The Comms Lead is the single author; templates remove the need to write under pressure.
- Timing commitments: status page posted within 15 minutes of SEV1 declaration and 30 minutes for SEV2; updates every 30/60 minutes; resolution notice within 30 minutes of mitigation; customer-facing summary within 5 business days for SEV1.
- Pre-approve 12–15 templates with Legal and Comms (degradation, delay in settlement, API errors, security event, third-party failure) so nothing needs legal review mid-incident.
- Subscription-based status page with per-component granularity mapped to the customer capabilities from S11, plus an email/webhook/RSS feed.
- Tiered outreach: top 100 accounts get a direct named call or email from their account manager within 30 minutes of SEV1, with a briefing pack from the Comms Lead; long tail gets the status page and a proactive email.
- Rules of language: state impact and next update time, never speculate on cause, never assign blame to a vendor before facts are confirmed.
- Run a quarterly customer-perception check with the top accounts on whether comms were timely and useful.
16. Regulatory, partner and legal notification playbook (depends on: 5, 15)
In payments, some incidents are reportable and the clock starts at detection. Build the assessment into the process so it is never an afterthought.
Work with Legal, Compliance and the CISO to produce a decision tree and contact matrix.
- Map obligations: NYDFS Part 500 (72-hour cybersecurity event notification), state breach laws, GLBA/FTC Safeguards, PCI DSS if card data is in scope, sponsor-bank and card-network contractual notice windows, and any FinCEN/OFAC implications.
- Add a mandatory regulatory-assessment checkpoint to every SEV1 and every security-related SEV2, owned by the Executive Liaison, completed within 2 hours of declaration and recorded even when the answer is "not reportable".
- Build the contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel, with 24x7 numbers and named backups.
- Pre-draft notification letters and hold them under legal privilege review.
- Check customer contracts for bespoke notification SLAs (often 1–4 hours for enterprise accounts) and encode them in the customer tiering.
- Test the playbook once per quarter as part of the simulation programme.
17. SLA credit and financial impact workflow (depends on: 5, 15)
Link incidents to money so severity, credits and prioritisation stay consistent — and so Finance stops being surprised.
Make credit calculation an automated output of the incident record, not a negotiation.
- Define the availability measurement method per contract, per component, and agree it with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and the per-capability telemetry; generate a proposed credit schedule within 5 business days of resolution.
- Decide the posture: proactive credits for the top tier (reputation upside) versus claims-based for the rest; document the approval chain.
- Track credits per incident and per root-cause family; feed a quarterly report showing which reliability investments would have prevented which credits.
- Set a target: reduce credits from $1.3M to under $400k in year one, and use that delta as the ongoing business case.
18. Postmortem standard and blameless review forum (depends on: 5, 6)
Replace "some incidents, various formats" with a mandatory, single-format, blameless process with fixed deadlines.
The discipline is in the deadlines and the forum, not in the template.
- Mandatory for: every SEV1 and SEV2, every incident where a customer detected it first, every incident over 2 hours, every repeat of a known cause, and every near-miss involving the ledger. Optional but templated for SEV3.
- Fixed timeline: draft within 3 business days, peer review within 5, published company-wide within 10. The IC owns delivery; the owning team's manager is accountable.
- One template: timeline, customer and financial impact, detection analysis (why not sooner), response analysis (why mitigation took as long as it did), contributing factors, what went well, action items with owner and due date.
- Blameless rules in writing: describe systems and decisions in the context available at the time; no individual named as a cause; HR and management commit that postmortems are never used in performance reviews.
- Weekly 60-minute Incident Review Board: reviews all postmortems from the prior week, challenges quality, ratifies severity, and approves or rejects action items. Attendance by engineering directors is mandatory.
- Publish a searchable postmortem library and a quarterly "top five recurring causes" analysis.
19. Action item ownership and tracking system (depends on: 18, 12)
11 of 64 actions closed is the single clearest symptom of a process nobody enforces. Give actions the same status as customer commitments.
Track them where engineering work already lives, with visible escalation.
- Every action gets: a named individual owner (not a team), a priority class, a due date and a Jira ticket auto-created from the postmortem.
- Priority classes with hard SLAs: P0 prevents recurrence of a SEV1, due in 30 days; P1 in 60 days; P2 in 90 days. P0s are committed into the next sprint before any roadmap work.
- Capacity rule: teams reserve a standing 15–20% of sprint capacity for reliability and incident actions. Without reserved capacity, the actions will not land.
- Escalation ladder for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Monthly reporting of closure rate by team in the engineering leadership review; include it in manager performance objectives.
- Target: 90% of P0/P1 actions closed on time within two quarters.
20. Runbooks, major-incident playbooks and the on-call readiness bar (depends on: 4, 8)
Nobody can respond well to unfamiliar systems at 3 a.m. without runbooks — and poor runbooks are a real part of the pager resistance.
Define a minimum readiness bar that a service must meet before it is allowed to page anyone.
- Readiness checklist per Tier 0/1 service: current architecture diagram, dependency map, dashboard link, alert-to-runbook mapping, rollback procedure, feature-flag kill switches, escalation contacts, and a data-loss/latency impact statement.
- Write major-incident playbooks for the top failure modes derived from S2: shared PostgreSQL ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank/third-party outage, settlement-window breach, and suspected security compromise.
- Prioritise the shared ledger: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and a clearly stated data-loss tolerance (RPO/RTO) signed off by the exec.
- Runbooks must be tested at least twice a year in a drill; untested runbooks are marked stale in the catalog.
- Enforcement: a service without readiness sign-off cannot create paging alerts, and the gap is reported to its director.
21. Training, certification and the commander academy (depends on: 6, 13, 14, 18)
Command is a skill, not a title. Build a certification path so 24x7 coverage is staffed by people who have practised.
Use a tiered curriculum with real assessment.
- **Scribe** (2 hours): timeline discipline and tooling. The entry point for everyone.
- **Responder** (half a day): severity scale, declaration, escalation, runbook use, comms hygiene. Mandatory for every engineer joining an on-call rotation.
- **Incident Commander** (two days plus shadowing): command presence, delegation, decision-making under uncertainty, severity calls, handover, exec management. Requires two shadowed incidents and one simulated SEV1 before certification.
- **Communications Lead** (one day): status-page writing, customer tiering, legal boundaries, regulator triggers.
- Certification is valid 12 months and renewed via a simulation; the register of certified people is an audit artefact.
- Add on-call onboarding per team: a new joiner shadows two shifts before holding primary, and never holds primary in their first 90 days.
22. Simulation programme: game days, drills and wheel of misfortune (depends on: 21, 12, 20)
The process must be rehearsed before it meets a real SEV1. Simulations also build the commander pool and expose runbook gaps cheaply.
Run a standing calendar rather than one-off exercises.
- Monthly 60-minute tabletop ("wheel of misfortune") per engineering group, using a real past incident from the 31.
- Quarterly full-scale game day in production or a production-like environment: regional failover, ledger replica promotion, dependency failure, with the whole role structure activated and timed.
- Twice-yearly unannounced paging drill to measure real acknowledgement times at night.
- One security-incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Every exercise produces a lightweight postmortem and action items in the same system as real incidents.
- Measure and publish drill metrics: time to IC, time to first status update, time to correct mitigation decision.
23. Pilot with wave 0 teams (depends on: 22, 9, 11)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Run a six-week pilot with tight measurement and a public verdict.
- Select 5–6 teams: core payments, ledger/database, platform/Kubernetes, API gateway, plus two of the 12 teams already on-call.
- Activate the full stack for them: new severity scale, Duty IC rotation, single paging tool, alert budget, status-page policy, mandatory postmortems, paid on-call.
- Hold a weekly pilot retro; expect and document 20–40 process defects, and fix them in the standard before rollout.
- Validate the hard questions: does the 5-minute IC rule hold at 3 a.m.? Do cross-team pulls get answered? Is the severity tree unambiguous?
- Exit criteria: MTTD under 10 minutes for pilot services, IC assigned within 5 minutes in 95% of incidents, page volume down 50%, all postmortems on time, positive on-call sentiment.
- Publish a one-page pilot result to the whole company — this is the main adoption argument for the remaining teams.
24. Metrics, dashboards and the review cadence (depends on: 12, 18, 23)
Instrument the process itself, so improvement is visible and the audit has evidence of monitoring and review.
Define a small set of metrics with owners and a fixed meeting rhythm.
- Response metrics: MTTD, time to declare, time to IC assigned, MTTA, MTTM, MTTR, incidents per month by severity, % detected by customers first.
- Quality metrics: page volume per person per week, alert actionability rate, page budget breaches, postmortem on-time rate, action closure rate and ageing.
- Business metrics: SLA credits paid, availability against 99.95% per capability, error-budget consumption, repeat-incident rate.
- People metrics: on-call load distribution across teams, out-of-hours pages per person, on-call sentiment and attrition among on-call staff.
- Cadence: weekly Incident Review Board (postmortems and actions), monthly Reliability Review (metrics per team, alert hygiene, on-call load), quarterly Executive/Board review (credits, trends, investment asks), annual policy review.
- Every metric gets a target and a named owner; dashboards are self-serve and public inside the company.
25. Wave rollout across all 28 teams with readiness gates (depends on: 23, 24)
Roll out in four waves of six to eight teams, every three weeks, ordered by criticality. Each wave passes an explicit gate rather than a deadline.
Gates keep quality high and make the standard credible.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with 6+ certified responders, one IC candidate nominated, one drill passed.
- Assign each wave a named coach from the program team for three weeks of hands-on support.
- Gate criteria are checked and signed by the director; teams that fail are re-scheduled, not waived.
- Freeze legacy tooling per wave: after onboarding, the old alerting paths are disabled, not left as a fallback.
- Publish a live adoption scoreboard by team so progress is social, not administrative.
26. SOC 2 control mapping, evidence automation and internal dry-run audit (depends on: 18, 19, 25)
Design the process so audit evidence is a by-product of doing the work, then test it before the auditors do.
Engage the auditor early to confirm the interpretation of controls.
- Map the process to the Trust Services Criteria: CC7.3 and CC7.4 (incident identification, response, recovery), CC7.2 (monitoring), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), plus availability criteria A1.2.
- Produce and approve formal policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard — versioned, signed, annually reviewed.
- Automate evidence: incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, postmortem library, action-item closure reports, training and certification register, drill records.
- Confirm the observation window with the auditor and ensure the process is operating for a minimum of three months before fieldwork.
- Run an internal dry-run audit at month six: sample 15 incidents and walk the full evidence chain; fix gaps with 8 weeks to spare.
- Keep a remediation log for any incident where the process was not followed, with the corrective action — auditors respond better to documented exceptions than to a claim of perfection.
27. Change management, incentives and communications campaign (depends on: 3, 9, 23)
Run this in parallel from day one. The process will be judged by engineers on fairness, and by executives on visible results.
Communicate the deal explicitly and repeatedly.
- The deal in one sentence: **you are paid for on-call, you are paged only for what you own, a trained commander runs the incident, and your postmortem actions get real sprint capacity**.
- Launch communications: CTO all-hands, per-team roadshows, a one-page process card for laptops, an internal wiki hub and a Slack support channel with a 4-hour answer SLA.
- Recognition: incident-response contribution in promotion criteria and performance frameworks, quarterly awards for best postmortem and biggest alert-noise reduction, public thanks after every SEV1.
- Manager accountability: adoption, alert hygiene, action closure and on-call load in each engineering manager's quarterly objectives.
- Handle the exceptions: a written path for engineers who cannot do nights (caring responsibilities, health), covered by stipended volunteers elsewhere.
- Track sentiment quarterly and publish the results, including bad news, to keep credibility.
28. Program risk register and contingency planning (depends on: 1)
Name the ways this program fails and pre-commit the response. Review it monthly in the steering group.
The main risks are predictable.
- **Volunteer shortfall for the IC pool**: contingency is to make command a rostered duty for engineering managers and senior engineers until the pool reaches 25.
- **Compensation not approved in time**: fall back to time-off-in-lieu plus a phased stipend, but do not launch mandatory night on-call without some compensation.
- **Tool migration slipping**: keep the single-queue requirement and cut scope on the incident-management layer, not on paging consolidation.
- **Alert pruning causing a missed incident**: prune from paging to ticket first, observe for 30 days, then delete; keep a recovery path.
- **Burnout or attrition among the 12 experienced on-call teams**: monitor load weekly and cap individual page counts.
- **A major SEV1 mid-rollout**: pre-agree that the program lead becomes a full-time responder and the wave schedule slips by one wave, with the steering group informed the same day.
29. Continuous improvement, maturity roadmap and post-audit sustainability (depends on: 24, 25, 26)
Protect against the classic failure: the process decays once the audit is passed. Build the second-year plan before the first year ends.
Set a maturity model and a forward roadmap with owners.
- Quarterly process retrospective with the IC pool: what in the process itself slowed us down, what needs simplifying, is the severity scale calibrated?
- Re-baseline targets every six months; a process that hits all targets is under-ambitious.
- Year-two roadmap candidates: follow-the-sun coverage cell, automated mitigation and self-healing for the top three recurring causes, error-budget policy that gates releases, per-customer real-time impact reporting, and blast-radius reduction for the shared ledger cluster (the largest single structural risk).
- Move from lagging metrics (MTTR) to leading ones (error-budget burn, near-miss rate, drill performance).
- Make the annual policy review, certification renewal and drill calendar permanent calendar items owned by the Head of Reliability, independent of the audit cycle.
- Report to the board quarterly on availability, credits and incident trends so the process keeps executive attention after SOC 2 is signed.
--- PROPOSAL 2 (agent gpt5.6-sol_initial_2, openai/gpt-5.6-sol) ---
Estimated complexity: high
Success metrics: - Within 7 days, every suspected SEV0–SEV2 has one incident record, one channel, and a named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named incident commander for more than 10 minutes.
- By day 30, 100% of Tier 0 and Tier 1 services have a named owner, primary escalation, secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and subject-matter coverage.
- By day 120, 100% of production services have sustainable ownership and tested escalation paths.
- At least 95% of SEV0 and SEV1 pages are acknowledged within 5 minutes by month 3.
- At least 95% of SEV2 pages are acknowledged within 10 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to 10 minutes by day 90 and 5 minutes by month 6.
- The proportion of incidents first detected by customers falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to below 90 minutes by day 120 and below 60 minutes by month 6.
- At least 95% of qualifying incidents meet their initial customer-communication deadline by month 3.
- At least 95% of published incidents meet their required update cadence by month 3.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without loss of Tier 0 or Tier 1 detection coverage.
- Monthly pages fall from 3,400 to no more than 1,500 by day 90, with alert actionability and missed-detection reviews used as countermeasures against unsafe suppression.
- 100% of new paging alerts satisfy the owner, runbook, dashboard, action, severity, and escalation quality rules by day 60.
- 100% of required SEV0 and SEV1 postmortems are drafted within 3 business days and reviewed within 5 business days by month 2.
- At least 90% of postmortem actions are completed by their approved due dates by month 6.
- All 53 currently open historical actions are triaged within 30 days; all unaccepted high-risk items are completed within 90 days.
- Repeat incidents with the same unaddressed contributing factor decline by at least 50% within 6 months.
- Every primary rotation has at least six trained responders or a documented, time-limited executive exception by day 120.
- No responder is routinely scheduled more frequently than one primary week in six by day 120.
- Two end-to-end cross-company exercises, including regional and ledger scenarios, are completed before the audit, with all critical findings assigned and tracked.
- Monthly availability meets or exceeds the 99.95% contractual target by month 6, with exceptions reviewed at the executive reliability meeting.
- SLA credits decline by at least 50% on an annualized trailing basis by month 8.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (19):
1. Establish ownership, authority, and funding
Launch the program within 48 hours under an executive sponsor. Give one program owner authority to standardize incident management across all 28 teams.
- Name the CTO or equivalent as executive sponsor and a Head of Incident Management or Reliability as directly accountable owner.
- Form a working group with Engineering, SRE or Platform, Product, Support, Customer Success, Communications, Security, Legal, Compliance, Risk, HR, Finance, and Internal Audit.
- Approve authority for an incident commander to stop deployments, roll back releases, disable features, shift traffic, invoke continuity plans, and pause payment processing when integrity is at risk.
- Preserve financial controls. The incident commander may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access requirements.
- Fund paging tools, compensation, training, observability work, exercises, and dedicated reliability capacity.
- Reserve engineering capacity for incident remediation. Start with 10% of capacity and adjust through quarterly risk reviews.
- Record the current baselines: 31 customer-impacting incidents, 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M in credits, 3,400 monthly alerts, 85% noise, and 11 of 64 actions closed.
- Maintain a risk register for staffing gaps, shared-ledger concentration, regional failover, alert coverage, third parties, and audit readiness.
2. Install immediate minimum controls (depends on: 1)
Put an interim process in place during the first seven days. Do not wait for tool consolidation, policy perfection, or the SOC 2 audit.
- Publish a one-page interim severity guide and incident declaration procedure.
- Establish one continuously monitored incident declaration path through chat, telephone, and the paging system.
- Create a standard incident channel, conference bridge, incident document, and event naming convention.
- Staff an interim primary and backup incident commander at all times. Compensate this duty retroactively under the final compensation policy.
- Give trained duty personnel access to the status page, paging system, dashboards, support queue, service catalog, and emergency contacts.
- Require an incident commander to be named within 10 minutes for every suspected major incident.
- Direct Support to escalate credible customer reports immediately rather than waiting for engineering confirmation.
- Triage all 53 open historical postmortem actions. Complete, re-plan, or formally risk-accept the items affecting ledger integrity, payment duplication, regional resilience, security, and detection first.
- Hold a daily 15-minute operational review until permanent controls are working.
3. Create the service and dependency catalog (depends on: 1)
Build a reliable ownership map for all production services and customer journeys. This is the basis for paging, escalation, impact assessment, and audit evidence.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, data stores, queues, external processors, banking partners, and customer-facing endpoints.
- Assign each component a single accountable team, primary responder group, secondary escalation group, engineering manager, and product owner.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 according to financial integrity, customer impact, dependency centrality, and contractual obligations.
- Treat the ledger, payment orchestration, authentication, settlement, reconciliation, and critical shared infrastructure as Tier 0 or Tier 1.
- Map every important customer journey to its service, database, cloud-region, and third-party dependencies.
- Record SLOs, RTOs, RPOs, data classification, dashboards, runbooks, deployment controls, feature flags, and failover methods.
- Assign separate but coordinated responders for the ledger application and the shared PostgreSQL platform.
- Document whether each service is active-active, active-passive, or region-bound. Identify dependencies that make nominal regional redundancy ineffective.
- Make missing ownership or missing runbooks a release-blocking risk for Tier 0 and Tier 1 services.
4. Adopt severity and incident lifecycle standards (depends on: 1)
Approve one impact-based severity model for operational, security, data, and third-party incidents. When evidence is incomplete, start at the higher credible severity and downgrade later.
- SEV0, crisis: Use for actual or credible unauthorized, lost, duplicated, or corrupted movement of money; ledger integrity loss; material security compromise; material data exposure; both-region failure; or an event likely to require crisis or regulatory management. Page all roles immediately, engage executives, Security, Legal, Compliance, and Risk, and consider pausing payment activity.
- SEV1, critical: Use for widespread inability to initiate, process, settle, or reconcile payments; a core journey failing without a viable workaround; material regional impact; fast SLA-budget exhaustion; or an imminent integrity risk. Staff all incident roles, notify the executive duty officer, and publish customer communications.
- SEV2, major: Use for a material customer subset, one or more critical customers, significant degradation with a workaround, partial transaction failure, or a likely contractual impact. Assign an incident commander and subject-matter responders; add communications and scribe roles whenever customers are affected.
- SEV3, minor: Use for localized, low-impact degradation with no financial-integrity, security, regulatory, or material contractual risk. The owning team leads the response and keeps an internal record; external communication is not normally required.
- Base severity on actual or credible impact, not the seniority of the reporter, number of alerts, or presumed complexity of the fix.
- Permit any employee to declare an incident. Only the incident commander may lower severity after recording the evidence and rationale.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of customer impact. Define resolution only after stability, backlog processing, transaction recovery, and required ledger reconciliation are complete.
- Start measurement from the earliest reliable indication of impact, including telemetry, customer reports, and partner notifications.
5. Define roles and sustainable 24x7 staffing (depends on: 3, 4)
Separate command from technical remediation. This allows trained commanders to coordinate any incident without asking engineers to debug code they do not own.
- Incident commander: Owns severity, priorities, role assignment, escalation, decision cadence, mitigation strategy, handoffs, and final closure. One person has command at a time.
- Communications lead: Owns internal notices, status-page updates, account-manager briefs, approved customer language, and coordination with Legal or regulators.
- Scribe: Maintains a timestamped timeline of observations, decisions, commands, owners, and status changes. Automation may assist but does not replace human validation for SEV0 and SEV1.
- Subject-matter responders: Diagnose and mitigate only services or domains for which they have accepted ownership, training, access, and runbooks.
- Executive duty officer: Removes organizational obstacles and approves exceptional business decisions. This role does not take command unless a formal transfer occurs.
- Security, Legal, Compliance, Support, Vendor Management, and Business Continuity join according to predefined triggers.
- Create a company-wide incident-command rotation with at least eight certified primary commanders and eight qualified backups. Use weekly rotations with explicit handoffs.
- Create similarly sustainable communications and scribe pools using Engineering Operations, Support, Customer Operations, and Communications personnel.
- Group service responders into approximately 8–12 coherent product or platform domains rather than creating 28 fragile rotations. Each domain rotation should normally contain at least six trained responders.
- Do not place an engineer into another team's responder pool without training, access, runbooks, shadow shifts, and explicit acceptance by both teams.
- Maintain dedicated database-platform and ledger-application escalation coverage for the shared PostgreSQL environment.
- Require distinct people for commander, communications, and primary technical lead during SEV0 and SEV1 incidents.
- Require a verbal and written handoff when any incident role changes. Record the exact time and the new role owner.
6. Implement compensation and fatigue safeguards (depends on: 5)
End unpaid on-call before expanding coverage. Treat availability, interrupted personal time, and overnight recovery as compensable work.
- Pay a fixed stipend for each primary on-call week and a secondary stipend equal to a defined percentage of the primary amount.
- Pay a higher holiday stipend. Apply overtime and call-out rules to non-exempt employees as required by law.
- Give exempt employees a minimum call-out credit or equivalent paid recovery time for material after-hours work.
- Provide a paid recovery day after prolonged overnight work, a SEV0, or a qualifying SEV1. Managers must arrange daytime coverage rather than expecting normal output.
- Have HR, Finance, and employment counsel publish dollar amounts, tax treatment, eligibility, and payroll procedures within 14 days. Apply the policy consistently across teams and locations.
- Target rotations no more frequent than one week in six. Exceptions require a time-limited staffing plan and executive risk acceptance.
- Avoid consecutive primary and secondary weeks. A person must not be primary for two simultaneous domain rotations.
- Track after-hours pages, sleep interruptions, swaps, missed acknowledgements, and reported burnout by rotation.
- Trigger a staffing or alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
- Permit responders to declare themselves temporarily unfit after overnight work without performance penalty.
7. Consolidate incident and paging tooling (depends on: 4)
Create one operational system of record while migrating safely from the six current alerting tools. Consolidation must reduce ambiguity without creating a monitoring gap.
- Select one enterprise paging and escalation platform and one integrated incident record.
- Initially ingest events from all six tools. Deduplicate, correlate, and route them through the new platform before retiring sources.
- Integrate paging with chat, conference bridges, ticket tracking, the service catalog, observability tools, and the customer status page.
- Automatically capture declaration time, acknowledgements, role assignments, severity changes, messages, decisions, mitigated time, and resolved time.
- Use role-based access, multifactor authentication, break-glass controls, immutable audit logs, and periodic access reviews.
- Provide mobile and telephone fallback paths if chat, identity, or the primary paging tool is unavailable.
- Test paging, escalation, status publication, and conference access every week.
- Retire a legacy alert path only after its signals have named owners, successful end-to-end tests, and at least two weeks of verified operation in the new platform.
8. Improve detection and enforce alert quality (depends on: 3, 7)
Shift detection toward customer journeys, payment outcomes, and ledger integrity. Infrastructure metrics alone will not solve the current customer-first detection problem.
- Instrument payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, and reporting with SLOs and business-level success metrics.
- Run external synthetic transactions and API checks from outside the production boundary and from both AWS regions.
- Monitor transaction failure rates, processing latency, queue age, unprocessed volume, reconciliation breaks, unexpected ledger balances, duplicate identifiers, regional asymmetry, and third-party response quality.
- Correlate application telemetry with Kubernetes, AWS, PostgreSQL, network, deployment, and feature-flag events.
- Route high-priority support cases and credible partner notifications into the same incident declaration path within five minutes.
- Define noise as a page that is duplicate, informational, unactionable, non-production, or requires no timely human action.
- Require every paging alert to name an owner, affected service, urgency, customer or SLO risk, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Send non-urgent conditions to a ticket queue rather than a pager.
- Run new alerts in shadow mode for at least seven days unless an emergency risk exception is approved. Test both firing and recovery behavior.
- Review any alert with less than 50% actionability or more than three firings in seven days within two business days.
- Never silently disable a noisy alert. Verify compensating detection, record the decision, and assign a correction owner first.
- Review alert actionability, false positives, missed detection, and page load with every responder group each month.
9. Codify acknowledgement and escalation paths (depends on: 3, 4, 5, 7, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts when a qualifying signal or customer report enters the system.
- For SEV0 and SEV1, page the owning primary immediately; page the secondary after five unacknowledged minutes; page the domain manager and company incident commander at 10 minutes; and engage the executive duty officer by 15 minutes.
- For SEV2, require primary acknowledgement within 10 minutes and incident-command assignment within 15 minutes. Escalate to the secondary and manager when either target is missed.
- For SEV3, require acknowledgement within 30 minutes when immediate production action is needed. Otherwise create a prioritized work item.
- Automatically page the company incident commander for any credible integrity or security concern, cross-team event, customer-visible Tier 0 failure, regional event, or unresolved ownership question.
- If impact remains unknown after 15 minutes, raise severity rather than waiting for certainty.
- Let the incident commander summon dependency owners, cloud support, database support, payment processors, banking partners, and vendors through maintained escalation contacts.
- Test vendor contacts and premium-support entitlements quarterly.
- Route alerts with no valid owner to the central command rotation, then treat the missing ownership record as a control defect.
- Require human acknowledgement. Delivery to a device or chat channel does not count.
- Record every missed acknowledgement, failed escalation, and manual contact workaround for review.
10. Standardize live incident execution (depends on: 4, 5, 7, 9)
Give responders one concise operating procedure for the first minutes through resolution. Prioritize limiting customer and financial harm before proving a root cause.
- Open a dedicated channel, bridge, incident record, and timeline immediately for SEV0 through SEV2.
- Have the incident commander state severity, known impact, current hypothesis, immediate objective, assigned roles, and next update time.
- Freeze unrelated production changes during SEV0 and SEV1 incidents. Record exceptions approved by the incident commander.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Use pre-approved runbooks for region failover, Kubernetes recovery, PostgreSQL failover, credential rotation, queue recovery, and payment suspension.
- Guard against split-brain, replay, duplication, and out-of-order processing during regional or database recovery.
- Require reconciliation and controlled backlog processing before declaring payment or ledger incidents resolved.
- Keep diagnosis and mitigation workstreams separate when enough responders are available.
- State decisions and owners aloud and in the incident record. Avoid unrecorded direct-message command paths.
- Require stability for a severity-specific observation period before closure. Reopen the incident if impact recurs during that period.
- Conduct an explicit operational handback to the owning team, Support, and Customer Success.
11. Standardize internal, customer, and regulatory communications (depends on: 4, 5, 7, 10)
Communicate known impact early without waiting for a root cause. Use approved facts, acknowledge uncertainty, and give the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish an initial customer status within 15 minutes when disclosure is operationally and legally appropriate, then update every 15 minutes.
- For SEV1, notify internal stakeholders within 15 minutes, publish an initial customer status within 15 minutes, and update at least every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send an initial notice within 30 minutes. Update at least every 60 minutes.
- Do not normally publish SEV3 events. Notify specifically affected customers if contracts or material impact require it.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. State affected capabilities, customer symptoms, workarounds, regions, and next update time.
- Do not speculate about root cause, blame, security scope, recovery time, or data integrity.
- Give account managers a single approved briefing and an affected-customer list. Prohibit contradictory or improvised incident explanations.
- Maintain templates for outages, delays, data-integrity investigation, third-party failure, regional failure, security events, and resolution.
- Issue a resolution notice only after operational recovery and required reconciliation. Provide a customer-facing incident summary within five business days for qualifying events.
- Have Legal and Compliance maintain a jurisdiction, regulator, sponsor-bank, network, cyber-insurer, partner, and contract notification matrix.
- Where applicable, explicitly track the current New York cybersecurity-event notification clock, including the 72-hour requirement, without assuming every incident is reportable.
- Have Legal record the reportability decision, decision time, evidence, approver, deadline, and submission confirmation.
- Allow Security or Legal to limit public detail during an active threat, but require the reason and an alternative stakeholder plan to be recorded.
- Coordinate service-credit calculations and contractual notices with Finance and Customer Success from the same incident record.
12. Make postmortems mandatory and actionable (depends on: 4, 7, 10)
Use postmortems to improve systems and controls, not to assign personal blame. Keep performance or misconduct processes separate from the learning review.
- Require a postmortem for every SEV0 and SEV1.
- Require one for a SEV2 that affected customers, incurred credits, breached an SLO or contract, involved financial or data integrity, repeated a prior failure, exposed a control gap, or lasted more than two hours.
- Permit incident command, Security, Compliance, or the service owner to require a review for a near miss.
- Produce a factual draft within three business days and hold the cross-functional review within five business days.
- Use one template covering executive summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, what worked, what failed, and lessons.
- Include why monitoring did or did not detect the event before customers.
- Avoid a single-root-cause assumption. Examine technical, organizational, process, dependency, testing, and incentive factors.
- Give every action one owner, due date, priority, expected risk reduction, verification method, and linked engineering item.
- Classify actions as containment due within 7 days, corrective work due within 30 days, or strategic work normally due within 90 days.
- Require director approval and documented residual-risk acceptance for overdue high-risk actions.
- Verify effectiveness after implementation. Closing a ticket without evidence does not close the action.
- Publish broadly useful reviews internally. Maintain access-restricted versions for security, privacy, personnel, or legally privileged details.
13. Measure performance and review it routinely (depends on: 7, 8, 11, 12)
Use outcome, process, quality, and human-sustainability measures together. Do not reward teams for suppressing declarations or hiding incidents.
- Measure detection time from first impact to first internal signal, declaration time, acknowledgement time, role-staffing time, mitigation time, resolution time, and recurrence.
- Report both median and 90th percentile. Break results down by severity, service tier, customer journey, region, detection source, and owning domain.
- Track customer-first detection, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, and role conflicts.
- Track availability, error-budget consumption, failed-payment volume, delayed value, reconciliation breaks, impacted customers, contractual breaches, and service credits.
- Track alert volume, actionability, duplicates, after-hours pages, missed pages, pages per responder, and tool-source distribution.
- Track required postmortems completed on time, actions completed by due date, action age, verified effectiveness, and repeat contributing factors.
- Track rotation size, on-call frequency, swaps, recovery days, attrition signals, and quarterly responder sentiment.
- Hold a weekly operational review for recent incidents, overdue actions, alert problems, and upcoming risk.
- Hold a monthly executive reliability review covering trends, investment decisions, accepted risks, and SLA exposure.
- Hold a quarterly resilience and control review with Security, Compliance, Risk, Internal Audit, and Product leadership.
- Use team scorecards to direct investment and assistance, not individual performance penalties.
- Reconcile dashboard data against a monthly sample of incident records and customer cases to detect metric gaming or missing incidents.
14. Train and certify participants (depends on: 4, 5, 9, 10, 11, 12)
Train people before assigning full independent duty. Use paid working time for training, shadowing, exercises, and certification.
- Train all employees to recognize impact, declare an incident, and find the incident channel and status page.
- Train engineers and Support on severity, escalation, customer-report handling, evidence preservation, and financial-integrity precautions.
- Certify incident commanders through instruction, tabletop exercises, shadow incidents, and observed command performance.
- Train communications leads in status writing, contractual communications, regulator escalation, and avoiding unsupported claims.
- Train scribes in timestamping, decision capture, evidence hygiene, and separating fact from hypothesis.
- Require subject-matter responders to demonstrate dashboard, runbook, rollback, failover, and access competence for their assigned domain.
- Add training to new-hire onboarding and repeat role-specific certification annually.
- Appoint an incident-management champion in each of the 28 teams to collect feedback and support local adoption.
- Conduct listening sessions focused on pager fairness, cross-team boundaries, tooling friction, and psychological safety.
- Publish that command duty is process coordination, not responsibility for understanding or repairing another team's code.
15. Pilot and expand the on-call model (depends on: 3, 5, 6, 8, 9, 14)
Pilot the model on the highest-risk customer journeys before expanding it. Correct staffing, alert, access, and compensation defects at each stage gate.
- Start with the ledger, payment orchestration, Kubernetes platform, PostgreSQL platform, authentication, settlement, and Support intake.
- Run the command, communications, and domain rotations in parallel with existing paths for two weeks.
- Verify primary and secondary coverage, handoffs, access, runbooks, paging, conference access, status publication, and compensation processing.
- Require at least two shadow shifts before independent primary duty.
- Review every pilot page within one business day for routing accuracy, actionability, responder load, and missing context.
- Expand by customer journey and dependency domain, not by arbitrary team order.
- Provide central first-line triage if useful, but keep technical remediation with the accepted service owner.
- Do not use contractors or a managed service as the sole incident commander or sole owner of payment and ledger remediation.
- Permit temporary shared domain rotations only after service owners document training, access, runbooks, and escalation boundaries.
- Set an executive-reviewed deadline and remediation plan for any production service that cannot provide sustainable 24x7 ownership.
16. Exercise regional, ledger, and communication failures (depends on: 10, 11, 14, 15)
Validate the process under realistic conditions before relying on it. Begin in tabletop and staging environments, then use controlled production tests where risk permits.
- Run a company-wide incident-command tabletop within 30 days of policy approval.
- Exercise loss of one AWS region, Kubernetes control-plane degradation, shared PostgreSQL failure, payment-processor failure, queue backlog, credential compromise, and suspected duplicate payments.
- Exercise a simultaneous operational and security event to test command boundaries and disclosure control.
- Exercise status-page failure and loss of the primary chat or paging provider.
- Exercise overnight staffing, role handoff, executive escalation, account-manager messaging, and a potential regulator-notification decision.
- Validate backups, restore procedures, RPO, RTO, failover prerequisites, and post-recovery reconciliation.
- Do not inject uncontrolled changes into the production ledger. Use replicas, staging, simulations, or tightly governed production tests.
- Record exercise observations as tracked actions under the same ownership and due-date rules as incident actions.
- Run at least one domain exercise per quarter and two cross-company exercises before the SOC 2 audit.
17. Execute a time-boxed enterprise rollout (depends on: 2, 6, 7, 11, 12, 13, 15)
Use fixed implementation waves so the audit deadline does not become the start date. Report progress weekly and escalate missed stage gates as business risks.
- Days 0–7: Establish governance, interim command coverage, one declaration path, provisional severity, and daily operational reviews.
- By day 14: Approve the core policy, role definitions, communications timings, compensation design, and historical-action triage.
- By day 30: Complete Tier 0 ownership, certify the first command roster, begin the on-call pilot, enable standard incident records, and run the first tabletop.
- By day 60: Provide 24x7 coverage for all Tier 0 and Tier 1 customer journeys, integrate the six alert sources, and enforce postmortem tracking.
- By day 90: Migrate critical paging, implement customer-journey detection, complete status and regulatory playbooks, and materially reduce alert noise.
- By day 120: Assign sustainable ownership and escalation for every production service and complete the first controlled regional or continuity exercise.
- By day 180: Complete tool retirement decisions, verify action closure, rerun weak scenarios, and demonstrate improving detection and mitigation trends.
- In month 7: Conduct a mock audit and executive readiness review, leaving at least one month to correct evidence or operating defects.
- Use exception records with owners, expiry dates, compensating controls, and executive approval. Do not allow indefinite verbal exceptions.
18. Build SOC 2 evidence as the process operates (depends on: 1)
Design evidence collection at the start rather than reconstructing it before the audit. Demonstrate both control design and sustained operation.
- Map the incident process to applicable SOC 2 criteria with Compliance and the auditor, including detection, response, communication, change management, access, availability, and corrective action.
- Maintain approved, version-controlled policies, procedures, severity definitions, role descriptions, and exception records.
- Preserve rotation schedules, compensation activation, training attendance, certification, paging tests, access reviews, and exercise results.
- Preserve incident declarations, timestamps, role assignments, communications, decisions, status updates, postmortems, and corrective-action evidence.
- Record regulatory and contractual notification assessments, including decisions that no notification was required.
- Define retention, confidentiality, legal-hold, and access requirements for operational and security records.
- Sample evidence monthly and trace incidents from initial signal through action verification.
- Have Internal Audit or an independent control owner test the process in months 4 and 6.
- Correct control failures through tracked actions rather than editing historical records.
- Conduct the formal mock audit in month 7 using the same evidence populations expected for the external audit.
19. Sustain accountability and continuous improvement (depends on: 13, 17, 18)
Make incident management an operating discipline rather than an audit project. Keep policy, staffing, tools, and investment aligned with changing customer and system risk.
- Assign permanent owners for the incident policy, paging platform, status page, service catalog, training program, and metrics.
- Review severity thresholds, communication timings, compensation, and staffing at least annually and after material incidents.
- Use incident trends to prioritize architectural work on the shared ledger, regional independence, deployment safety, dependency isolation, and graceful degradation.
- Review repeat incidents and repeat contributing factors quarterly. Require executive action when remediation repeatedly loses priority.
- Survey responders quarterly and publish actions addressing fatigue, fairness, psychological safety, and tool friction.
- Recognize effective incident leadership, early declaration, useful postmortems, and preventive work.
- Prohibit retaliation for good-faith incident declaration or escalation.
- Provide the board or risk committee a quarterly summary of severe incidents, SLA exposure, regulatory events, overdue high-risk actions, and resilience investment.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
--- PROPOSAL 3 (agent qwen3.8-max_initial_3, alibaba/qwen3.8-max) ---
Estimated complexity: high
Success metrics: - **MTTD reduced from 22 minutes to <5 minutes** within 6 months of full rollout.
- **Customer-first detection rate reduced from 40% to <10%** within 6 months.
- **MTTM for SEV1 incidents reduced from 3 h 10 min to <60 minutes** within 9 months.
- **Monthly SLA credit payouts reduced from ~$325K to <$100K per quarter** within 12 months.
- **Alert volume reduced from 3,400/month to <600 actionable alerts/month** within 6 months; signal-to-noise ratio >80%.
- **Postmortem completion rate: 100% of SEV1/SEV2 incidents** have a blameless postmortem within 5 business days.
- **Postmortem action-item completion rate >90% within 30 days** of the postmortem (up from ~17%).
- **Zero incidents with >15 minutes of unowned command** (down from 2 incidents with >1 hour).
- **100% on-call coverage**: all 28 teams staffed with primary + secondary on-call 24×7 within 14 weeks.
- **On-call compensation adopted**: 100% of on-call engineers receiving stipends and page pay; on-call satisfaction score ≥4/5 in quarterly survey.
- **Status-page first-update within 15 minutes for SEV1 and 30 minutes for SEV2**, 100% compliance.
- **SOC 2 Type II audit passed** at month 8 with zero incident-response findings.
- **All 260 engineers trained**; 56+ certified ICs and 28+ certified CLs active within 14 weeks.
- **Six alerting tools consolidated to one** within 6 months; legacy tools decommissioned.
- **Regulator notification process tested**: at least one tabletop exercise includes a NY DFS / FinCEN notification drill, and the legal/compliance playbook is documented and approved.
- **Quarterly IMC reviews held consistently** with published KPI dashboards and action-item tracking.
- **On-call participation resistance resolved**: <10% of engineers report 'unwilling to participate' in the 6-month pulse survey (baseline to be measured in S1).
Steps (13):
1. Assess Current State and Baseline Metrics
Build an evidence-based picture of the current incident management reality before designing anything new.
Collect and catalog the last 12 months of incident data: all 31 customer-impacting incidents, 3,400 monthly alerts, on-call coverage gaps across the 28 teams, and the 11 of 64 closed postmortem action items. Interview one lead from each of the 28 teams to surface pain points, political concerns (the 'carrying a pager for other teams' pushback), and tool sprawl.
Deliverables to produce:
- **Alert inventory**: which of the six alert tools feed which teams, alert volume per team, noise rate per tool, and overlap between tools.
- **Incident timeline analysis**: median detection-to-notification-to-mitigation-to-resolution times, who detected first (internal vs. customer), who mitigated, and where handoff gaps occurred.
- **On-call coverage map**: which 12 teams have on-call, which 16 do not, rotation length, compensation status, and escalation paths (or lack thereof).
- **Postmortem audit**: format variance, action-item tracking gaps, and the two incidents with no clear owner for over an hour.
- **Tooling and integration audit**: Kubernetes observability stack, six alerting tools, status-page provider, communication channels (Slack, email, phone), and any existing runbooks.
- **Compliance gap analysis**: SOC 2 Type II CC7.3/CC7.4 requirements vs. current practice, with a risk register for the eight-month window.
- **Peer benchmarking**: incident management practices at 3–4 comparable B2B fintech platforms (e.g., Plaid, Stripe, Adyen) for severity scales, on-call comp, and MTTR targets.
2. Secure Executive Sponsorship and Form the IM Governance Body (depends on: 1)
Anchor the program with visible, top-down authority so that 28 teams adopt changes they did not individually request.
The CEO email about 'outages we hear about from clients' is a ready-made mandate. Convert it into a formal sponsorship structure.
- Appoint an **executive sponsor** (CTO or VP Engineering) who owns the program end-to-end and reports to the CEO monthly.
- Create an **Incident Management Office (IMO)**: one dedicated senior incident management lead, one tooling/platform engineer, and one part-time data analyst.
- Establish an **Incident Management Council (IMC)**: one engineering manager from each of the 28 teams, plus the VP of Customer Success, a compliance lead, and a security lead. The IMC meets bi-weekly during rollout, monthly thereafter.
- Draft and circulate an **executive mandate memo** that states: incident response is a shared operational obligation, not a per-team favor; participation in on-call rotations is a condition of employment for production-facing roles; and the program is not optional pending the SOC 2 audit.
- Allocate a dedicated budget line for on-call compensation, tooling consolidation, status-page licensing, training, and external facilitation.
3. Define Severity Levels and Automatic Triggers (depends on: 1)
Replace the current ad-hoc triage with a five-level severity taxonomy that every engineer, support agent, and account manager can apply in under 60 seconds.
- **SEV1 – Critical**: Ledger data corruption or loss, complete payment processing halt, confirmed data breach affecting customer PII or funds, or regulatory reporting breach. Triggers: automatic all-hands page to on-call, CTO + CEO paged within 10 minutes, dedicated bridge call within 5 minutes, status-page update within 15 minutes, regulator notification assessment within 1 hour, customer comms within 30 minutes.
- **SEV2 – High**: Payment processing degraded >30% throughput or >5% error rate, single-region failover failure, ledger read-only mode, or any condition likely to breach the 99.95% SLA within the current window. Triggers: primary + secondary on-call paged, incident commander assigned within 10 minutes, bridge call within 15 minutes, status-page update within 30 minutes, VP Engineering notified within 20 minutes.
- **SEV3 – Medium**: Non-critical service degradation affecting <30% of customers, single non-ledger microservice outage with working fallback, or elevated latency above SLA threshold but below halt. Triggers: primary on-call paged, IM notified within 30 minutes, status-page update within 1 hour if customer-visible, daily standup update.
- **SEV4 – Low**: Degraded internal tooling, minor UX bug with workaround, non-customer-facing alert. Triggers: next-business-day response, ticket created, no page unless on-call agrees.
- **SEV5 – Informational / Noise**: Cosmetic issues, planned-maintenance notifications, alert misfires logged for tuning. Triggers: no page, logged for weekly alert-quality review.
Define **escalation rules**: any SEV3 unresolved after 4 hours auto-escalates to SEV2; any SEV2 unresolved after 2 hours auto-escalates to SEV1. Severity can be **downgraded** only by the incident commander with IMC notification.
Publish the taxonomy as a one-page decision tree, a Slack slash-command (`/sev`), and an integration into the alerting tool so that every alert carries a suggested severity.
4. Define Incident Roles and Staffing Model (depends on: 3)
Codify four mandatory roles for every SEV1/SEV2 incident and optional roles for SEV3, then solve the 24×7 staffing problem across 28 teams.
**Roles**
- **Incident Commander (IC)**: owns the incident end-to-end, declares severity, assigns tasks, authorizes mitigations, decides when to escalate or stand down. Never writes code during the incident.
- **Communications Lead (CL)**: owns status-page updates, internal Slack channels, account-manager briefings, and regulator notifications. Separate from the IC so the IC can focus on mitigation.
- **Scribe / Timeline Keeper**: logs every decision, action, and timestamp in the incident channel and the incident-management tool. Produces the raw timeline for the postmortem.
- **Subject-Matter Responders (SMRs)**: 1–3 engineers from the owning team(s) who diagnose and fix. For the shared PostgreSQL ledger, a dedicated DBA responder is always required.
**24×7 Staffing via a Three-Tier Follow-the-Sun Model**
- **Tier 1 – Front-line on-call**: Primary + secondary responder per team, paged first. Covers the team's own services.
- **Tier 2 – Platform / SRE on-call**: A dedicated 6-person SRE rotation covering cross-cutting infrastructure: Kubernetes, the shared PostgreSQL ledger, networking, and the two AWS regions. This tier directly addresses the 'carrying a pager for other teams' concern by absorbing infrastructure incidents.
- **Tier 3 – IMC escalation**: Engineering managers and the IMO on-call for multi-team or SEV1 incidents. Provides the IC and CL when no team-level IC is available.
**Follow-the-Sun**: If any engineering hub exists in a second timezone, use it for overnight Tier 1 coverage. If not, partner with a managed on-call service for overnight first-response triage (severity declaration + paging the correct team), reducing 3 a.m. pages for NY-based engineers.
**IC and CL pools**: Nominate at least 2 ICs and 1 CL per team (56 ICs, 28 CLs minimum). ICs are trained and certified before they rotate. For SEV1 incidents, the IC must be a certified IC from the IMC pool, not just 'whoever is around'.
**Ledger-specific rule**: Because the PostgreSQL ledger is shared, a **Ledger Duty Officer** from the SRE Tier 2 is always on the bridge for any incident touching ledger services, regardless of which team owns the failing microservice.
5. Design On-Call Rotations, Compensation, and Alert-Quality Rules (depends on: 4)
Make on-call sustainable, fairly compensated, and free of alert noise so engineers stop resisting participation.
**Rotation Design**
- 7-day rotations, one primary + one secondary per team per week. No engineer is on-call more than one week in four.
- Minimum 48-hour rest between rotations. No on-call during approved PTO.
- All 28 teams participate. Teams without current on-call get a 90-day ramp with a shadow rotation before going live.
- Tier 2 SRE rotation: 6 engineers, one week on / five weeks off, with a dedicated backup.
**Compensation Package**
- **Base on-call stipend**: $500 per week of primary on-call, $250 for secondary, paid regardless of whether pages fire.
- **Page pay**: $75 per acknowledged page outside business hours; $150 if the page leads to active incident work.
- **Time-off-in-lieu (TOIL)**: Any engineer who works >4 hours overnight (00:00–06:00 local) gets a full TOIL day. >8 hours in a single incident gets 1.5 TOIL days.
- **SEV1 bonus**: $300 flat bonus for every engineer who actively works a SEV1 incident, paid within the next pay cycle.
- **Annual on-call cap**: No engineer exceeds 13 weeks of on-call per year. Exceeding the cap triggers a mandatory team-staffing review.
- Budget estimate: ~$420K/year for stipends and page pay across 28 teams; present this to the CFO as a fraction of the $1.3M annual SLA credit cost.
**Alert-Quality Rules (the '85% noise' problem)**
- Every alert must carry: owning team, suggested severity, runbook link, and a 30-day noise score.
- **Alert budget**: each team gets a maximum of 100 actionable alerts per month. Exceeding the budget triggers a mandatory alert-tuning session with the IMO.
- **Noise threshold**: any alert that fires >10 times in 7 days with no human action is auto-flagged for suppression or tuning within 14 days.
- **Alert review cadence**: weekly 30-minute alert-quality review per team; monthly cross-team alert review in the IMC.
- **Sunset rule**: alerts with no runbook are demoted to SEV5 after 30 days and suppressed after 60 days unless a runbook is written.
- Target: reduce monthly alert volume from 3,400 to <600 actionable alerts within 6 months.
6. Build Detection, Escalation, and Communication Paths (depends on: 3, 4)
Eliminate the 22-minute median detection gap and the 40% customer-first-detection rate with layered monitoring and a single escalation spine.
**Detection Layers**
- **Synthetic transactions**: run a payment end-to-end through the full stack (API → service → ledger → confirmation) every 60 seconds from both AWS regions. Alert if latency >2× baseline or any step fails. This catches what per-service metrics miss.
- **Customer-traffic anomaly detection**: monitor API error rates, payment success rates, and latency percentiles per customer cohort. Alert on >2σ deviation.
- **SLO-based alerting**: define SLIs for the 99.95% SLA (availability, latency p99, ledger consistency). Alert when error budget burn rate exceeds threshold, before the SLA actually breaches.
- **Infrastructure health**: Kubernetes node/pod health, PostgreSQL replication lag, disk I/O, and cross-region latency.
- **Support-ticket spike detection**: if >5 customers open tickets about the same symptom within 10 minutes, auto-create a SEV3 candidate.
**Escalation Path**
- Alert fires → PagerDuty routes to Tier 1 primary → 5-min no-ack → Tier 1 secondary → 10-min no-ack → Tier 2 SRE → 15-min no-ack → IMC on-call manager → 20-min no-ack → VP Engineering auto-page.
- Any SEV1 declaration auto-pages the CTO, opens a dedicated Slack channel + Zoom bridge, and notifies the CL.
- **No incident goes unowned for >15 minutes.** If no IC is assigned by minute 15, the IMC on-call manager assumes IC role by default.
**Internal Communications**
- Dedicated Slack channels: `#inc-sev1`, `#inc-sev2`, `#inc-sev3` (auto-created per incident), plus `#inc-updates` for broadcast.
- IC posts a structured update every 15 minutes (SEV1), 30 minutes (SEV2), 1 hour (SEV3) into the incident channel.
- CL posts a summary to `#inc-updates` and notifies relevant engineering managers.
**Customer Communications**
- **Status page**: auto-updated via API. SEV1: first update within 15 minutes, then every 30 minutes until resolved. SEV2: first update within 30 minutes, then every hour. SEV3: within 1 hour if customer-visible.
- **Account managers**: CL briefs AMs via a dedicated Slack channel within 30 minutes (SEV1) or 1 hour (SEV2). AMs contact their top-20 revenue accounts directly.
- **Customer email/SMS**: for SEV1 and SEV2, automated notification to all 2,100 customers via the status-page subscription system within 30 minutes.
- **Regulator notification**: Legal/Compliance assesses within 1 hour whether NY DFS, FinCEN, or card-network notification is required. If yes, file within the regulatory deadline (typically 72 hours for NY DFS cybersecurity events). Log the decision and filing in the incident record.
**Post-resolution**: CL publishes a 'resolved' update within 30 minutes of mitigation. For SEV1/SEV2, a preliminary customer-facing RCA summary is published within 5 business days.
7. Standardize Postmortems with Tracking and Accountability (depends on: 3, 6)
Fix the '11 of 64 action items closed' problem with a mandatory, uniform, blameless postmortem process backed by engineering-manager accountability.
**When Mandatory**
- All SEV1 and SEV2 incidents: postmortem required within 5 business days.
- SEV3 incidents: postmortem required if >50 customers affected, if the incident lasted >4 hours, or if it was customer-detected.
- SEV4/SEV5: optional, but any recurring SEV4 (≥3 times in 30 days) triggers a mandatory review.
**Format (single template, enforced by tooling)**
- Incident summary (severity, duration, customers affected, revenue impact, SLA credit exposure).
- Timeline (auto-generated from scribe notes + alert timestamps).
- Detection analysis: how was it detected, why did it take X minutes, could it have been faster.
- Root cause analysis using 5-Whys or fault-tree, not blame.
- Contributing factors (process, tooling, staffing, knowledge gaps).
- Impact quantification: customers affected, transactions failed, SLA credits triggered.
- Action items: each with a **named owner**, **due date**, **priority**, and **ticket in Jira**.
- Lessons learned and what went well.
**Blameless Review Meeting**
- Held within 5 business days, facilitated by the IMO or a trained facilitator (never the IC of that incident).
- All responders, the CL, relevant engineering managers, and an IMC representative attend.
- Ground rules: focus on system and process failures, not individual mistakes. The facilitator enforces this.
- Meeting recorded; notes published to the engineering-wide wiki within 48 hours.
**Action-Item Tracking and Accountability**
- Every action item is created as a Jira ticket with a due date and a named owner.
- **Engineering managers are accountable**: action-item completion is a standing agenda item in the bi-weekly IMC meeting. Any action item >7 days overdue is escalated to the VP Engineering.
- **Completion gate**: no team may close its postmortem until 100% of its action items have Jira tickets. Postmortem is 'closed' only when all tickets are resolved.
- **Quarterly audit**: the IMO audits action-item completion rates and reports to the IMC and the executive sponsor. Target: >90% completion within 30 days of the postmortem.
- Link postmortem quality and action-item completion to team health scores and engineering-manager performance reviews.
8. Define KPIs, Dashboards, and Governance Reviews (depends on: 7)
Create a measurable feedback loop so leadership can see whether the program is working and where to intervene.
**Primary KPIs (tracked weekly, reported monthly)**
- **MTTD** (Median Time to Detect): target <5 minutes (from 22).
- **MTTA** (Median Time to Acknowledge): target <5 minutes.
- **MTTM** (Median Time to Mitigate): target <60 minutes for SEV1 (from 3 h 10 min), <4 hours for SEV2.
- **Customer-first detection rate**: target <10% (from 40%).
- **SLA compliance**: maintain 99.95%; track monthly SLA credit payouts, target <$100K/quarter (from ~$325K/quarter).
- **Alert signal-to-noise ratio**: target >80% actionable (from ~15%).
- **Alert volume**: target <600/month (from 3,400).
- **Postmortem completion rate**: 100% for SEV1/SEV2 within 5 business days.
- **Action-item completion rate**: >90% within 30 days (from ~17%).
- **On-call health**: pages per engineer per week (target <5), TOIL usage, on-call satisfaction survey score.
- **Unowned incident duration**: target 0 incidents with >15 minutes without an IC.
**Dashboards**
- Real-time operational dashboard (Grafana): current incidents, active alerts, on-call roster, SLA error-budget burn.
- Weekly leadership dashboard (auto-generated): KPI trends, open action items, alert-noise report, on-call load distribution.
- Quarterly IMC scorecard per team.
**Review Cadence**
- **Weekly**: IMO publishes KPI snapshot to `#inc-updates`.
- **Bi-weekly IMC**: review open incidents, overdue action items, alert-quality exceptions, and on-call load.
- **Monthly executive review**: CTO presents KPI trends, SLA credit cost, and risk register to the CEO.
- **Quarterly incident-management review**: deep-dive into trends, training gaps, tooling needs, and process improvements. Output fed into the next quarter's roadmap.
9. Consolidate Tooling and Build the Incident Management Platform (depends on: 2, 3)
Replace six alerting tools and ad-hoc status-page updates with a single, integrated incident management stack.
**Target Tool Architecture**
- **Single alerting and on-call platform** (e.g., PagerDuty or Opsgenie): ingest all alerts, apply severity routing, manage on-call schedules, handle escalations, and send pages. Retire the other five tools within 6 months.
- **Observability consolidation**: standardize on one APM/metrics stack (e.g., Datadog or Grafana Cloud) for all 180 Kubernetes services across both AWS regions. Ensure the shared PostgreSQL ledger has dedicated dashboards.
- **Status page**: a dedicated, branded status page (e.g., Statuspage.io or Instatus) with API integration for auto-updates. Subscribe all 2,100 customers.
- **Incident coordination tool**: integrate incident-management workflows into Slack (auto-create channels, invite responders, post templates) and a dedicated incident record system (e.g., Jira Service Management, incident.io, or Rootly) for timelines, postmortems, and action-item tracking.
- **Runbook repository**: a central wiki (Confluence or Notion) with mandatory runbooks for every alert. No alert goes live without a linked runbook.
**Implementation Tasks**
- Migrate all 28 teams' alert rules into the single platform in three waves (highest-volume teams first).
- Build the severity-based routing rules and escalation policies per S3 and S6.
- Automate status-page updates triggered by severity declaration.
- Build the synthetic-transaction monitor and SLO-based alerting per S6.
- Integrate Jira for automatic action-item ticket creation from postmortems.
- Decommission legacy tools only after all teams have completed training on the new stack.
- Budget: allocate $150K–$250K/year for licensing, plus engineering time for migration.
10. Prepare for the SOC 2 Type II Audit (depends on: 7, 8, 9)
Ensure the incident management process produces the evidence the auditor will need, well before the audit window opens in eight months.
**SOC 2 Requirements to Address (CC7.3, CC7.4, CC7.5)**
- Documented incident response procedures (the severity taxonomy, role definitions, communication templates).
- Evidence of incident detection, response, and recovery for every SEV1/SEV2 incident during the audit period.
- Postmortem records with action-item tracking.
- On-call schedules, training records, and escalation evidence.
- Status-page update logs and customer notification records.
- Regulator notification logs (if any).
**Preparation Tasks**
- The IMO maintains a **SOC 2 evidence folder**: every incident record, postmortem, action-item ticket, status-page update, and training completion certificate is stored and indexed.
- Conduct a **mock SOC 2 audit** at month 5: an internal or external auditor reviews the incident management process end-to-end and identifies gaps.
- Remediate mock-audit findings before month 7.
- Ensure the incident management tool retains all records for at least 12 months (the SOC 2 Type II observation window).
- Document the **chain of custody** for incident records: who accessed, modified, or closed each record.
- Prepare a **narrative document** describing the incident management process, roles, and controls for the auditor.
- Coordinate with the compliance lead to align incident management evidence with the broader SOC 2 scope (access controls, change management, etc.).
11. Design and Deliver Training, Runbooks, and Change Management (depends on: 4, 5, 9)
Equip all 260 engineers, 28 team leads, account managers, and support staff with the knowledge and muscle memory to execute the new process.
**Training Tracks**
- **All 260 engineers** (2-hour session): severity taxonomy, how to acknowledge a page, how to join an incident bridge, how to hand off to an IC, and how to write a postmortem contribution. Delivered in team-level sessions over 4 weeks.
- **IC pool (56+ engineers)** (8-hour certification): incident command techniques, severity declaration, escalation decision-making, bridge facilitation, and blameless postmortem facilitation. Includes two tabletop exercises. Certification valid for 12 months, renewed annually.
- **CL pool (28+ staff)** (4-hour session): status-page writing, customer communication templates, regulator notification triggers, and AM briefing protocol.
- **Account managers and support staff** (1-hour session): how to read the status page, how to escalate a customer report into an incident, and what information to collect.
- **SRE Tier 2** (16-hour onboarding): Kubernetes and PostgreSQL ledger deep-dive, cross-region failover runbooks, and escalation authority.
**Runbooks**
- Every alert must have a runbook before it is routed to on-call. The IMO provides a runbook template and audits compliance weekly.
- Priority runbooks to write first: shared PostgreSQL ledger failover, Kubernetes cluster degradation, payment-processing pipeline failure, cross-region failover, and ledger data-integrity check.
- Runbooks are peer-reviewed and version-controlled.
**Change Management for Adoption**
- Address the 'carrying a pager for other teams' concern directly: publish an FAQ explaining the three-tier model, the SRE Tier 2 absorbing cross-team infrastructure, the compensation package, and the TOIL policy.
- Run **office hours** weekly for the first 8 weeks where any engineer can ask questions or raise concerns.
- Identify **team champions**: one engineer per team who volunteers as an early adopter and peer mentor.
- Publish a **weekly 'incident management newsletter'** during rollout: what changed, what improved, KPI trends, and success stories.
- Make on-call participation a documented expectation in job descriptions and performance reviews for production-facing roles.
12. Execute Phased Rollout, Tabletop Exercises, and Continuous Improvement (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11)
Introduce the process in three waves so teams are not overwhelmed, then validate with exercises and iterate continuously.
**Phase 1 – Weeks 1–6: Foundation**
- Publish the severity taxonomy, role definitions, and communication protocols (S3, S4, S6).
- Launch the single alerting platform for the 12 teams already on-call; begin migration for the other 16.
- Activate the SRE Tier 2 rotation for the shared PostgreSQL ledger and cross-cutting infrastructure.
- Deploy the status page and test the API integration.
- Begin IC and CL training (first cohort of 20 ICs, 10 CLs).
- Publish the on-call compensation package; HR integrates stipends into payroll.
- Write the top 10 priority runbooks.
**Phase 2 – Weeks 7–14: Expansion**
- All 28 teams live on the single alerting platform; legacy tools in read-only mode.
- All 28 teams on the on-call rotation schedule (the 16 new teams in shadow mode for the first 4 weeks).
- Second and third IC/CL training cohorts completed.
- First **tabletop exercise**: simulate a SEV1 ledger corruption scenario with all roles, test the escalation path, status-page updates, and AM briefings. Debrief and fix gaps.
- Postmortem template and Jira integration live; all new incidents use the standard process.
- Alert-tuning sprint: each team reduces its alert volume by 50%.
**Phase 3 – Weeks 15–24: Optimization**
- All teams fully live; legacy alerting tools decommissioned.
- Second **tabletop exercise**: simulate a SEV1 cross-region failure with regulator notification.
- First quarterly IMC review with full KPI dashboard.
- Mock SOC 2 audit (month 5) and remediation.
- Retrospective on the rollout: survey all 260 engineers for feedback, adjust compensation or rotation rules if needed.
- Establish the **continuous improvement cadence**: quarterly process review, annual severity-taxonomy review, and annual on-call compensation benchmarking.
**Ongoing Governance**
- The IMC owns the process document and approves changes.
- The IMO tracks all KPIs and reports to the CTO monthly.
- Any process change requires IMC approval and a 2-week notice period before enforcement.
- Annual external benchmarking against peer B2B payments platforms.
13. Establish Ongoing Governance, Annual Review, and Audit Readiness Cycle (depends on: 12)
Embed incident management as a permanent organizational capability, not a one-time project.
- **Annual process review**: the IMC reviews the severity taxonomy, role definitions, on-call structure, and compensation against industry benchmarks and internal KPIs. Update as needed.
- **Bi-annual tabletop exercises**: one SEV1 infrastructure scenario, one SEV1 data-breach/regulator scenario. Rotate the IC and CL assignments so everyone gets practice.
- **Quarterly alert-quality audit**: the IMO reviews alert volumes, noise rates, and runbook coverage across all 28 teams.
- **On-call health survey**: quarterly anonymous survey measuring burnout, fairness, and compensation satisfaction. Results reviewed by the IMC.
- **SOC 2 readiness cycle**: begin evidence collection immediately after each audit ends. The IMO maintains a rolling evidence folder. Mock audit at month 5 of every 12-month audit cycle.
- **Postmortem maturity tracking**: track the action-item completion rate monthly. If it drops below 80%, the VP Engineering intervenes.
- **Incident management maturity model**: adopt a 5-level maturity model (ad-hoc → defined → managed → optimized → predictive). Assess annually. Target: Level 3 within 12 months, Level 4 within 24 months.
- **Budget review**: annually review on-call compensation, tooling costs, and training budget against the reduction in SLA credits and incident frequency.
--- PROPOSAL 4 (agent grok4.6_initial_4, xai/grok-4.6) ---
Estimated complexity: high
Success metrics: - Median time to detect customer-impacting incidents ≤ 5 minutes within 6 months of go-live.
- Share of SEV-1/SEV-2 incidents first detected by customers ≤ 5% (from 40%).
- Median time to mitigate SEV-1/SEV-2 ≤ 45 minutes (from 3 h 10 min).
- Named Incident Commander assigned within 5 minutes for ≥ 95% of SEV-1/SEV-2.
- First status-page update within policy time for ≥ 95% of SEV-1/SEV-2.
- SLA credits down ≥ 80% versus the trailing $1.3M within 12 months.
- Paging volume ≤ 500 per month and noise ≤ 15% (from 3,400 and 85%).
- 100% of production services have a named owning team and a paging policy.
- Postmortems filed within 5 business days for 100% of SEV-1/SEV-2; action-item close rate ≥ 80% within 30 days.
- 24×7 IC and critical-path coverage with zero unfilled shifts per quarter.
- Paid on-call live for every rotation before that rotation pages humans.
- SOC 2 Type II incident-response controls evidenced for ≥ 5 months before the auditor's report.
- On-call pulse: ≥ 70% of engineers agree rotations are fair and limited to their services.
Steps (30):
1. Secure executive mandate and budget
Get a written CEO/CTO mandate that incident command is a company process, not a team hobby.
The mandate must state that **paid on-call** is required for production ownership. "No pager for other teams' code" is solved by named ownership, not by refusing coverage.
- Approve budget for tooling, stipends, training, and a dedicated program lead for six months.
- Name an executive sponsor (CTO or VP Engineering) who will chair the weekly incident review.
- Tie the clock to SOC 2 Type II: the process must be live in about 10 weeks so ~6 months of evidence remain.
- Commit that the CEO will hear about outages from this process, not from customers.
2. Form the working group and decision rights (depends on: 1)
Stand up a small group that can decide. Do not form a 28-team committee.
**Core seats:** SRE/platform lead, payments/ledger engineering manager, support lead, legal/compliance, HR, one rotating team EM, and a program manager.
- Meet twice a week for 10 weeks, then weekly.
- RACI: the group proposes; the sponsor decides in 48 hours; teams implement.
- Publish one Slack channel and one source-of-truth doc on day one.
- Time-box design to four weeks. Ship v1 rather than wait for consensus.
3. Inventory services, owners, and on-call gaps (depends on: 1)
Build a living catalog of all ~180 services: owning team, criticality, current on-call, alert sources, and runbook link.
Walk the last 31 customer-impacting incidents and the two events where **nobody was in charge**. Record who detected, who led, time to mitigate, and which alerts fired.
- Tag each service as critical-path, customer-visible, or internal.
- List the 16 teams with no on-call and every orphan service with no owner.
- Map all six alerting tools and the 3,400 monthly alerts onto services.
- Flag the shared PostgreSQL ledger and two-region failover as named special cases.
4. Map regulatory and contractual notification duties (depends on: 1)
Legal and compliance list every duty an incident can trigger. Do not invent clocks that violate a contract.
Cover SOC 2 CC7, customer MSA/SLA credit terms, money-transmitter rules, NYDFS 23 NYCRR 500 if applicable, PCI if in scope, and breach clocks.
- Extract **notification timings** from the largest customer contracts (status page, named AM, written notice).
- Define when legal, regulators, insurers, or the board must be told.
- Feed these clocks into severity triggers and the communications playbook.
5. Approve paid on-call and incident pay (depends on: 1, 4)
Unpaid on-call is why 16 teams refuse the pager and why nights are uncovered. Fix the money before asking for coverage.
HR, legal, and finance design a New York–compliant package: weekly stipend for primary and secondary, extra stipend for company IC and comms, and **after-hours incident pay** or comp time.
- Treat exempt vs non-exempt staff explicitly under NY wage-hour rules.
- Put stipend in the next pay cycle after policy publish, not "later."
- Cap consecutive night weeks. Fund hiring if a team cannot rotate fairly (minimum six people for 24x7 primary plus secondary).
- Publish the package before any new rotation starts. This is the main answer to pager pushback.
6. Ratify four severity levels and their triggers (depends on: 3, 4)
Adopt a business-impact scale. Engineers do not invent severity in the moment.
**SEV-1:** material payments failure; ledger down or inconsistent; security or customer-data incident; both regions impaired; or many customers already in SLA-credit territory.
**SEV-2:** degraded payments or a contracted feature down for multiple customers; SLA at risk.
**SEV-3:** narrow or single-customer impact with a workaround; no fleet-wide SLA risk.
**SEV-4:** no customer impact; ticket only.
- SEV-1 pages IC, comms, scribe, owning SMEs, and an exec; war room in 5 minutes; status page in 10; AM outreach in 20.
- SEV-2 pages IC and owning SMEs; comms may be the IC; status page in 15 minutes; updates every 30 minutes.
- SEV-3 pages the owning team only; customer notice only if that customer is affected.
- Anyone may declare. Only the IC may downgrade. When unsure, start high.
7. Define incident roles and their authority (depends on: 6)
Four roles. Separate coordination from debugging so "who is in charge" cannot stall for an hour again.
- **Incident Commander:** owns severity, the room, the clock, and the next action. Does not write code. May page anyone, freeze deploys, and invoke failover. Staffed from a company-wide trained pool, not from the failing team.
- **Communications Lead:** status page, customers, AMs, execs, regulators. Speaks only from IC-approved facts.
- **Scribe:** timeline in the incident tool. Required for SEV-1 and SEV-2.
- **SME responders:** the owning team's on-call. They mitigate. They do not run the room.
Publish a one-page authority card. The IC stays in charge even if a VP joins.
8. Design 24x7 coverage without 28 night rotations (depends on: 3, 5, 7)
Do not put 28 teams on 24x7. That is what engineers are rejecting.
Use **three layers** so people page for their own code, plus a trained commander.
- Layer A — company IC and SEV-1 comms: 24x7; about 24 trained people; week-long primary and secondary.
- Layer B — critical-path team on-call (ledger, payments processing, auth, API edge, platform/Kubernetes, data stores): 24x7 primary plus secondary.
- Layer C — all other teams: business-hours on-call; after hours the IC pages the team EM, who has a written escalation list.
Platform on-call is the safety net for unknown-owner pages, never the permanent owner. Every service must have a named team within 60 days or be scheduled to shut off.
9. Set rotation, handoff, and load rules (depends on: 8)
Write mechanical rules so rotations are fair and load is visible.
Primary week, then secondary week, then at least two weeks off. No one holds primary on two rotations at once.
- Handoff is a 30-minute overlap covering open incidents, silenced alerts, and upcoming changes.
- Page-load SLO: p50 ≤ 4 pages per 12-hour night shift; p95 ≤ 10. A breach opens an alert-quality action.
- Require a shadow week before a first IC shift or a first critical-path rotation.
- Swaps live in the paging tool. Managers own coverage gaps, not the last person on the roster.
10. Write detection and escalation paths (depends on: 6, 9)
Customers currently detect 40% of incidents and median time to detect is 22 minutes. That is the first failure mode.
Detection path: synthetic full-payment probes in both regions, SLO burn-rate alerts, support-to-incident intake, and one customer callback path that can create a SEV.
- A page must be acked in **5 minutes** or it auto-escalates to secondary, then IC, then the EM, then the VP.
- Support may declare SEV-2 or higher without engineering permission.
- If ownership is unclear for 10 minutes, the IC keeps the incident and assigns a temporary owner. Never wait.
- An exec bridge auto-opens for every SEV-1 at T+15 minutes.
11. Write internal, customer, and regulator communications (depends on: 4, 6, 7)
Stop "whoever is around" from writing the status page. Comms follow the clock, not convenience.
Timings from declaration:
- Internal war room: immediate. Exec summary for SEV-1/2 at 15 minutes, then every 30 minutes.
- **Public status page:** SEV-1 in 10 minutes, SEV-2 in 15. Updates at least every 30 minutes until resolve. Templates only. No speculation.
- Account managers get an affected-customer list and a script at T+20 minutes for SEV-1/2.
- Resolve notice and credit assessment within one business day.
- Comms pages legal on SEV-1 security, ledger integrity, or any outage that will breach contractual notice. Legal owns outbound regulatory letters; the IC owns facts.
12. Standardize blameless postmortems and action tracking (depends on: 6)
A written postmortem is mandatory for every SEV-1 and SEV-2 within 5 business days. SEV-3 if the IC or EM requests it.
Use one template: timeline, customer impact (volume, duration, credits), detection gap, what went well, what did not, process-focused five whys, and numbered actions with owner and due date.
- Review is **blameless** and scheduled. The IC attends. The exec sponsor reads every SEV-1.
- Actions live in one tracker, not in the doc. No action without an owner and a date. Default due date 14 days; 30 days max unless architecture work with a milestone.
- Close rate is a published metric. The old 11-of-64 pattern is a process failure.
13. Set alert quality rules that make paging acceptable (depends on: 3, 6)
3,400 alerts a month and 85% noise is why on-call feels like punishment. Pages are a product with a quality bar.
A page (not a ticket) must map to a customer-facing SLO or a hard dependency of one. It must have an owner team, a runbook link, and a default severity. It must be actionable at 3 a.m. by the person who is paged.
- Ban parallel paging from six tools. One paging policy: symptom-based; burn-rate preferred over raw thresholds.
- Every team gets a monthly noise budget. Exceeding it is a sprint task, not heroics.
- A human may silence a flapping alert only with a linked ticket.
14. Publish Incident Management Policy v1 (depends on: 5, 6, 7, 8, 9, 10, 11, 12, 13)
Collapse the design into a short policy people will open during an outage.
Ten pages or fewer, plus one-page cards for severity, roles, and comms timings. Host it where the incident tool can link it.
- Include the compensation summary and the rule: you are **not on-call for other teams' services**.
- Version it. v1 is mandatory from the pilot start date.
- Legal, HR, and the exec sponsor sign. Announce in all-hands, not only in Slack.
15. Implement a single incident command tool (depends on: 7, 10, 11)
Put one tool in the path that creates the room, pages roles from severity, records the timeline, and prompts status-page updates.
Requirements: Slack (or equivalent) incident bot, severity in one click, role assignment, stakeholder groups, and timeline export for postmortems and auditors.
- Integrate with the pager so IC, comms, and SME pages are automatic.
- Retain artifacts at least one year for SOC 2.
- Ad-hoc Zoom/Slack threads are no longer the system of record.
16. Consolidate six alerting tools onto one pager (depends on: 9, 13)
Pick one paging product. Connect existing monitors to it. Migrate **pages** first, tickets second.
- Inventory every page-producing rule. Delete or downgrade the noisy majority in S23.
- Route by service label → owning team schedule → the escalation policy from S10.
- IC and comms schedules live in the same product.
- Set a hard date after which pages outside the chosen tool are not valid on-call obligations.
17. Operationalize the status page and AM path (depends on: 11, 15)
Put the status page behind the Comms Lead role. Use templates for investigating, identified, mitigating, and resolved.
Subscribe AMs and customers whose contracts require it. Generate the affected-customer list from the incident (tenant, region, payment method).
- Dry-run a SEV-2 update before the pilot goes live.
- Record every public update in the incident timeline for the audit.
- Define partial vs full-outage wording so impact cannot be understated.
18. Stand up one postmortem repo and action board (depends on: 12)
Create the template, the filing location, and one Jira/Linear board with states: open, in progress, blocked, done, and won't-do (with exec reason).
Wire the incident tool so a SEV-1/2 automatically opens a draft postmortem and action tickets.
- Each week the program manager reports actions past due to the exec sponsor.
- If it is not in the board, it does not exist.
19. Assign owners and write critical-path runbooks (depends on: 3, 9)
Close the ownership gaps that cause "pager for other people's code."
Every production service gets a team in the catalog. Unowned services get an owner in 30 days or a decommission date.
- Write runbooks for ledger Postgres, regional failover, payments API, auth, and the Kubernetes control plane: symptoms, dashboards, mitigate vs escalate, customer impact.
- Link runbooks from alerts. If there is no runbook, the alert cannot page at night unless the EM accepts the gap in writing.
20. Detect payments failures before customers do (depends on: 3, 13)
Build or finish end-to-end synthetics: create payment, ledger write, webhook, both regions, both critical payment methods.
Alert on SLO burn, not on a single 500. Page SEV-2 or SEV-1 from these probes. This is the fastest lever on the 22-minute MTTD and the 40% customer-detected rate.
- Add ledger lag, replication, disk, and failover-readiness as first-class pages to the ledger team.
- Every postmortem asks first: **why did a customer see this first?**
21. Train the first cadre of ICs, comms leads, and scribes (depends on: 7, 14, 15)
Train about 24 ICs and 12 comms leads before the pilot. Classroom plus a recorded shadow of a simulated SEV-1.
Curriculum: severity, authority card, tool, comms timings, when to call legal, how to run a room of 20, how to hand off at 2 a.m.
- Certification: pass a tabletop. No certificate, no rotation.
- Recertify yearly and after any SEV-1 where process failed.
- Managers of ICs protect calendar time. This is part of the job.
22. Pilot on the payments critical path for six weeks (depends on: 5, 14, 15, 16, 17, 18, 19, 21)
Go live with policy, tool, paid rotations, and IC coverage for ledger, payments, API, platform, and support intake.
Keep old paths as backup for one week, then cut over. Real incidents use the new process only.
- Staff the program lead in every SEV-2+ as coach, not as secret IC.
- Collect friction daily. Fix tooling and wording in 48 hours.
- Expansion gate: named IC in under 5 minutes, first status update on time, no unpaid pages, postmortem filed.
23. Cut alert noise with a forced burn-down (depends on: 13, 16)
Give every team a numbered list of their noisiest alerts. Move noise from 85% to **under 15%**, and monthly pages from 3,400 toward 500.
Each sprint, critical-path teams must delete, debounce, or convert to ticket a fixed quota. Platform provides burn-rate and grouping libraries.
- Publish a weekly noise leaderboard. Shame systems, not people.
- After eight weeks, any page without a runbook or with >30% false pages in 14 days is auto-downgraded until fixed.
24. Roll remaining teams onto the model by risk (depends on: 22)
After the pilot gate, add teams in waves of four to five every two weeks. Highest customer-impact first.
Layer C teams get business-hours schedules and the EM night list. Do not surprise anyone with a pager.
- Each wave: ownership confirmed, alerts routed, runbooks for paging alerts, paid rotation in HR, one tabletop.
- Finish all 28 teams at least five months before the SOC 2 report date so the observation window covers the company.
- Orphan services still unowned at wave end are escalated to the sponsor for shutdown or reassignment.
25. Run tabletops and multi-region game days (depends on: 21, 22)
Schedule a monthly tabletop: SEV-1 ledger, SEV-1 region loss, SEV-2 degraded payments, customer-detected incident, and a "who is in charge" chaos drill.
Quarterly game day: fail a region or a ledger replica in staging or a controlled production drill.
- Include support, AMs, legal, and an exec. Process fails if only engineers show up.
- Capture actions on the same board as real postmortems.
- Use results as SOC 2 evidence that IR is tested.
26. Resolve ownership fights and pager culture (depends on: 5, 8, 22)
Treat pushback as design input, not defiance. Repeat the contract in office hours: you carry a pager for **your** services; IC is coordination; nights are paid; noise is a defect.
- EMs who cannot staff a fair rotation get headcount or have services reassigned. Do not run two-person 24x7.
- Publicly close the two historical "nobody in charge" incidents with what would be different now.
- Pulse-survey on-call at 60 and 120 days. If load or fairness is red, stop expansion until fixed.
27. Launch metrics, weekly review, and error budgets (depends on: 15, 22)
The CEO email exists because there was no operating rhythm. Stand up a dashboard and use it.
Track incident count by SEV, MTTD, MTTA, MTTM, customer-detected %, time-to-IC, time-to-status, pages per shift, noise %, SLA credits, postmortem SLA, and action close rate.
- **Weekly 30-minute review** chaired by the exec sponsor. Last week's incidents, open actions, alert SLO breaches. Mandatory for EMs of teams that had SEV-2+.
- Error budgets on critical SLOs: burn too fast and the team stops features and pays down reliability.
- Monthly report to the CEO so customer-first news does not recur.
28. Collect SOC 2 Type II evidence as a side effect (depends on: 4, 14, 22, 27)
Do not create a parallel "audit process." The real process is the evidence.
From pilot start, retain policy versions, on-call schedules, incident records, status-page history, postmortems, action tickets, training attendance, tabletop notes, and access reviews of the incident and pager tools.
- Map each artifact to SOC 2 CC7.2–CC7.5. Compliance owns the matrix; engineering owns completeness.
- Internal audit or a consultant runs readiness reviews at week 12 and week 20, not the week before the auditor.
29. Inspect and adapt the process after 90 days (depends on: 22, 27)
After 90 days live, revise v1 using data, not opinions.
Look at MTTD/MTTM, uncovered shifts, IC burnout, missed status updates, action close rate, and survey results.
- Change severity definitions if teams inflate or deflate them.
- Adjust Layer B vs C membership from real pages.
- Cut steps nobody follows. Add only what the last 90 days proved missing.
- Freeze v2 as the SOC 2 described process unless a defect is found.
30. Rehearse the audit and freeze the observation window (depends on: 24, 28, 29)
Six to eight weeks before the auditor, run a mock interview with an IC, a random engineer, support, and compliance.
Show one SEV-1 and one SEV-2 end-to-end: detect, page, command, comms, postmortem, closed actions.
- Fix evidence gaps immediately.
- Freeze process wording for the rest of the observation window; log exceptions.
- Brief the CEO and customer success with the metrics so the "outages we hear about from clients" story is retired.
--- PROPOSAL 5 (agent deepseek-v4-pro_initial_5, deepseek/deepseek-v4-pro) ---
Estimated complexity: high
Success metrics: - Customer-detected incidents decrease from 40% to less than 15% within six months.
- Median time to detect (MTTD) is under 5 minutes for SEV1 and SEV2 incidents.
- Median time to mitigate (MTTM) is under 60 minutes for SEV1 and under 2 hours for SEV2.
- Alert noise decreases from 85% to below 10% within three months.
- 100% of SEV1 and SEV2 incidents have a completed blameless postmortem within 5 business days.
- 100% of postmortem action items are tracked with an owner and due date; 90% are completed on time.
- 24x7 on-call coverage achieved across all 28 teams with no unpaid on-call.
- 99% of pages are acknowledged within 5 minutes.
- SLA credits paid reduce by at least 50% over the next 12 months.
- SOC 2 readiness: all incident response controls are documented, tested, and evidence is produced by month 7.
Steps (14):
1. Baseline current incident response and align stakeholders
Collect data from the last 12 months of incidents, all six alert tools, on-call practices, and team interviews. Identify gaps against the target incident management process and secure executive sponsorship.
- Gather incident timeline, detection source, mitigation time, customer impact, SLA credit, and postmortem status for all 31 incidents.
- Survey 28 teams on on-call burden, alert quality, and operational pain.
- Map current tools, escalation paths, and communication workflows.
- Create baseline metrics and a stakeholder map with executive sponsor and audit owner.
2. Define severity levels and response triggers (depends on: 1)
Define a four-level severity scale with objective business-impact criteria so any engineer can classify an incident consistently.
- SEV1: widespread transaction processing outage, data breach, security incident, or severe SLA breach; triggers full incident command, executive notification, and a 5-minute status page update.
- SEV2: major feature outage, significant degradation without workaround, or customer financial risk; triggers incident commander, full communications role, and status page updates.
- SEV3: partial impairment with workaround or limited customer impact; triggers on-call response, internal communication, and optional status page update.
- SEV4: minor or internal issue, no customer impact; handled during business hours through ticketing.
- Include an escalation matrix showing who can declare, downgrade, and invoke regulatory or legal involvement.
3. Define incident roles and decision authority (depends on: 1, 2)
Define incident roles, responsibilities, and decision authority using RACI to remove ambiguity about who is in charge.
- Incident Commander: owns the incident, declares severity, and coordinates resolution.
- Communications Lead: owns internal and external messaging, status page updates, and account manager notifications.
- Scribe: maintains timeline, incident log, and postmortem notes.
- Subject-matter responders: diagnose and fix the incident; may come from multiple teams.
- Executive sponsor: optional for SEV1; customer liaison: handles account managers.
- Define decision rights for severity declaration, escalation, rollback, customer communications, and incident closure.
4. Design 24x7 staffing model across 28 teams (depends on: 3)
Design 24x7 coverage across 28 teams without overloading engineers. Use service-based on-call plus a central incident command pool.
- Each service or domain team assigns primary and secondary on-call for its own services.
- Create central incident commander, communications, and scribe rotations staffed from a trained incident response guild across all teams; use follow-the-sun between the two AWS regions and time zones.
- Define escalation layers: service on-call to team lead or manager to service owner to executive.
- Define handoff times, shadow shifts, and load balancing; target at most one week of on-call per engineer per month.
- Bridge the current 12-team paid on-call to 28-team paid coverage; no team remains uncovered.
5. Define on-call rotations, compensation, and alert quality rules (depends on: 4)
Define sustainable rotations, pay, and rules that eliminate noisy pages.
- Rotations: weekly or biweekly, at least one primary and one secondary, with 12-hour shifts where possible or 24-hour for low-volume services.
- Compensation: monthly on-call stipend for all on-call engineers, additional incident response bonus for after-hours work, and time off in lieu; align with market rates.
- Alert quality rules: every page must be actionable, have a runbook link, specify a service owner, include severity, and be based on SLO burn or known failure signals; no dashboard-only alerts.
- Noise budget: reject or downgrade non-actionable alerts; all pages must go to on-call only after suppression and deduplication.
- Weekly alert review removes the top noisy alerts.
6. Design detection, escalation, and alert routing (depends on: 2, 4, 5)
Define how incidents are detected, routed, and escalated so nothing waits on a human to notice.
- Consolidate the six alert tools into one alerting and paging platform with routing by service, severity, and tags.
- Detection sources: infrastructure metrics, application synthetic transactions, log-based anomalies, business transaction SLI monitoring, and customer-reported issues through support or account managers.
- Routing: alert is paged to service on-call within 30 seconds; primary must acknowledge within 5 minutes; if no ack, page secondary then on-call manager.
- Escalation timeouts: unresolved SEV1 escalates to service owner at 15 minutes and to leadership at 30 minutes; any engineer can escalate to the incident commander.
- Define customer-reported incident intake and classification in the same tool.
7. Define internal and external communication protocols (depends on: 2, 3)
Define communication channels, templates, and timing for internal, customer, and regulator audiences.
- Internal: dedicated incident Slack channel, internal status page mirror, and war room bridge for SEV1; incident commander and communications lead own these channels.
- Status page: SEV1 post within 5 minutes, updates every 30 minutes or on material change, resolution within 60 minutes of mitigation; SEV2 post within 15 minutes, updates hourly; SEV3 optional.
- Account managers: SEV1 and SEV2 notify account managers within 15 minutes with an approved customer-facing description and expected impact.
- Regulators: legal or compliance determines notification for data breaches, security incidents, funds availability issues, or regulatory reportable events; criteria and timing follow legal and regulatory requirements; communications lead coordinates.
- Use pre-approved message templates and an approval chain; no ad-hoc wording.
8. Define postmortem policy and action tracking (depends on: 3)
Define mandatory blameless postmortems and action tracking.
- Mandatory for all SEV1 and SEV2 incidents, and any SEV3 that breaches SLA or is customer-detected.
- Format: impact, timeline, root causes, contributing factors, detection and response gaps, what worked well, and action items.
- Blameless: focus on system and process causes, not individual blame; use trained facilitators.
- Ownership: each action has an owner, due date, and tracking ID in a single backlog.
- Review postmortems at the weekly incident review; track action closure; expect 100% completion.
- Complete postmortems within 5 business days for SEV1 and SEV2 incidents.
9. Define metrics, dashboards, and review cadence (depends on: 2, 3, 8)
Define metrics and review cadence to measure process health.
- Metrics: MTTD, MTTM, customer detected percentage, alert noise percentage, on-call response time, on-call load, SLA credits paid, and postmortem action completion.
- Dashboards: real-time operational dashboard for on-call engineers and management.
- Weekly incident review: review all SEV1 and SEV2 incidents, action items, and noisy alerts.
- Monthly trends with leadership; quarterly review against SLOs and audit controls.
- Success thresholds: MTTD under 5 minutes, MTTM under 60 minutes for SEV1, customer detected under 15%, and alert noise under 10%.
10. Configure incident tooling and integrations (depends on: 5, 6, 7, 8, 9)
Implement and integrate the tools that automate the defined process.
- Aggregate alerts from the existing six tools into PagerDuty, Opsgenie, or a similar platform.
- Configure on-call schedules, escalation policies, and paging targeted at service owners.
- Integrate status page API for automated or one-click updates.
- Add Slack commands to declare incidents, start war rooms, assign roles, and post status updates.
- Integrate runbook and service catalog access; create postmortem templates in Jira or Notion with action item tracking.
- Ensure audit trails and role assignments are logged for SOC 2.
11. Pilot with 2-3 volunteer teams and iterate (depends on: 10)
Run a controlled pilot before full rollout to validate and refine the process.
- Select 2-3 volunteer teams with representative services and on-call patterns.
- Run the new severity, roles, on-call, alerting, and communication process for 2 weeks.
- Track metrics and gather feedback from on-call engineers, incident commanders, and communications leads.
- Iterate severity thresholds, alert rules, templates, and runbooks based on findings.
- Exit criteria: no SEV1 without a declared incident commander, alert noise below target, and positive on-call survey results.
12. Train and certify all 28 teams (depends on: 11)
Prepare all 28 teams so the process is not a paper exercise.
- Create role-specific training: all engineers, incident commanders, communications leads, scribes, subject-matter responders, and managers.
- Deliver interactive workshops and online modules.
- Run at least two game-day drills per quarter simulating SEV1 and SEV2 incidents, testing escalation, communication, and postmortem.
- Certify incident commanders and communications leads; only certified staff can hold those roles in rotations.
- Include scenario-specific training for account managers and legal or regulatory handling.
13. Full rollout and enforce process (depends on: 12)
Move every team onto the new process and retire old habits.
- Migrate all 28 teams into consolidated on-call schedules and alert routing.
- Retire old alerting tools or methods that bypass the process.
- Activate status page templates and postmortem action tracking globally.
- Monitor the first 30 days with a daily operational review and quick fixes.
- Communicate rollout status to leadership and all engineers; provide office hours and support.
14. SOC 2 readiness and continuous improvement (depends on: 13)
Map the process to SOC 2 and keep improving after go-live.
- Map each process component, including severity, roles, detection, communications, postmortems, and metrics, to SOC 2 control requirements.
- Produce runbooks, evidence of training, incident logs, postmortems, and action item tracking.
- Conduct an internal readiness assessment at month 6 and remediate gaps before the audit at month 8.
- Maintain weekly and monthly metrics reviews to drive continuous improvement.
- Refresh on-call compensation and alert quality rules based on operational data.
HOW THE ROUNDS WENT (from the round analyses):
Round 0: All five agents converge on the same skeleton — exec mandate, baseline of the 31 incidents, a 4–5 level severity scale, separated IC/comms/scribe/SME roles, paid on-call, one paging tool, synthetic payment probes, timed status-page updates, mandatory blameless postmortems with tracked actions, pilot-then-waves, SOC 2 evidence as a by-product — so they differ mainly in staffing design, pace, and operational depth. P1 and P4 are the most granular (29–30 steps), P2 is the most payments-literate and the only one with day-1 interim controls, P3 is mid-weight with some internal inconsistencies, and P5 is the thinnest and least dated.
shared: **Command separated from debugging**: every plan defines IC / comms lead / scribe / SMEs, forbids the IC from typing, allows anyone to declare, and lets only the IC downgrade (P1 S6, P2 S5, P3 S4, P4 S7, P5 S3).
shared: **Pay for on-call and page only for owned code** as the explicit answer to the pushback, with a service-ownership catalog and criticality tiers as the routing foundation (P1 S4/S9, P2 S3/S6, P3 S4/S5, P4 S5/S18, P5 S4/S5).
shared: **Alert-quality gate plus customer-journey detection**: every page needs owner, runbook, severity and SLO linkage; synthetic end-to-end payment probes from both regions; consolidation of the six tools into one pager (P1 S10/S11/S12, P2 S7/S8, P3 S6/S9, P4 S13/S16/S20, P5 S5/S6/S10).
shared: **Mandatory postmortems with enforced action tracking**: SEV1/SEV2 always, 3–5 day drafts, single template, named individual owners, due dates in one tracker, escalation of overdue items, and SOC 2 evidence produced from the live process rather than reconstructed.
differences: **24x7 staffing design** is the biggest split: P1 S7/S8 uses a 25–35-person Duty IC pool plus tiered team obligations; P2 S5 collapses 28 teams into 8–12 domain rotations of ≥6; P3 S4 adds an SRE Tier 2 that absorbs infra pages and floats a managed overnight triage vendor; P4 S8 puts Layer C teams on business hours only with an EM night list; P5 S4 invokes "follow-the-sun between the two AWS regions", which does not create time-zone coverage.
differences: **Pace and interim cover**: P2 S2 installs interim IC staffing, one declaration path and a provisional severity guide within 7 days and triages the 53 open actions immediately; P4 S1 targets live in ~10 weeks and S30 freezes the process before the observation window; P1 runs design 1–6, pilot 7–12, rollout 13–24; P3 ends rollout at week 24 with a month-5 mock audit; P5 gives almost no dates and pilots only at step 11 of 14.
differences: **Payments/ledger rigour**: P2 S10 guards against split-brain, duplication and replay, requires reconciliation and backlog processing before "resolved", and separates mitigation from resolution; P1 adds ledger double-entry assurance (S11), a regulator matrix incl. NYDFS/FinCEN/sponsor banks (S16) and an automated SLA-credit workflow (S17); P4 S4 extracts notification clocks from contracts before setting severity triggers; P5 S7 defers regulator timing entirely to legal with no clock.
differences: **How failure and resistance are handled**: P1 S28 is a standing risk register with pre-committed fallbacks (volunteer shortfall, comp not approved, pruning a real alert); P4 S26 stops wave expansion if the 60/120-day pulse is red; P2 S17 requires exception records with owners and expiry dates; P3 and P5 have no comparable contingency mechanism.
Proposal 1: 29-step program run by a Head of Reliability with exec steering, starting from a forensic re-coding of all 31 incidents and the six-tool alert estate. Builds a service catalog with Tier 0–3, a five-level severity scale with auto-escalation on any ledger-touching incident, a 25–35-person Duty IC rotation plus tiered team on-call, and an on-call readiness bar that blocks a service from paging anyone until runbooks exist. Unusually complete on money and law: stipend ranges, FLSA/NY review, a regulator playbook, and an automated SLA-credit workflow targeting $1.3M → <$400k. Ends with a month-6 dry-run audit over 15 sampled incidents and a year-two maturity roadmap.
Proposal 2: 19 dense steps that put interim controls live in the first 7 days — one declaration path, staffed interim commanders, and triage of the 53 open historical actions — before any tool consolidation. Uses SEV0–SEV3, domain rotations of ≥6 instead of 28 fragile ones, compensation and fatigue rules (one week in six, paid recovery day, self-declared unfitness) and shadow-mode alerts for 7 days before they can page. Strongest on payments correctness: dual-control preserved during ledger recovery, split-brain/duplication guards, reconciliation required before "resolved". Milestones are day-anchored (7/14/30/60/90/120/180) with a month-7 mock audit and metric-gaming checks.
Proposal 3: 13 steps run by a new Incident Management Office and a 28-manager Council, starting with a baseline plus peer benchmarking against Stripe/Adyen. Distinctive is the three-tier staffing where a 6-person SRE Tier 2 absorbs Kubernetes/PostgreSQL pages and a Ledger Duty Officer joins any ledger-touching incident, explicitly to defuse the cross-team pager objection. Compensation is fully costed ($500/$250 weekly, $75–150 per page, $300 SEV1 bonus, ~$420K/yr). Weak spots: a per-team budget of 100 alerts/month contradicts the <600/month target, "56 certified ICs and all 260 engineers trained in 14 weeks" is optimistic, and step 12 funnels all eleven prior steps into one rollout.
Proposal 4: 30 short, opinionated steps that time-box design to four weeks and aim to be live in ~10 weeks so five-plus months of SOC 2 evidence accumulate. Organises coverage as Layer A (24 trained company ICs, 24x7), Layer B (critical path 24x7), Layer C (business hours plus an EM escalation list), with mechanical rotation rules and a page-load SLO of p50 ≤4 per night shift. Maps contractual and regulatory notification clocks (S4) before setting severity triggers, runs a forced noise burn-down with a weekly leaderboard, and gates wave expansion on a red on-call pulse survey. Closes with a mock auditor interview and a deliberate wording freeze for the observation window.
Proposal 5: 14 conventional steps covering the full brief: SEV1–SEV4, RACI roles, service on-call plus a central IC guild, stipends and TOIL, one paging platform, comms timings, mandatory postmortems, metrics and a month-6 readiness assessment. Correct in outline but the thinnest on specifics — no dates until months 6–8, no wave plan, no page budget, no risk register, and no targeted answer to the "other teams' code" objection beyond a survey. Two realism problems: a 5-minute SEV1 status-page post, and a 24x7 model built on "follow-the-sun between the two AWS regions" for a single New York organisation. The pilot lands at step 11 with training and full rollout after it, leaving little room to iterate before the audit.
Round 1: All five plans converged on a common architecture — 7-day interim command, forensic baseline, service catalog with tiers, 4–5 level severity, IC/Comms/Scribe/SME roles, a central command pool plus 8–12 domain rotations, paid on-call before mandatory nights, page budgets, money-path detection, one pager, 3/5/10-day postmortems with tracked actions, critical-path pilot, gated waves, month-6/7 audit rehearsal. The real remaining spread is in severity thresholds, compensation concreteness, and how much program scaffolding (policy v1, risk register, 90-day revision) each keeps.
differences: **Severity design.** P2 (S5) alone gives quantitative declaration guardrails (>10% payment failures for 5 min = SEV1; 1–10% = SEV2) plus a SEV0 crisis tier and no time-based auto-escalation; P5 (S5) uses five levels with aggressive clocks (SEV2 unresolved >1h becomes SEV1); P1/P3/P4 (S6) use SEV1–SEV4 with ledger-touch and 2-hour auto-escalation.
differences: **Compensation and load caps.** P1 (S9) and P5 (S9) publish dollar figures (~$800–1,200/primary week, $150/night page); P4 (S9) budgets $600k–1M with structure only; P2 (S8) and P3 (S7) defer amounts to HR in 14 days. Caps diverge: one primary week in six (P1, P2, P4) versus one in four (P5 S8; P3 S7 says "four or six depending on staffing").
differences: **Program scaffolding.** P1 keeps a standalone Policy v1 with exception register (S24), risk register (S31) and a 90-day inspect-and-adapt (S32); P3 keeps a risk register (S26). P4 dropped its own Policy v1 and 90-day revision; P2 and P5 have no contingency/risk step at all.
differences: **Interim protection and density.** P1 (S2), P2 (S2), P4 (S2) and P5 (S1) stand up interim command in 48h–7 days; P3 only lists "interim controls week 1" in S1 with no step behind it. On density, P4 is tightest at 24 steps, P1 heaviest at 32 with overlapping alert work (S10 standard vs S28 burn-down) and an 11-dependency join at S24.
influences: P2's 7-day interim minimum controls, retroactive pay for interim duty, triage of the 53 open actions and the daily 15-minute ops review (R0 S2) were taken by P1 (S2), P4 (S2) and P5 (S1); P3 left it as a single timeline bullet.
influences: P4's three-layer coverage and the "deal" fairness contract (R0 S8, S26) were taken by P1 (S8, S29), P3 (S8, S24) and P5 (S7–S8); P2 reached the same result via 8–12 responder domains (S7) without the layer naming.
influences: P1's forensic baseline, page budget, SLA-credit workflow, Incident Review Board and gated waves (R0 S2, S10, S17, S18, S25) were adopted almost wholesale by P3 (S2, S9, S14, S16, S22) and selectively by P4 (S3, S10, S16, S17, S22) and P2 (S3, S11, S15, S18, S21).
influences: P2's safety rails travelled widely: shadow-mode alerts and "never silently disable", two weeks of verified operation before retiring a legacy path, human acknowledgement required, mitigation-vs-resolution semantics and ledger dual-control (R0 S1, S4, S7, S8, S9) now appear in P1 (S6, S7, S10, S12), P4 (S7, S10, S12, S14), P3 (S9, S11) and P5 (S13).
influences: Nobody took P3's outsourced overnight triage vendor or its per-page pay table (R0 S4, S5) — P3 dropped both itself — and nobody kept P5's 5-minute SEV1 status-page deadline (R0 S7); all five settled on 15 minutes for the top availability tier.
Proposal 1 (improved): Closed its biggest R0 gap — six weeks of design with no protection — by adding a 7-day interim command bridge (S2). Replaced "14–18 teams at 24x7" with a Layer A/B/C model over 10–12 domains (S8), added live-execution doctrine (S14), Policy v1 (S24), a noise burn-down campaign (S28) and staged day-90/month-6 metric targets. Now the most complete plan, but at 32 steps it is also the heaviest.
Proposal 2 (improved): Grew from 19 to 24 steps by adding the pieces it was missing — readiness/runbook standard (S9), detection uplift (S12), pilot (S20), waves (S21), exercises (S22) — while keeping its distinctive strengths: compliance designed on day one (S4, before severity in S5), a balanced anti-gaming scorecard (S18), and now numeric severity guardrails.
Proposal 3 (improved): Doubled from 13 to 27 steps and adopted most of P1's R0 architecture: baseline, listening tour, layered staffing, escalation ladder, regulatory playbook, credits, runbooks, exercises, gated pilot and waves, risk register, SOC 2 dry-run. Far more complete and better sequenced, but largely derivative and it lost its own concrete compensation numbers.
Proposal 4 (improved): Consolidated 30 loose steps into 24 denser ones (postmortems and actions merged into S17, runbooks and readiness folded into S14, all comms into S15) and imported the financial and training machinery it was missing. Remains the most readable plan, but it dropped two of its own best R0 steps.
Proposal 5 (improved): Expanded from 14 thin steps to 26, closing nearly every R0 gap: interim duty officer, forensic baseline, listening tour, escalation ladder, regulatory playbook, credits, runbooks, exercises, pilot, waves and SOC 2 dry-run. The gains come almost entirely from copying P1's R0 structure and P2's SEV0, so it is now complete but the least differentiated plan.
Round 2: All five plans converged onto a near-common template: four-level severity, a central command corps plus 10–12 domain rotations plus business-hours Layer C, paid on-call gated before any mandatory night rotation, a two-page-per-week budget, one paging platform, 3/5/10-day blameless postmortems, wave rollout with gates and a month-7 mock audit. Real differentiation narrowed to severity anchoring (P4's integrity flag, P1's payment anchors), plan weight and audit sequencing, while P3 essentially re-issued P1's round-1 plan.
differences: **Severity anchoring.** P4 step 6 attaches a financial-integrity/security flag to any level (SEV0's forcing function without a fifth level); P1 step 6 adds payments anchors (value at risk per hour, settlement-deadline proximity, named accounts); P2 step 6 demotes percentages to guardrails; P3 step 7 and P5 step 6 stay purely qualitative. P1/P4 also tighten the SEV3 postmortem trigger (customer-detected, >2h, repeat), while P3/P5 keep "postmortem on request".
differences: **Audit sequencing.** P4 confirms the observation window in week 1 (step 1), P2 in step 3, P3 in step 5; P1 leaves it in step 30, which depends on steps 24/26/28, and cuts independent testing from two rounds to one at month 5. P2/P3/P5 keep tests in months 4 and 6.
differences: **Coverage economics.** P1 alone funds an overnight Duty Triage Desk owning the first ten minutes (step 8) but never sizes or budgets it; P4 alone shows the arithmetic (260 engineers sustain 10–12 rotations, not 28, step 8); P2 alone refuses to merge small teams for schedule convenience, scopes the six-responder rule to direct 24x7 rotations (steps 8, 22) and keeps a surge roster for concurrent incidents.
differences: **Weight and pace.** P1/P3 run 32 steps with standalone noise campaign, Policy v1 and risk register; P4 compresses to 26 with an 11-bullet comms mega-step (15) and the risk register demoted to last (26); P5 merges regulatory+credits (16) and postmortems+actions (17) and has no policy-publication step; P2 has 26 steps, no risk register at all, and the fastest schedule (Tier 0/1 by week 10, all teams by week 18 versus week 24 elsewhere).
influences: P2 and P5 abandoned their SEV0 tiers for the SEV1–SEV4 scale of P1/P3/P4 (P2 step 6, P5 step 6); P4 instead converted SEV0 into a financial-integrity/security flag (step 6) and P1 into payments anchors (step 6) — the round's one genuine design argument, resolved three different ways.
influences: P3 imported P1's round-1 plan almost wholesale (21 matching steps: Layer A/B/C 9, noise burn-down 26, Policy v1 24, risk register 31, 90-day inspect 32) and grafted on P2's day-one compliance step (P3 step 5 from P2 step 4) and P2's fake-redundancy check (P3 step 12 from P2 step 12).
influences: P5 took P1's Layer A/B/C coverage (step 8), readiness bar and ledger blast-radius workstream (step 10), risk register (step 25) and culture step (23), and tightened its own 1-in-4 rotation cap to P1's 1-in-6 (step 9).
influences: P2 took the standalone listening tour/fairness contract from P1 step 4, P3 step 3 and P4 step 4 (now P2 step 4) and concrete pay bands from P1/P5 step 9; in the other direction P4 took P2's week-1 observation-window confirmation (P4 step 1) and P1 added P2's 99.95% availability metric and an Internal Audit seat (step 1).
influences: Nobody took P5's named vendor shortlist (PagerDuty/incident.io/FireHydrant, round-1 step 12) — P5 dropped it itself; nobody adopted P2's numeric declaration thresholds (10% / 1–10% failure rates), which P2 also demoted; and only P2 (step 8) still plans for simultaneous incidents with a surge roster.
Proposal 1 (improved): Structure is unchanged at 32 steps; the gains are inside steps. Step 8 adds a paid overnight Duty Triage Desk, step 6 adds payments-specific severity anchors, step 11 adds an incident-replay test with "replay coverage" as a leading metric. Independent control testing shrank from two rounds to one.
Proposal 2 (improved): Renamed the scale to SEV1–SEV4, ending the SEV0/SEV1 ambiguity, and split seven previously buried topics (fairness contract, alert contract, customer comms, actions, policy, metrics, sustainability) into their own steps. Compensation moved from "fixed weekly stipends" to actual dollar bands and an annual budget.
Proposal 3 (improved): Effectively re-issues P1's round-1 plan: 21 of its 32 steps match P1 and only 5 match its own previous version. It gains the machinery it lacked (Policy v1, noise campaign, risk register, 90-day inspect) plus a day-one compliance step from P2, but contributes nothing new and leaves two internal inconsistencies.
Proposal 4 (improved): Compressed to 26 steps while adding the round's neatest severity idea and the only staffing arithmetic. Gains: integrity/security flag (step 6), 260-engineer rotation math (step 8), week-1 audit clock (step 1), postmortems split from action tracking. Losses: the change-management step disappeared and the risk register is demoted to the end.
Proposal 5 (improved): Dropped SEV0 for a four-level scale, adopted P1's Layer A/B/C model, metric set and risk register, and added the readiness bar, simulations and a culture step. Still 26 steps, but two of them are now overloaded and the policy-publication step is missing.
Round 3: Round 3 converged hard: all five plans now run the same skeleton (charter → 7-day floor → forensic baseline → fairness contract → catalog/tiers → severity → roles → three-layer 24x7 → paid on-call → alert budget → detection → single paging platform → escalation → comms → regulatory/credits → postmortems/actions → training → exercises → policy → pilot → gated waves → metrics → SOC 2 → 90-day inspect), with near-identical numbers ($1,000 primary week, 2 out-of-hours pages/responder/week, 15% reserved capacity, 30 ICs / 18 comms leads). Differentiation now comes from a few genuine additions — P1's change intelligence and Support-as-detection tier, P4's Duty Triage Desk carve-out and vendor-incident class, P2's defined noise/actionable metrics and severity modifiers — while P3 and P5 mostly merge material authored by P1 and P4.
differences: **Unique content**: only P1 has change intelligence/deployment safety (step 13), Support and account managers as an instrumented detection tier (step 20) and a concurrency doctrine with a Multi-Incident Coordinator (step 8); only P4 has a vendor-incident class for processor/bank/cloud failures (step 13) and a rule for who writes the 3 a.m. status page (step 8); only P2 defines "actionable page" and "noise" and uses five modifiers instead of flags (steps 12, 6). P3 and P5 contribute nothing the others do not have.
differences: **Overnight first line**: P1 (step 9), P4 (step 8) and P5 (step 8) fund a paid Duty Triage Desk; P4 adds the safety carve-out that it never holds suspected ledger, payment-halt or security pages. P2 (step 8) and P3 (step 9) refuse a triage desk and route unknown-owner pages to Duty Command plus platform, P2 additionally reclassifying lower-tier services that can cause severe overnight harm.
differences: **Schedule**: P2 (step 1) compresses design to weeks 1–4, pilot 5–10, rollout by week 18, control tests in months 4 and 6; P1, P3, P4 and P5 keep design weeks 1–6, pilot 7–12, rollout 13–24 and a month-5 internal test. P2 buys a longer observation window at the cost of a 4-week design for 180 services.
differences: **Failure handling and metric honesty**: P1 (step 34), P3 (step 28) and P5 (step 29) keep a standing monthly risk register; P4 compresses contingencies into bullets in step 26; P2 has none. On claims, P2 targets "no unresolved material exception" and P1 warns credits may rise before they fall (step 22), while P3 and P5 still promise "SOC 2 passes with zero exceptions".
influences: P1's replay test (R2 step 11) went everywhere: P2 step 13 plus a replay-coverage metric, P3 step 12, P4 steps 3 and 11 (as a "replay catalog" with a gap owner), P5 step 11. It is now the field's standard proof that detection actually improved.
influences: P4's severity flag (R2 step 6) was taken by P1 (step 7, financial-integrity/security flag forcing dual control and the reportability checkpoint) and generalised by P2 into five modifiers — FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, VENDOR (step 6). P3 and P5 declined it and kept four levels plus auto-escalation.
influences: P1's Duty Triage Desk (R2 step 8) was adopted by P4 (step 8) and P5 (step 8); P4 improved it with the carve-out that ledger, payment-halt and security events page commander and domains in parallel, plus a half-day triage-desk training module (step 18). P2 and P3 explicitly did not.
influences: P4's "confirm the SOC 2 observation window in week 1" (R2 step 1) was taken by P3 (step 1) and P5 (step 3), and expanded by P1 into a full week-one audit, legal and evidence scoping step (step 4) covering privilege, legal hold and the obligation matrix.
influences: P4's staffing arithmetic (R2 step 8 — 260 engineers cannot sustain 28 night rotas, never a two-person rotation) was taken verbatim by P3 (step 9) and restated by P1 (step 9, ~100 of 260 engineers carry a night obligation at one week in six). Nobody took P2's R2 caution against merging small teams for schedule convenience — all five consolidate into 10–12 domains.
Proposal 1 (improved): Four substantive new steps and three hardening additions, all addressing gaps the field had left open. It is now the most complete plan, at the cost of being the longest (35 steps).
Proposal 2 (improved): It decomposed its round-2 mega-steps into discrete, individually gated steps and added the sharpest measurement definitions in the field. It remains the only plan without a risk register.
Proposal 3 (improved): A competent merge: it absorbed P4's audit clock and staffing math, P1's replay test and P2's accommodations, and split the readiness bar out of the execution doctrine. It adds nothing the other plans do not already have, and it created two new mega-steps.
Proposal 4 (improved): It imported P1's Duty Triage Desk and replay test and then improved both, and added two things nobody else has: a vendor-incident class and a rule for who writes the status page overnight. It stays the most compact serious plan at 26 steps.
Proposal 5 (improved): It closed two real holes from its round-2 version — there was no policy-publication step and no separated postmortem/action or internal/customer comms steps — but it did so by copying P1's round-2 plan almost step for step. It contributes no idea of its own and picks up none of P1's round-3 advances.
THE PROPOSALS OF THE FINAL ROUND (3), to assess:
--- PROPOSAL 1 (agent opus5_refine_1, anthropic/claude-opus-5) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90 a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95% internally and externally.
- Top-100 account outreach completed within 30 minutes of SEV1 in 95% of qualifying cases, using the approved briefing pack.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- Change-correlated incidents are identified within 10 minutes of declaration in 90% of cases, and the change-correlated share of incidents declines quarter over quarter.
- SLA credits fall to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, one concurrency drill and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and a tested read-only mode by month 6, with milestones reported to the board quarterly.
Steps (35):
1. Charter the programme: one owner, one mandate, funded, dated against the audit
Turn the CEO email into a chartered company programme with a single accountable owner and authority over all 28 teams. Incident response stops being a per-team preference and becomes a **company operating process**.
- Name the CTO as executive sponsor and a full-time Head of Reliability & Incident Management as accountable owner, supported by a programme office of three: programme lead, incident-platform engineer, reliability analyst.
- Form a decision group (Engineering, Platform/SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables on day one: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, named service ownership, and **paid on-call**.
- Publish the timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal control test month 5, mock audit month 7, SOC 2 fieldwork month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.3M/yr, training and exercises, plus 15% of engineering capacity reserved for reliability work.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence immediately.
- Publish a one-page interim severity card and a single declaration path: one chat command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim 24x7 Duty Incident Commander roster (primary plus backup) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document and one naming convention for every major incident, starting now.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is the design input, the executive narrative and the frozen "before" picture for the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the **detection backlog**.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Correlate incidents with deployments, config changes and feature flags to quantify how many started with a change we made.
- Quantify true cost beyond credits: failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost, churn risk on named accounts.
- Freeze the baselines in a signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Audit, legal and evidence scoping in week one (depends on: 1)
Design evidence as a by-product of operating, never as a reconstruction before fieldwork. Settle the scope and the legal handling of incident records now, not in month seven.
- Confirm with the external auditor the Type II observation window, the incident population definition, and the evidence they will sample. Everything from the day-7 floor onwards must count.
- Map incident response to the Trust Services Criteria with Compliance: CC7.2–CC7.5 (monitoring, identification, response, recovery), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), A1.2 (availability).
- Define the evidence set and where it is produced automatically: incident records, paging and acknowledgement logs, role assignments, status-page history, reportability decisions, postmortems, action closure with verification, training register, drill records, access reviews.
- Agree retention, confidentiality, legal hold and access rules. Decide with counsel which postmortem content is privileged and how privileged material is segregated **without making the ordinary postmortem secret**.
- Start the obligation matrix with Legal: NYDFS 23 NYCRR 500, state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network contractual windows, cyber-insurer notice.
- Begin monthly evidence sampling from month 1 so control operation is visible long before the audit.
5. Listening tour, resistance map and the written on-call deal (depends on: 1, 3)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written deal signed by the sponsor.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
- State the hard gate publicly: no new mandatory night rotation starts before compensation, training, runbooks and staffing rules are live.
6. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- Record for every service: one accountable team, engineering manager, chat channel, escalation policy, dashboards, runbook link, dependency list, regions and data stores.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record SLO, RTO, RPO, active-active or region-bound, failover method, and the dependencies that make nominal two-region redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
7. Severity standard, declaration rights and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning responders and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and responders paged, comms and scribe if customer-visible, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, credit-generating or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Attach a financial-integrity flag or a security flag to any severity. The flag forces dual control, Legal engagement and the reportability checkpoint without inventing a fifth level.
- Anchor on payments reality alongside error rates: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including a customer report.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
8. Roles, authority, concurrency and handover discipline (depends on: 7)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice to the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1 and SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, Finance and Vendor Management join on defined triggers rather than by invitation.
- Rules: command claimed within 5 minutes and stated in channel; distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- **Concurrency doctrine:** two simultaneous SEV1/SEV2 incidents activate the secondary commander and a surge roster; a designated Multi-Incident Coordinator arbitrates shared resources such as the ledger, the database platform and the deploy freeze.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
9. Three-layer 24x7 coverage: command corps, domain rotations, overnight triage desk (depends on: 5, 6, 8)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving roughly one primary week per person every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement and reconciliation, auth, API edge, Kubernetes platform, data and reporting, partner integrations). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Publish the staffing arithmetic: 10–12 domain rotations of six to eight people plus a 30-person command corps means roughly 100 of 260 engineers carry a night obligation, at about one week in six. Twenty-eight independent rotations would be unstaffable and is therefore rejected on the numbers.
- Teams too small for a fair rotation get headcount, service reassignment, or a time-limited executive exception. **Never a two-person 24x7 rotation.**
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: a human acknowledgement within 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
10. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 5, 9)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour, and a premium for the overnight triage desk.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per responder per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty since day 7 is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons without career penalty, and on-call load is published quarterly by team.
11. Alert quality contract and page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression can never hide an incident.
12. Detection uplift on the money path, validated by incident replay (depends on: 6, 11)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action and a missed-detection review.
13. Change intelligence and deployment safety on the critical path (depends on: 6, 12)
Most of these incidents start with something we changed. Make change the first hypothesis the tooling answers, and make changes safer to reverse.
- Stream every deployment, configuration change, feature-flag flip, schema migration and infrastructure change into the incident timeline with service labels and owners.
- Give the commander an automatic "what changed in the last 60 minutes on the affected journey" panel at declaration time.
- Require Tier 0/1 changes to be progressively delivered with a documented rollback that is tested, time-bounded and executable by the on-call responder without the author.
- Treat ledger schema migrations and settlement-affecting changes as a separate class: dual approval, rehearsed rollback, no deploys inside the settlement window.
- Enforce a change freeze during SEV1 and SEV2, lifted only by the commander and logged.
- Report change-correlated incidents monthly; a rising ratio is a signal to strengthen release safety, not to blame a team.
14. One pager, one incident record, one status page — migrated without a detection gap (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy paging path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access, MFA and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page delivered outside this platform creates no on-call obligation.
15. Escalation ladder and the five-minute command rule (depends on: 8, 9, 14)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts and support entitlements: AWS and database premium support, processors, sponsor banks, card networks.
16. Live execution doctrine and payments safety rules (depends on: 8, 14, 15)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed in advance rather than improvised.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and automatic reopening if impact recurs.
17. Service readiness bar, major-incident playbooks and ledger blast-radius reduction (depends on: 6, 9, 16)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human, and the shared ledger is the largest structural risk in the estate.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as a concentration risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers, stronger regional independence — with board-visible quarterly milestones.
- Enforcement: no readiness sign-off, no night paging alerts unless the manager accepts the gap in writing with an expiry date and a compensating control. Never respond to a gap by turning detection off.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
18. Internal communications protocol (depends on: 8, 14)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, Sales and Security.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change. First internal brief within 10 minutes of SEV1 and 15 of SEV2.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
19. Customer communications, status page and account-manager outreach (depends on: 7, 18)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; monitoring notice after mitigation; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, monitoring, resolution.
- Component-level status page mapped to customer journeys rather than internal service names, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1 and 60 of SEV2, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state observed impact, affected capabilities, any workaround and the next update time; **never speculate on cause, recovery time, data integrity or blame**.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
- Record any legally required restriction or delay of public detail, its approver, and the alternative stakeholder plan.
20. Support and account management as a detection and intake tier (depends on: 7, 12, 19)
Customers detected 40% of incidents first, which means the front line already holds the signal. Turn Support and account managers into an instrumented detection channel rather than a bystander.
- Give Support explicit declaration rights, a one-page trigger card, and a macro that opens an incident candidate directly in the incident platform.
- Automate the clustering rule: five similar tickets or calls in ten minutes auto-creates a triage incident assigned to the duty commander.
- Route credible partner, processor and sponsor-bank notifications into the same declaration path within five minutes.
- Generate the affected-customer list automatically from journey telemetry and the incident record, and push it to Support, CS and the account-manager briefing.
- Train Support and account managers on approved language and prohibit independent technical explanations to customers.
- Measure and publish "signal was in Support before it was in monitoring" as a detection defect, and feed each instance into the detection backlog.
21. Regulatory, partner and legal notification playbook (depends on: 4, 7, 19)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Complete the obligation matrix started in scoping and have counsel validate the triggers, deadlines, channels and submitting authority for each obligation.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with facts considered, decision-maker, timestamp and reassessment trigger.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, card networks, outside counsel, cyber-insurer, critical vendors.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the operational facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly inside the exercise programme, including one full regulator-notification simulation per year.
22. SLA credit workflow, true cost model and incentive guardrails (depends on: 7, 19)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component, including how partial degradation counts.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Guardrail against the perverse incentive: better detection will surface incidents that previously went unbilled, so credits may rise before they fall. Publish this expectation to the executive team in advance, and make it a written rule that no finance or commercial pressure may influence severity or declaration.
23. Blameless postmortem standard and Incident Review Board (depends on: 4, 7, 8)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger; any control failure.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning engineering director is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions where security, privacy or privileged content requires it.
24. Action ownership, reserved capacity and enforcement (depends on: 14, 23)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, accountable manager, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness through tests, telemetry, exercises or production evidence before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
25. Training and certification academy (depends on: 8, 15, 18, 23)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, the 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, contractual clocks, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days of joining a domain.
- Certification is valid 12 months and renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
26. Exercise programme: tabletops, game days, night drills and vendor rehearsals (depends on: 14, 17, 25)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander. First company-wide tabletop within 30 days.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times, plus a concurrency drill with two simultaneous incidents.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down. Rehearse joint escalation with AWS, a processor and a sponsor bank.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
27. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 15, 18, 19, 21, 23, 24)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, escalation and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix, Readiness Bar.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
28. Pilot on the payments critical path (depends on: 10, 12, 13, 14, 17, 25, 27)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement and reconciliation, PostgreSQL platform, Kubernetes platform, API edge, auth and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, action tracking, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The programme lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time in 95% of cases, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
29. Alert noise burn-down campaign (depends on: 11, 14, 28)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → under 1,500 pages a month by day 90, under 500 with noise below 15% by month 6.
30. Metrics, dashboards, review cadence and anti-gaming (depends on: 14, 23, 28)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time from first impact to detect, declare, commander, acknowledge, mitigate, resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors, change-correlated incident ratio.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
31. Wave rollout to all 28 teams with readiness gates (depends on: 28, 30)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the programme office for three weeks; the director signs the gate.
- **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard by team, tier and control gap.
32. Change management, fairness and pager culture (depends on: 5, 10, 28)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
33. SOC 2 evidence by design, internal control test and mock audit (depends on: 4, 27, 30, 31)
The real process is the evidence. Never build a parallel audit process, and never reconstruct records after the fact.
- Maintain the evidence set defined in scoping, produced automatically and indexed: versioned policies and exceptions, catalog records, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Sample incidents monthly from first signal through verified corrective action, deliberately including customer-reported events, downgraded incidents and missed timelines.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling end to end and reporting gaps to the sponsor.
- Run a formal mock audit in month 7 using the populations, evidence requests and interviews the auditor will use: a commander, a random engineer, Support, Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
34. Programme risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps and immediate staffing intervention.
- **A major SEV1 mid-rollout:** the programme lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
35. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 30, 31, 33)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, exercise calendar and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and quarterly board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 2 (agent gpt5.6-sol_refine_2, openai/gpt-5.6-sol) ---
Estimated complexity: high
Success metrics: - By day 7, every suspected major incident uses one record, one coordination channel, and a named Incident Commander within 10 minutes.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes.
- By month 3, at least 95% of SEV1 and SEV2 incidents have a named commander within 5 minutes.
- By month 3, at least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and appropriate technical coverage.
- By week 18, all 180 services have an owner, tier, escalation path, and tested coverage model.
- No new mandatory night rotation begins before compensation, training, access, runbooks, and minimum staffing are active.
- Every direct 24x7 technical rotation has at least six qualified responders or an approved, expiring executive exception.
- No responder is routinely primary more often than one week in six or assigned to two simultaneous primary rotations.
- Median impact-to-detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- By day 90, incident replay identifies a current detector for at least 90% of the 31 historical incidents, including its expected detection minute.
- Median time to mitigation falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of applicable SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet the required internal and customer update cadence by month 3.
- Monthly human notification episodes fall from the current 3,400 alert events to no more than 1,500 by day 90 and 500 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without reducing Tier 0 or Tier 1 replay coverage.
- Average after-hours load remains at or below two notification episodes per responder per week; every sustained breach receives a dated remediation plan.
- All six monitoring sources route human pages through the controlled paging platform by week 12, with direct legacy routes retired by week 18.
- All required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 from month 3 onward.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due dates with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Reportability is assessed and recorded within 1 hour for 100% of SEV1 and qualifying SEV2 incidents, including not-reportable decisions.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Approved ledger RTO and RPO, controlled failover, read-only mode, and post-recovery reconciliation are exercised before month 6.
- Quarterly responder surveys reach at least 75% favorable responses for fairness, ownership boundaries, compensation, and sustainability by month 6.
- Monthly contracted availability meets or exceeds 99.95% by month 6 using the contractually authoritative measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- The month-7 mock audit finds no unowned high-risk control gap and at least 95% of sampled incidents contain complete operating evidence.
- The SOC 2 Type II incident-response controls complete external testing without an unresolved material exception.
Steps (28):
1. Charter the program and fund immediate action
Make incident management a **company operating process** within 48 hours. Give one accountable leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Include Engineering, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit in a decision group. The sponsor resolves blocked decisions within 48 hours.
- Approve non-negotiables: one severity model, one human-paging path, one incident record, paid on-call, named service ownership, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, resilience, and incident actions.
- Fund tooling, compensation, training, exercises, and program staff. Compare the cost with the existing $1.3M in annual SLA credits.
- Set the schedule: operating floor by day 7, design during weeks 1–4, pilot during weeks 5–10, rollout during weeks 11–18, control tests in months 4 and 6, and mock audit in month 7.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for the final policy or new tooling. Put a minimum viable process into operation immediately and retain evidence from the first day.
- Publish a one-page interim severity guide, declaration procedure, role card, and communications clock.
- Provide one monitored declaration route through chat, telephone, and the current paging environment.
- Create one channel, bridge, timeline, and incident identifier for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced managers and engineers.
- Require a named commander within 10 minutes. The duty engineering director assumes command if the command page is unclaimed.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Stop uncompensated mandatory after-hours expansion. Pay interim duty under a temporary stipend, retroactive to program launch.
- Hold a 15-minute daily operational review until the permanent process is active.
3. Create the factual, contractual, and control baseline (depends on: 1)
Build one defensible baseline for process design, investment decisions, and SOC 2 testing. Preserve the original data so progress cannot be created by changing definitions.
- Reconstruct all 31 incidents from impact start through detection, declaration, command, mitigation, recovery, communications, credits, and corrective actions.
- Reconstruct the two incidents with unclear command minute by minute.
- Identify the missing signal for every customer-first detection.
- Inventory all six alert sources, 3,400 monthly alert events, duplicates, noisy rules, missing owners, and missing runbooks.
- Record current rotations, unpaid duty, overnight activations, schedule size, and uncovered services.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, 85% noise, $1.3M in credits, and 11 of 64 actions closed.
- Inventory customer-specific availability definitions, notice periods, credit terms, sponsor-bank obligations, and other incident-related contracts.
- Confirm the required SOC 2 Type II observation period and evidence expectations with the auditor during week 1.
4. Publish the on-call fairness contract (depends on: 1)
Treat pager resistance as a legitimate design constraint. Adoption depends on a written agreement that separates command from technical ownership.
- Interview representatives from all 28 teams and from Support, Customer Success, Security, and Operations.
- Separate concerns about unpaid work, sleep loss, unfamiliar systems, noisy alerts, inadequate runbooks, and blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for systems they own or have formally accepted and trained to support, and assisted by a separate Incident Commander.
- State that Platform may temporarily triage unknown ownership but does not inherit another team's service.
- Provide confidential accommodations for health, disability, pregnancy, or caregiving constraints without career penalty.
- Measure baseline trust, fairness, fatigue, and alert confidence. Repeat at days 60 and 120, then quarterly.
5. Build the service catalog and customer-journey map (depends on: 3)
Make a machine-readable catalog the source of truth for routing, impact analysis, status-page components, and control evidence. Every production service must have one accountable owner.
- Record the owning team, manager, business capability, repository, channel, dashboard, runbook, escalation policy, dependencies, regions, and data stores for all 180 services.
- Map initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, reporting, and onboarding to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing systems that can create material contractual impact.
- Classify Tier 2 as deferrable internal or batch systems and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional recovery mode, contractual commitments, and critical vendors for Tier 0 and Tier 1.
- Keep separate but coordinated ownership for the ledger application and PostgreSQL platform.
- Give orphan services an owner or approved decommission date within 30 days. Treat unowned Tier 0 services as release-blocking executive risks.
6. Adopt the severity model and incident modifiers (depends on: 3, 5)
Classify incidents by credible customer, financial, security, regulatory, and contractual harm. Start at the higher plausible severity while scope or integrity remains unknown.
- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; broad loss of a core payment journey; both regions impaired; or processing must be stopped. Immediately page all response roles and the Executive Duty Officer. Open the bridge within 5 minutes, freeze unrelated changes, issue internal notice within 10 minutes, publish applicable customer status within 15 minutes, start legal assessment within 1 hour, and require a postmortem.
- **SEV2 — major:** material payment degradation; settlement deadline at risk; regional impairment with reduced resilience; a critical customer or material cohort unavailable; or an SLA breach is likely. Page command and technical roles immediately. Issue internal notice within 15 minutes, applicable customer status within 30 minutes, and require a postmortem.
- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads. Page only when immediate action can reduce harm.
- **SEV4 — operational event:** no current customer impact and no credible imminent harm. Create a ticket and handle during normal hours.
- Add FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, and VENDOR modifiers. These invoke specialist controls without distorting customer-impact severity.
- Use at least SEV2 posture for unknown impact lasting 15 minutes, credible ledger-integrity risk, or a cross-domain incident without clear ownership.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Permit anyone to declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Publish a decision tree and examples based on the 31 historical incidents.
7. Define roles, authority, and handoffs (depends on: 6)
Separate coordination, communications, recordkeeping, and technical repair. One named person holds command continuously throughout every SEV1 and SEV2.
- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not act as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, executive updates, and coordination with Legal.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- **Subject-Matter Responders:** diagnose and mitigate only systems for which they have ownership, access, training, or a formally accepted support agreement.
- **Executive Duty Officer:** removes organizational obstacles and makes exceptional business decisions without displacing the commander.
- Add Security, Legal, Compliance, Finance, and Vendor Management on modifier-specific triggers.
- Require distinct commander, Communications Lead, scribe, and technical lead for SEV1. Communications and scribe may combine for the first 10 minutes of a bounded SEV2.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, reconciliation, and change evidence for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, including the exact transfer time and unresolved risks.
8. Create sustainable 24x7 coverage (depends on: 4, 5, 7)
Use central command coverage and risk-based technical coverage instead of creating 28 fragile night rotations. Responders do not carry pagers for unfamiliar code.
- Create a 24x7 Incident Command corps of approximately 30 certified people with primary and backup schedules. Two schedules create about 104 weekly assignments a year, or roughly three to four weeks per person annually.
- Create a 24x7 Communications pool of 18–24 trained people from Support, Customer Operations, Engineering Operations, and management.
- Create a similarly sized scribe pool. The backup commander temporarily records the first minutes if a scribe has not joined.
- Maintain a 24x7 Executive Duty Officer schedule and specialist contact paths for Security and Legal.
- Group Tier 0 and Tier 1 systems into roughly 8–12 coherent response domains only where responders share training, access, runbooks, and explicit support acceptance.
- Staff every critical domain with primary and secondary responders and at least six qualified people. Target eight where overnight activation is frequent.
- Give Tier 2 and Tier 3 services business-hours ownership plus tested manager and director escalation.
- Reclassify any lower-tier service capable of causing severe overnight harm rather than hiding the risk behind manager callback.
- Route unknown-owner incidents to Duty Command and Platform temporarily. Record each occurrence as a catalog control defect.
- Prohibit simultaneous primary assignments and two-person 24x7 rotations.
9. Implement compensation and fatigue protections (depends on: 4, 8)
End unpaid on-call before expanding mandatory coverage. Use fixed duty compensation so responders are not rewarded for alert volume.
- Use planning bands of $900–$1,200 per Tier 0/1 primary week and $300–$500 per secondary week.
- Use planning bands of $1,000–$1,300 per Duty Commander week and $400–$700 for Communications or scribe primary duty.
- Pay holiday premiums. Compensate all legally compensable active and waiting time for non-exempt staff, including overtime where required.
- Have HR, Finance, Payroll, and employment counsel approve final bands, tax handling, FLSA classification, New York wage-hour treatment, and schedule constraints within 14 days.
- Provide a protected recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal delivery commitments by about 15% during a primary week.
- Prohibit consecutive primary weeks, on-call during leave, hidden schedule swaps, and primary duty more often than one week in six.
- Allow responders to declare temporary fatigue-related unfitness without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Make active payroll setup, training, access, and readiness hard gates before any new mandatory night rotation starts.
10. Codify the live incident lifecycle (depends on: 6, 7)
Give every incident the same operational flow from first signal through verified recovery. The first objective is limiting customer and financial harm, not proving root cause.
- Use the states Detected, Declared, Triaged, Mitigating, Mitigated, Monitoring, Resolved, and Reviewed.
- Record impact start separately from detection. Use the earliest defensible evidence and revise it transparently when later facts emerge.
- Open with a standard command message: severity, known impact, assigned roles, immediate objective, workstreams, and next update time.
- Freeze unrelated production changes during SEV1 and normally during SEV2. Record every exception.
- Prefer reversible mitigation: rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, partner rerouting, or controlled processing suspension.
- Separate mitigation and diagnosis workstreams when staffing allows.
- Keep decisions in the shared incident record rather than direct messages.
- Require formal command and technical handoffs for shift changes, fatigue, or incidents exceeding four hours.
- For payment incidents, verify backlog handling, duplicate protection, customer state, settlement exposure, and ledger reconciliation before resolution.
- Require a severity-specific stability period and explicit handback to the owning team, Support, and Customer Success.
11. Establish one paging and incident system of record (depends on: 5, 7, 8)
Monitoring tools may remain specialized, but every human page and major-incident record must enter one controlled platform. This provides consistent routing and an audit trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page within two weeks.
- Ingest events from all six monitoring tools before disabling their direct human-paging routes.
- Route pages using the service catalog and deduplicate events belonging to the same symptom.
- Provide one declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, decisions, severity changes, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status publishing, and conference facilities.
- Apply MFA, role-based access, periodic access reviews, and tamper-evident history.
- Keep privileged legal or security material in restricted linked records rather than exposing it in the general timeline.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, the status provider, or an AWS region.
- Retire each legacy paging route only after ownership review, end-to-end testing, and two weeks of verified operation.
12. Enforce the alert-quality contract (depends on: 3, 5)
Treat every human page as a production interface with an owner and a required action. Measure human notification episodes rather than raw monitoring events.
- Require every paging rule to identify the service, owner, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Define an actionable page as one that causes or materially informs a timely intervention or risk decision.
- Define noise as duplicate, non-urgent, unactionable, stale, test-generated, or incorrectly routed notification.
- Page on payment outcomes, error-budget burn, queue age against deadlines, and financial-integrity risk rather than raw CPU, memory, pod, or log thresholds.
- Run new rules in shadow mode for seven days unless a documented emergency exception applies.
- Review repeated no-action pages within two business days.
- Set a page budget of no more than two after-hours notification episodes per responder per week, measured over four weeks.
- Make a sustained budget breach trigger a tuning sprint and block additional non-emergency paging rules.
- Require compensating detection and central approval before suppressing a Tier 0 or Tier 1 rule.
- Never disable an existing critical detector solely because its metadata or runbook is incomplete. Track the gap with a dated remediation owner.
13. Detect payment failures before customers (depends on: 5, 12)
Move detection from infrastructure health to customer journeys and ledger truth. Validate coverage against actual historical failures.
- Define SLIs and internal SLOs for initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, APIs, webhooks, and reporting freshness.
- Set internal objectives with enough headroom to protect the contractual 99.95% availability commitment.
- Run external synthetic transactions through critical journeys at least once per minute from paths independent of the production platform.
- Test each region and expose dependencies that defeat nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed payment value relative to settlement deadlines.
- Add tenant and cohort anomaly detection for high-value customers and critical payment methods.
- Convert credible Support, account-manager, processor, bank, and network reports into incident candidates within five minutes.
- Replay all 31 historical incidents. Record which current detector would fire and at what minute.
- Treat customer-first detection as a mandatory missed-detection review with a tracked action.
14. Implement the detection and escalation ladder (depends on: 7, 8, 11)
Create one time-bound path from the first credible signal to named command and the correct technical owner. Device delivery does not count as acknowledgement.
- Converge automated alerts, engineer observations, support cases, account-manager reports, partner notices, and customer calls on the same declaration path.
- Page the Duty Incident Commander and owning critical-domain primary immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at 5 minutes, manager at 10 minutes, and director at 15 minutes.
- Escalate unclaimed command to the backup commander at 5 minutes. The Executive Duty Officer assumes temporary command at 10 minutes until a certified transfer occurs.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Keep command with the current commander when service ownership remains unclear. Assign a temporary technical lead and record the ownership gap.
- Maintain tested external escalation routes for AWS, database support, processors, sponsor banks, networks, and critical vendors.
- Test the full declaration, acknowledgement, fallback, conference, and status-publishing path weekly.
- Treat missed acknowledgement, failed routing, and unowned incidents as control failures requiring review.
15. Standardize internal communications (depends on: 7, 10, 11)
Give responders one working room and stakeholders one controlled source of truth. Executives must not interrupt the technical command path.
- Maintain one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue the initial internal notice within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within the same initial-notice window.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record every material decision and outbound message in the incident timeline.
- Require the executive team to sign the communication behavior rules.
16. Standardize customer, account, and regulatory communications (depends on: 3, 6, 15)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate observed impact before root cause is known.
- Publish customer status within 15 minutes of a customer-visible SEV1 and within 30 minutes of a customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map public status components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, available workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Have account managers contact affected strategic accounts within 30 minutes for SEV1 and 60 minutes for SEV2 using the approved briefing.
- Offer status subscriptions to all customers. Auto-enroll only where contracts, consent, and applicable communication rules permit.
- Encode customer-specific notice deadlines and channels in the customer record.
- Have Legal and Compliance maintain a counsel-validated matrix covering applicable NYDFS, breach, GLBA or FTC, PCI, money-transmitter, sponsor-bank, network, insurance, and customer obligations.
- Complete and record a reportability assessment within one hour for every SEV1 and every security, privacy, or integrity-related SEV2, including decisions of not reportable.
- Let Legal own regulatory text and submission while the Incident Commander owns operational facts. Record any legally required restriction of public detail and its alternative stakeholder plan.
17. Tie incidents to SLA credits and financial exposure (depends on: 3, 16)
Link the incident record to contractual and financial outcomes. Finance should not discover outages through later credit claims.
- Define the authoritative availability calculation for each contract and customer journey with Legal and Finance.
- Calculate affected customers and minutes from incident scope and journey telemetry.
- Produce a preliminary credit and contractual-exposure estimate within five business days of resolution.
- Record failed payment count, failed or delayed value, settlement exposure, reconciliation breaks, support effort, and engineering effort.
- Establish a documented approval path for proactive credits and claims-based credits.
- Attribute credits and financial harm to recurring failure families.
- Use the quarterly credit analysis to prioritize detection, resilience, and architectural investment.
18. Establish mandatory blameless postmortems (depends on: 6, 7)
Use one learning standard with fixed deadlines. Keep learning separate from disciplinary and misconduct processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact lasting more than two hours, SLA credits, contractual breach, repeated contributing factors, major control failures, and ledger-integrity near misses.
- Produce a factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the owning engineering director accountable for completion. The commander owns response analysis, and the scribe supplies the timeline.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit answers to why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the commander or primary technical responder.
- Describe decisions using the context and information available at the time. Do not name an individual as the root cause.
- Keep HR, misconduct, and personnel matters in separate processes.
- Publish broadly useful findings internally while restricting security, privacy, personnel, and privileged content appropriately.
19. Make corrective actions enforceable risk commitments (depends on: 18)
An action is not complete when its ticket is closed. It is complete when the intended risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within 7 days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Prefer actions that remove hazards or reduce blast radius over vague actions such as retraining or adding monitoring.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of engineering capacity for approved reliability work.
- Escalate overdue high-risk actions to the manager after 7 days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit unrepaired severe conditions to block related releases.
- Verify effectiveness using tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days as complete, re-planned, superseded with evidence, or formally risk-accepted.
20. Set the service-readiness bar and payments playbooks (depends on: 5, 10, 12, 13)
A critical service must be supportable at 3 a.m. before it enters direct overnight coverage. Existing critical detection remains active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and integrity constraints.
- Require responders to demonstrate access and safe execution before independent primary duty.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional impairment, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document split-brain prevention, replay protection, failover, controlled backlog recovery, and post-recovery reconciliation.
- Require executive-approved RTO and RPO for the shared ledger cluster.
- Exercise critical runbooks at least twice a year and after material changes.
- Block new Tier 0 or Tier 1 releases and paging rules when readiness requirements are missing.
- Handle existing gaps using named owners, compensating controls, executive-approved expiry dates, and remediation plans.
- Run a parallel architecture workstream to reduce ledger blast radius, isolate non-critical readers, and strengthen regional independence.
21. Train and certify every response role (depends on: 7, 14, 15, 18, 20)
Command and communication are learned skills. Use paid working time and certify people before independent duty.
- Give all employees a 30-minute module on recognizing impact, declaring incidents, and locating the status page.
- Give responders a half-day module on severity, acknowledgement, escalation, evidence, runbooks, and financial-integrity precautions.
- Train scribes for two hours on timeline quality and fact-versus-hypothesis labeling.
- Train Incident Commanders for two days on delegation, uncertainty, severity, mitigation strategy, fatigue, handoffs, and executive management.
- Require commander candidates to complete a simulated SEV1 and two shadowed incidents or exercises.
- Train Communications Leads for one day on status writing, customer segmentation, legal boundaries, and contractual clocks.
- Require domain responders to demonstrate dashboards, access, rollback, failover, escalation, and relevant playbooks.
- Require two shadow shifts before independent primary duty.
- Renew certification annually through simulation.
- Maintain the training, assessment, and certification register as operational and audit evidence.
- Nominate an incident-management champion in each of the 28 teams.
22. Exercise command, recovery, and tool failure (depends on: 11, 20, 21)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first cross-company command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous operational and security events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors when relevant.
- Conduct at least one unannounced after-hours paging test before the audit and two annually thereafter.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure acknowledgement, command, customer notice, mitigation decision, handoff, and recovery times.
- Create tracked corrective actions for every material exercise finding.
23. Publish the signed policy and control set (depends on: 6, 7, 8, 9, 10, 12, 14, 15, 16, 18, 19)
Convert the design into concise documents that people can use during an incident. The actual operating process must also be the documented and audited process.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, evidence, and exception standards.
- Include one-page cards for severity, roles, authority, escalation, and communication timings inside the incident tool.
- State explicitly that responders support only owned or formally accepted and trained service portfolios.
- Include the compensation structure, fatigue rules, declaration rights, and non-retaliation commitment.
- Obtain approval from the CTO, HR, Legal, Security, Compliance, and Internal Audit.
- Announce the policy at an all-hands and through team briefings.
- Create an exception register with owner, rationale, compensating control, approver, review date, and expiry.
- Version every policy change. Do not rewrite historical records when the process changes.
24. Instrument the scorecard and review forums (depends on: 3, 11, 18, 19)
Measure the process before the pilot so failures become visible immediately. Report medians and 90th percentiles rather than averages alone.
- Measure impact-to-detection, detection-to-declaration, declaration-to-command, acknowledgement, mitigation, recovery, and resolution.
- Split results by severity, service tier, customer journey, region, detection source, and business-hours status.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track notification episodes, actionability, duplicates, after-hours load, missed detection, routing errors, and page-budget breaches.
- Track postmortem timeliness, action age, due-date performance, verified effectiveness, and repeated contributing factors.
- Track journey availability, error-budget burn, failed or delayed value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight work, recovery days, exceptions, sentiment, and responder attrition.
- Hold a weekly Incident Review Board chaired by the Head of Reliability with relevant directors.
- Hold a monthly executive reliability review and a quarterly control and resilience review with Internal Audit.
- Reconcile incident records monthly against support cases, customer complaints, status history, credits, and major operational anomalies to detect under-reporting.
- Use team-level scorecards to direct help and investment. Never penalize an individual for good-faith declaration.
25. Pilot the complete process on the payment path (depends on: 9, 11, 13, 20, 21, 23, 24)
Run a four-to-six-week pilot across the highest-risk journey before expanding. The interim operating floor remains active for the rest of the company.
- Include payment orchestration, ledger application, PostgreSQL platform, API edge, authentication, settlement, reconciliation, Kubernetes platform, and Support intake.
- Include teams with existing on-call experience and teams new to the model.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking together.
- Run legacy and new paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently taking command.
- Correct critical process or tooling defects within 48 hours.
- Exit only after 95% timely command assignment, 95% communications compliance, no unpaid pages, complete required postmortems, tested fallbacks, and at least 50% lower pilot noise.
- Publish the pilot results, defects, and policy changes company-wide.
26. Roll out by risk with readiness gates (depends on: 25)
Expand in controlled waves and finish early enough to accumulate operating evidence before the audit. A calendar date does not override a failed readiness gate.
- Roll out remaining Tier 0 domains first, followed by Tier 1, Tier 2, and Tier 3.
- Use four waves of six to eight teams, each lasting two to three weeks.
- Gate each service on catalog ownership, appropriate coverage, active compensation, trained responders, access, tested escalation, alert quality, runbooks, and a passed tabletop.
- Require at least six responders only for direct 24x7 technical rotations. Use business-hours coverage for lower tiers.
- Give each wave a named coach and director sign-off.
- Reschedule failed gates or use a time-limited executive exception with compensating controls. Do not create silent waivers.
- Disable legacy human-paging routes after verified cutover for each wave.
- Run a quota-based noise reduction sprint in every wave, starting with the highest-volume rules.
- Pair every suppression with a compensating-detection check.
- Publish an internal adoption dashboard by team, service tier, coverage, and control gap.
- Complete critical coverage by approximately week 12 and all 28 teams by week 18.
27. Prove SOC 2 operating effectiveness (depends on: 22, 23, 24, 26)
Generate evidence through normal operation rather than reconstructing it before fieldwork. Test both control design and consistent execution.
- Map controls to the applicable Trust Services Criteria with Compliance and the auditor, including monitoring, incident identification, response, recovery, communications, and availability.
- Retain approved policies, exceptions, service ownership, schedules, compensation activation, access reviews, training, incidents, communications, reportability decisions, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, missed timelines, non-reportable decisions, and exercises in the testing population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the evidence populations and interviews expected from the external auditor.
- Correct deviations through tracked actions with owners and dates. Never edit history to create apparent compliance.
- Verify that evidence retention covers the full auditor-defined observation period.
- Brief commanders, engineers, Support, and Compliance on the actual process without scripting inaccurate answers.
28. Inspect, adapt, and institutionalize ownership (depends on: 26, 27)
Prevent the process from decaying after rollout or the audit. Change it using measured operating evidence rather than opinion.
- Review the policy after 90 days of live operation using severity calibration, page load, missed detection, communication compliance, action closure, fatigue, and survey results.
- Remove steps that create work without reducing risk. Add controls only where incidents, exercises, or evidence show a gap.
- Reassess Tier 0 and Tier 1 classification and domain boundaries every six months.
- Assign permanent owners for policy, catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review compensation bands, rotation burden, accommodations, and staffing annually.
- Report severe incidents, credits, overdue high-risk actions, and ledger concentration risk to the board or risk committee quarterly.
- Maintain the ledger blast-radius program as an executive risk until failover, degraded mode, reconciliation, and regional independence meet approved objectives.
- Evaluate follow-the-sun coverage using one year of actual activation and staffing data.
- Build year-two plans for automated mitigation, safer deployments, graceful degradation, and error-budget release controls.
--- PROPOSAL 3 (agent qwen3.8-max_refine_3, alibaba/qwen3.8-max) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of 'not reportable'.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (29):
1. Charter the program, fund it, and start the audit clock
Convert the CEO email into a company operating process with one accountable owner, a budget, and a dated timeline that starts this week.
- Name the CTO as executive sponsor and appoint a Head of Reliability & Incident Management as the single accountable owner with full-time authority over all 28 teams.
- Stand up a three-person program office: program lead, platform engineer, reliability analyst.
- Form an eight-person steering group spanning Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, and HR. It proposes; the sponsor decides within 48 hours. Never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one human-paging path, one incident record, one postmortem format, paid on-call, mandatory action tracking, and named service ownership.
- Approve budget anchored against the $1.3M in SLA credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs, and 15% reserved engineering capacity.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Publish a one-page charter company-wide on day 2. Incident response is a company process, not a per-team preference.
- Timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day interim command floor (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place this week so the next outage has a named commander and the evidence clock starts immediately.
- Publish a one-page interim severity card and one declaration path: a Slack command, a phone number, and the existing pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs from the 12 teams already on-call. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical action items into complete, re-plan, or formally risk-accept within 30 days, prioritising ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline of incidents, alerts, and money lost (depends on: 1)
Rebuild the facts before locking any design. This is both the design input and the frozen 'before' picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines in one signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the written fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written, signed deal.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- State explicitly that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Map compliance, evidence, and audit requirements from day one (depends on: 1)
Design evidence as a by-product of operations, not a reconstruction before the auditor arrives. The interim process in week 1 is already evidence.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for internal and external communication; CC5 for control activities; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure with verification, training register, drill records, reportability decisions including 'not reportable'.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records. Require role-based access, MFA, and periodic access review.
- Version and approve all policy documents from day one: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard, Alert Quality Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has a named owner. A wrong owner recreates the pager objection. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Assign one accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. **Tier 1**: customer-facing but degradable. **Tier 2**: internal or batch. **Tier 3**: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for every Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform. They fail differently and need different hands.
- Orphan services get an owner within 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Missing ownership or runbooks blocks Tier 0/1 releases.
7. Adopt the severity scale, declaration rights, and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix. Anyone may declare. Nobody is penalised for over-declaring.
- **SEV1 (crisis):** money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; deploy freeze; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads; business-hours comms; postmortem if customer-detected, over two hours, a repeat, or credit-generating.
- **SEV4:** no customer impact. Ticket only. Never pages.
- Payments-specific anchors: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident become at least SEV2.
- Only the Incident Commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
8. Define incident roles, authority, dual-control, and handover discipline (depends on: 7)
Solve 'nobody in charge for an hour' by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Command claimed within 5 minutes and stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Staff 24x7 with a three-layer coverage model, not 28 night rotations (depends on: 6, 8)
Do not create 28 night rotations. That is precisely what engineers are rejecting. Centralise coordination in a trained command corps and keep technical ownership local.
- **Layer A — Incident Command corps:** approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. Paired with a Communications Lead pool of approximately 18 from Support, CS, and engineering management, and a scribe pool used as the training entry point.
- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat a Lisbon or APAC follow-the-sun cell as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
10. Approve paid on-call, New York labor compliance, and fatigue safeguards (depends on: 4, 9)
Unpaid on-call in New York is both a retention problem and a wage-hour exposure. Pay must be live in payroll before any new mandatory night rotation starts. Publish the actual numbers, then ask people to sign up.
- Indicative scheme locked by HR, Finance, and employment counsel within 14 days: approximately $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as approximately 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
- Publish on-call load by team quarterly.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
11. Set the alert-quality standard and a hard page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target: 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, noise below 15%, with no loss of Tier 0/1 detection coverage.
12. Detect payment and ledger failures before customers do (depends on: 6, 11)
The goal is blunt: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
- Run the replay test: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish replay coverage as a leading metric.
13. Consolidate to one paging platform, one incident record, one status page (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
14. Codify the five-minute escalation path and live execution doctrine (depends on: 8, 9, 13)
Write one unskippable path from 'something looks wrong' to 'someone is in charge'. The default action is never waiting. If nobody claims command within 5 minutes, the platform assigns it and announces it.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute.
- Separate diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Closure requires a stability observation window and explicit handback.
15. Set the service-readiness bar and write major-incident playbooks (depends on: 6, 9, 12)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human at night. The shared ledger cluster is the single largest structural risk.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction: partitioning, read replicas, isolation of non-critical readers, with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
16. Run one communications clock for internals, customers, and regulators (depends on: 7, 8, 13)
Replace 'whoever is around' with one timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure. State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is 'not reportable'.
- Obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
17. Tie incidents to SLA credits and true financial cost (depends on: 7, 16)
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
18. Make blameless postmortems mandatory with one format and fixed deadlines (depends on: 7, 8)
Replace 'some incidents, various formats' with one mandatory format, fixed deadlines, and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did?
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and are never the commander of the incident under review.
- Publish a searchable library and a quarterly top-five recurring causes analysis, with restricted versions for security or privileged content.
19. Enforce action ownership, reserved capacity, and tracking (depends on: 13, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity. Closing a ticket without evidence does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
- Target: 90% of high-priority actions closed by due date within two quarters.
20. Train and certify every role before independent duty (depends on: 8, 14, 16, 18)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it. The certification register is an audit artefact.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days plus shadowing): command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
21. Rehearse with tabletops, game days, and unannounced drills (depends on: 13, 15, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-mitigation.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
22. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 14, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 10, 12, 13, 15, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse.
- Publish a one-page result to the whole company.
24. Instrument metrics, dashboards, review cadence, and anti-gaming (depends on: 13, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
25. Roll out in risk-ordered waves with readiness gates (depends on: 23, 24)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, never waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave. Give every team its ranked list of noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Review missed detections monthly with the same seriousness as noise.
- Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
26. Run change management, fairness, and pager culture from day one (depends on: 4, 10)
Run this in parallel from day one. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
27. Produce SOC 2 evidence by operating, test internally, and mock-audit (depends on: 22, 24, 25)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence. Documented exceptions beat a claim of perfection.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5, CC2.2/CC2.3, CC5, and A1.2. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including 'not reportable', postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in months 4 and 6, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
28. Maintain the program risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and pre-commit the response. Review monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
29. Inspect at 90 days, lock year-two ownership, and prevent decay (depends on: 24, 25, 27)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 4 (agent grok4.6_refine_4, xai/grok-4.6) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- By day 7, every suspected major incident uses one record, one channel, and a named commander; Support can declare without engineering confirmation.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of customer-visible SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (26):
1. Charter the program, fund it, and start the audit clock
Turn the CEO email into a chartered company program within 48 hours. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability as full-time owner, with a three-person office: program lead, platform engineer, and analyst.
- Form a small decision group of Engineering, SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit. It proposes. The sponsor decides within 48 hours.
- Lock the non-negotiables now: one severity scale, one human-paging platform, one incident record, one postmortem format, mandatory action tracking, paid on-call, and named service ownership.
- Confirm the SOC 2 Type II observation window with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, exercises, and reserved engineering capacity against the $1.3M in SLA credits.
- Reserve 15% of engineering capacity for detection, runbooks, and incident actions, protected by the sponsor.
- Publish a one-page charter on day 2. Clock: floor by day 7, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for policy, tooling, or the audit. Put a crude but real process in place this week so the next outage already has an owner.
- Publish a one-page interim severity card and one declaration path: chat command, phone number, and current pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline, and one naming convention for every suspected major incident.
- Triage the 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
- Replay one of the two nobody-in-charge incidents as a tabletop within 14 days using this floor process.
3. Rebuild the forensic baseline and incident-replay catalog (depends on: 1)
Rebuild the facts before locking design. This is the design input, the frozen before-picture for the CEO, and the test set for detection work.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact missing signal. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Name the top 50 noisy rules and every rule with no owner or runbook.
- Quantify true cost: credits, failed payment volume, delayed value, reconciliation breaks, and engineering hours lost.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11 of 64 actions closed, on-call in 12 of 28 teams.
- Build a **replay catalog**: for each historical incident, the detector that should now fire, at which minute, and the owner of the gap.
4. Publish the on-call fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders. Do not punish them by making them the company's permanent night watch.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the **fairness contract** in writing: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive exemption path for health, disability, or caregiving.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and journey tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has one owner. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Record owning team, manager, chat channel, escalation policy, dashboards, runbook, dependencies, regions, and data stores for all 180 services.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, refund, report, onboard — to services, data stores, regions, sponsor banks, and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform. They fail differently.
- Orphan services get an owner in 30 days or an approved decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Coverage follows the journey, not the org chart. Small teams that own Tier 0 pieces get headcount, service reassignment, or membership in a domain rotation. Never a two-person 24x7 rota.
6. Lock severity levels, integrity flags, and the incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Classify on actual or credible customer, financial, security, and contractual harm, never on who reported it or how hard the fix looks.
- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 contained: narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages a human.
- Attach an **integrity, security, settlement, or regulatory flag** to any severity. The flag forces dual-control, Legal, and the reportability checkpoint without inventing a fifth level. A one-customer ledger corruption is still a flagged crisis.
- Auto-escalate to at least SEV2: any ledger-cluster event, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident.
- Anyone may declare. Nobody is punished for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- Communications Lead: single voice for the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists. A human validates at SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory text. The commander owns the facts.
- Distinct people for command, comms, and technical lead at SEV1. Comms and scribe may combine only for bounded SEV2.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
- Every role assignment and handover is announced verbally and in writing with the exact time and open risks.
8. Staff 24x7 with command, domains, and a Duty Triage Desk (depends on: 4, 5, 7)
Do not create 28 night rotations. That is what engineers are rejecting. Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified people from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to eight months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate ownership into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- **Duty Triage Desk**: a small paid overnight first-line rotation that owns the first ten minutes of ambiguous, unowned, or low-confidence pages. It verifies, enriches, and applies only the runbook's safe steps, then wakes the owning domain. It never sits on a suspected ledger, payment-halt, or security event. Those page commander and likely domains immediately, in parallel.
- Overnight comms: SEV1 pages a Communications Lead 24x7. Customer-visible SEV2 lets the commander publish the first status from a template; Comms is paged if the incident is still open at 30 minutes or a top-100 account is affected.
- Staffing math: 260 engineers can sustain 10–12 domain rotations, one command corps, and one triage desk. They cannot sustain 28 night rotas. Role exclusivity: nobody is primary on two rotations in the same week. Commanders may also be domain responders in different weeks.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
- Evaluate follow-the-sun as a 12-month option, not a year-one dependency. Seed Layer B from the 12 teams already on-call.
9. Pay on-call, meet New York labor rules, and cap fatigue (depends on: 4, 8)
Unpaid on-call in New York is a retention problem and a wage-hour exposure. Pay must be in payroll before any new mandatory night rotation starts.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, $400–$800 for Comms or triage-desk duty, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- **Hard gate: no mandatory night rotation starts before compensation is live in payroll.** Interim duty is paid retroactively. Budget roughly $0.9M–$1.2M a year, then refine with actual rotation count.
10. Enforce an alert-quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers, then replay the past (depends on: 5, 10)
Stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. A detector is not done until it would have caught the last 12 months.
- Define SLIs and SLOs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner, processor, and sponsor-bank notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a named defect class with a mandatory tracked action.
- Run the **replay test** against the S3 catalog. For each of the 31 incidents, name the detector that would now fire and at which minute. Close gaps the replay exposes before calling detection improved.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One chat command creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy. Unowned pages go to the Duty Triage Desk and Duty Command, and log a catalog defect.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Codify escalation, the five-minute command rule, and vendor incidents (depends on: 7, 8, 12)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting. Payments also fail at processors and banks, which you cannot patch.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Vendor-incident class: processor, sponsor-bank, card-network, or cloud-control-plane failure. Command is still required. Success is time-to-customer-notice, time-to-failover-decision, and queue management, not root cause at the vendor.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Write live execution doctrine, readiness bars, and ledger playbooks (depends on: 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving root cause. A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis. Split diagnosis and mitigation workstreams once staffing allows.
- Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Drain backlogs under control. Complete reconciliation before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date. Existing detection stays on while gaps are repaired.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers, with board-visible milestones and executive-signed RPO/RTO.
15. Run one communications clock for internals, customers, and account managers (depends on: 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for customer-visible SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys, not internal service names. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer record.
- Language rules: state impact and the next update time. **Never speculate** on cause, recovery time, data integrity, or blame.
16. Operationalize regulatory notice, partner clocks, and SLA credits (depends on: 6, 15)
In payments, some incidents start a legal clock at detection. Tie incidents to money so Finance does not learn about outages from invoices.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when not reportable**, with facts, decision-maker, and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families. Track failed payment volume, delayed value, and reconciliation breaks as the true incident cost.
17. Make postmortems mandatory and actions enforceable (depends on: 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth. Eleven of 64 closed is a process nobody enforces.
- Mandatory for every SEV1 and SEV2, any incident a customer detected first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and gap, timeline, response analysis, contributing conditions, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items. Facilitators are trained and are never the commander of the incident under review.
- Every action gets one named owner, priority, due date, expected risk reduction, verification method, and a ticket created automatically. If it is not on the board, it does not exist.
- Classes: containment in 7 days, corrective in 30, strategic in 90. SEV1 recurrence-prevention enters the next sprint ahead of roadmap work. Reserve 15% capacity.
- Overdue ladder: manager at +7 days, director at +14, CTO at +30. Overdue high-risk items need written residual-risk acceptance and can block related releases. Verify effectiveness before closing.
18. Train and certify every response role before independent duty (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Duty Triage Desk, half day: enrichment, safe-step limits, when to wake a domain immediately, when not to delay.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
19. Publish Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17)
Collapse the design into a document people will actually open mid-outage, and make it official. Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
20. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 18, 19)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, Duty Triage Desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
21. Instrument metrics, reviews, and anti-gaming (depends on: 12, 17, 20)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- Response: time from first impact to detect, declare, commander, acknowledge, mitigate, and resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
22. Rehearse with tabletops, game days, and night drills (depends on: 12, 14, 18, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced overnight paging drill to measure real acknowledgement times.
- One combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO, before the audit.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
23. Burn down alert noise and roll out by risk with gates (depends on: 20, 21)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Noise reduction is a quota inside each wave, not a background hope.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy paging paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Give every team its noisiest rules from the baseline. Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed. Pair every suppression with a compensating-detection check.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
- If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
24. Operate the fairness and culture program in parallel (depends on: 4, 9)
Run this from the moment the deal is published. Engineers judge the process on fairness. Executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
25. Prove SOC 2 operating effectiveness before fieldwork (depends on: 19, 21, 22, 23)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (depends on: 21, 23, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, metrics, and the ledger blast-radius workstream, independent of the audit cycle.
- Pre-committed contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 5 (agent deepseek-v4-pro_refine_5, deepseek/deepseek-v4-pro) ---
Estimated complexity: high
Success metrics: - Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (30):
1. Charter the program and start the audit clock
Turn the CEO email into a chartered company program with one accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
- Confirm the SOC 2 Type II observation window with the auditor; map the future response controls to applicable Trust Services Criteria; define evidence set and retention requirements.
4. Listening tour, resistance map and the written on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (depends on: 5, 7)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 4, 8)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (depends on: 5, 10)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (depends on: 7, 13)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 12, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
- Agree with Legal and Finance the availability measurement method per contract and per component; compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide credit posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain; attribute credits to root-cause families for investment decisions.
18. Blameless postmortem standard and Incident Review Board (depends on: 6, 7, 12)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
19. Action ownership, reserved capacity and enforcement (depends on: 12, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
20. Training and certification academy (depends on: 7, 13, 15, 18)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
21. Exercise programme: tabletops, game days and unannounced drills (depends on: 12, 14, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
22. Publish Incident Management Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
24. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
25. Alert noise burn-down campaign (depends on: 10, 12, 23)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
26. Wave rollout to all 28 teams with readiness gates (depends on: 23, 24, 25)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
27. Change management, fairness and pager culture (depends on: 4, 9, 23)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
28. SOC 2 evidence by design, internal testing and mock audit (depends on: 22, 24, 26)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
29. Risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
30. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 24, 26, 28)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
You do not know which of them was voted for, and you must not guess it: judge them on their merits. Your answer has these parts:
- "summary": what the final round offers, as a Markdown list: one item per final proposal naming it (P1, P2…) and saying in one sentence what distinguishes it, then one item on what they all share. No paragraph.
- "assessments": one object per final proposal, with "proposal" (its number), "fitness" ("strong", "adequate" or "weak" for the task as stated), "strengths" and "weaknesses" (lists of concrete points: steps, order, metrics, realism, risk handling).
- "ranking": the numbers of the final proposals from best to worst.
- "ranking_reasons": why that order, naming what separates each one from the next.
- "versus_initial": one object per initial proposal, comparing YOUR FIRST CHOICE with it: "proposal" (its number in round 0), "verdict" ("better" if your first choice is better than that initial proposal, "worse" or "similar"), "why" and "how" (in what concrete ways).
- "improved_over_initial": true if your first choice is better than EVERY initial proposal.
- "improvement_summary": what the deliberation added or lost with respect to the initial proposals, overall, as a Markdown list of points.
[PROCESS]
[SYSTEM]
You are an expert reviewer of multi-agent planning processes.
Several LLM agents drafted plans for a task, refined them over a number of rounds while seeing each other's proposals, and finally voted for the best one.
Be exhaustive but precise: name concrete steps, ideas and metrics, never generalities. Judge plans by their fitness for the task as stated, their realism, their completeness, the soundness of their order and dependencies, how measurable their success is and how they handle things going wrong.
You are an impartial evaluator, not a chronicler: assess the proposals and the process on their merits, never rationalise what happened or assume that the outcome was right.
After your analysis, answer in the requested structure.
Every text field you write will be read by a busy person who skims. Make it easy to skim: short sentences and short paragraphs; when you name several things, prefer a list to a paragraph, with sub-items when an item has parts, but keep a single fact as a sentence; lead with the point and then the evidence; name proposals and steps by number (P2, step 4); no preamble, no repetition of the question, no closing summary; bold at most one key phrase per item or paragraph. Text fields accept Markdown: a blank line between paragraphs, "- " for lists, **bold**.
[HUMAN]
Task given to the agents: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
THE INITIAL PROPOSALS (round 0):
--- PROPOSAL 1 (agent opus5_initial_1, anthropic/claude-opus-5) ---
Estimated complexity: high
Success metrics: - Median time to detect falls from 22 minutes to under 5 minutes within 9 months.
- Customer-first detection falls from 40% of incidents to under 10% within 6 months and under 5% within 12.
- Median time to mitigate falls from 3 h 10 min to under 60 minutes within 12 months.
- An Incident Commander is assigned and announced within 5 minutes in 95% of SEV1/SEV2 incidents; zero incidents with unclear ownership beyond 10 minutes.
- Status page updated within 15 minutes of SEV1 declaration and 30 minutes of SEV2 in 95% of cases.
- Monthly alert volume falls from 3,400 to under 500 pages, with actionability above 75%; out-of-hours pages under 2 per person per week.
- All six legacy alerting tools consolidated into one paging platform, legacy paging paths disabled, by week 16.
- 100% of SEV1/SEV2 incidents have a published blameless postmortem within 10 business days, in the single approved format.
- Postmortem action closure rises from 17% (11/64) to 90% of P0/P1 actions closed by their due date.
- SLA credits fall from $1.3M to under $400k in the first 12 months.
- Customer-impacting incidents fall from 31 to under 15 per year; repeat incidents from a known cause under 10%.
- All 180 services have a named owning team and criticality tier; 100% of Tier 0/1 services meet the on-call readiness bar.
- All 28 teams onboarded by week 24, with 24x7 rotations of 6+ certified responders for every Tier 0/1 team.
- 30+ certified Incident Commanders and 20+ certified Communications Leads, giving 24x7 primary and secondary command cover.
- Paid on-call policy approved by HR, Legal and Finance and in payroll before any mandatory night rotation starts.
- Internal dry-run audit at month 6 passes a 15-incident evidence walkthrough; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call sentiment score improves quarter over quarter; no increase in attrition among engineers on rotation.
- Quarterly game day and twice-yearly unannounced paging drill executed on schedule, each producing tracked action items.
Steps (29):
1. Program charter, executive mandate and funding
Convert the CEO's frustration into a named program with one accountable owner, a budget and a deadline that is earlier than the audit.
The charter must state that incident management is a **company-level operating process**, not a per-team choice. Without that, the 28 teams will opt out.
- Appoint a single Incident Management Program Lead (Head of Reliability/SRE) with direct exec sponsorship from CTO and CEO.
- Form a steering group: CTO, VP Eng, Head of Support/CS, CISO/Compliance, Legal, Finance (SLA credits), HR (on-call pay).
- Set the non-negotiables: one severity scale, one paging tool, one postmortem format, mandatory action tracking, paid on-call.
- Fix the timeline: design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, audit-ready by week 28 (four weeks of buffer before the audit).
- Approve budget lines: tooling (~$150–250k/yr), on-call compensation (~$600k–1M/yr), 2–3 dedicated program FTEs. Anchor it against $1.3M of credits plus incident cost.
2. Forensic baseline of the 31 incidents and the alert estate (depends on: 1)
Before designing anything, rebuild the facts. Re-open all 31 incidents and profile the 3,400 monthly alerts so every later design decision is evidence-based.
This also creates the "before" picture the exec and the auditor will compare against.
- Re-code each incident: trigger, service, detection source (customer vs monitor), timestamps for detect/acknowledge/declare/mitigate/resolve, who led, credits paid, root cause family.
- Quantify the 40% customer-first detections: which signal was missing in each case.
- Classify the two "nobody in charge" incidents minute by minute; use them as the burning-platform story.
- Audit the six alerting tools: volume per tool, per team, per alert rule; identify the top 50 rules that produce most of the 85% noise; find rules with no owner and no runbook.
- Baseline the numbers formally: MTTD 22 min, MTTM 3h10, 31 incidents, $1.3M credits, 11/64 actions closed. Freeze them as the reference line.
3. Stakeholder listening tour and resistance map (depends on: 1)
Engineer pushback against "carrying a pager for other teams' code" is the main delivery risk. Treat it as a design input, not an attitude problem.
Run structured interviews across all 28 teams, plus Support, CS and Sales, in two weeks.
- Test the real objection: is it unpaid work, night sleep, unfamiliar code, poor runbooks, or fear of blame? Each has a different fix.
- Collect current informal practices — the 12 teams already on-call are the pilot candidates and the source of veterans.
- Document the promise that answers the objection: **you are only paged for services your team owns**, plus a trained commander who runs the incident and pulls in others.
- Map influencers and blockers by team; recruit 10–15 credible engineers as a design working group so the process is co-authored, not imposed.
- Survey baseline sentiment (trust in alerts, willingness to be on-call, burnout) to re-measure at 6 and 12 months.
4. Service ownership catalog and criticality tiering (depends on: 2)
You cannot page the right person across 180 services until each service has a named owning team. This is the foundation of both on-call fairness and severity mapping.
Build a machine-readable catalog (Backstage or equivalent) that is the single source of truth for routing.
- One owning team per service, a named engineering manager, a Slack channel, a paging escalation policy, a dependency list.
- Tier services by business impact: Tier 0 (money movement, ledger, auth, shared PostgreSQL cluster), Tier 1 (customer-facing but degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map each Tier 0/1 service to the customer-visible capability it supports (payment initiation, settlement, reporting, onboarding).
- Flag orphan services and cross-team shared components; force an ownership decision for each within 30 days, or schedule decommissioning.
- Publish coverage gaps to the steering group: any Tier 0 service without an owner is an executive escalation.
5. Severity scale and declaration criteria (depends on: 2, 4)
Define a five-level scale with objective, payments-specific triggers so declaration is a lookup, not a debate. Anyone may declare; only the Incident Commander may downgrade.
Each level triggers a fixed bundle of response, comms and postmortem obligations.
- **SEV1**: money movement stopped or incorrect, ledger integrity in doubt, data breach, full region loss, >10% of customers impacted. Triggers: immediate 24x7 page of IC + comms + exec, bridge within 5 min, status page within 15 min, mandatory postmortem, regulator assessment.
- **SEV2**: severe degradation, settlement at risk of missing a window, single large/strategic customer fully down, SLA breach likely. Triggers: IC paged, status page within 30 min, mandatory postmortem.
- **SEV3**: partial or workaround-available degradation, no credit exposure. Team-led, business-hours comms, postmortem optional but encouraged.
- **SEV4/5**: minor or internal-only; ticket-tracked, no paging.
- Add auto-escalation rules: any SEV3 open >2 h, or any incident touching the shared ledger cluster, becomes SEV2 automatically. Include a severity decision tree and 12 worked examples drawn from the 31 real incidents.
6. Incident roles, decision authority and handover rules (depends on: 5)
Solve the "nobody in charge for an hour" failure by making command explicit, transferable and logged.
Define five roles with written responsibilities, entry criteria and explicit authority.
- **Incident Commander**: owns the incident, not the fix. Authority to declare severity, pull any engineer, approve customer-impacting mitigations, invoke failover and authorise spend. The IC never types in the terminal.
- **Communications Lead**: owns status page, internal updates, account-manager briefings and the exec summary. Single voice to customers.
- **Scribe**: maintains the timeline, decisions and open questions; feeds the postmortem and the audit evidence trail.
- **Subject-Matter Responders**: engineers from owning teams; they investigate and remediate, and report to the IC.
- **Executive Liaison** (SEV1 only): shields the IC from exec questions and owns regulator/board escalation.
- Rules: the IC role is assumed within 5 minutes of declaration, stated explicitly in the channel ("I am IC"), and any handover is announced and logged. Roles may be combined below SEV2; never at SEV1.
7. 24x7 incident command coverage model (depends on: 6, 4)
Command is staffed by a small, trained, cross-team pool — not by 28 teams individually. This is what makes 24x7 realistic in one New York time zone.
Design a central rotation that scales with a growing certified pool.
- Create a **Duty Incident Commander** rotation of 25–35 certified volunteers (target ~1 per team, plus managers and senior engineers), giving each person roughly one week per 6–8 months.
- Pair with a Duty Comms Lead rotation (Support/CS leads plus engineering managers, ~15–20 people) and a Scribe pool (rotating, lowest barrier, used as the training entry point).
- Coverage: primary + secondary IC at all times; hard 5-minute acknowledgement SLA with automatic failover to secondary, then to the on-call engineering director.
- Night coverage options to evaluate in writing: US-only rotation with paid night stipend now; a Lisbon/Dublin or APAC follow-the-sun cell as a 12-month option; a 24x7 NOC-style triage desk for first-line detection.
- Eligibility: certification required (S21); commanders are volunteers with manager approval and can step out with 30 days' notice.
8. Team on-call structure, rotations and routing rules (depends on: 4, 6)
Rebuild team on-call around the principle that answers the pushback: **you are only paged for code your team owns**.
Apply a tiered obligation so 28 teams are not treated identically.
- Tier 0/1 owning teams (expected 14–18 teams): 24x7 primary + secondary, minimum 6 people per rotation, one-week shifts, handover Wednesday mornings.
- Tier 2/3 teams: business-hours on-call with a best-effort out-of-hours escalation path, no night paging.
- Platform/Infrastructure and Database teams: 24x7, since they own the shared PostgreSQL ledger cluster and the Kubernetes/regional layer.
- Rotations under 6 people are merged across teams or backfilled by hiring; no rotation of fewer than 4 is approved.
- Routing: every page resolves through the service catalog to the owning team's escalation policy; cross-team pages are made by the IC, never by an alert.
- Guardrails: maximum one week in four, no on-call in the week after a SEV1 you led, protected recovery time after any night page, and a per-person page budget (see S10).
9. On-call compensation, labour compliance and fairness policy (depends on: 8, 3)
Unpaid on-call is both a retention risk and a legal exposure in New York. Paying for it is the fastest way to convert resistance into participation.
Design the scheme with HR, Legal, Finance and Payroll, and publish it before asking anyone to sign up.
- Base stipend per week on rotation, differentiated by tier: e.g. $800–1,200 for 24x7 Tier 0/1, $300–500 for business-hours rotations, with premiums for holidays and weekends.
- Per-incident payment for out-of-hours activation (e.g. $150 per night page plus hourly beyond one hour) and guaranteed time-off-in-lieu after night work.
- Separate Duty IC stipend, since command is a distinct and heavier burden.
- Verify FLSA exempt/non-exempt treatment, NY State wage rules and overtime exposure for non-exempt staff; document the legal review.
- Budget and model the annual cost; get board/CFO approval as a line item, benchmarked against $1.3M of credits.
- Add non-cash elements: on-call time counted as delivery load (teams reduce sprint commitment by ~15%), incident leadership recognised in promotion criteria, and a public quarterly report of on-call load per team.
10. Alert quality standard and page budget (depends on: 2, 4)
3,400 alerts a month at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a contractual condition of paging someone.
Publish a standard, then enforce it mechanically.
- Every paging alert must have: a named owning team, a documented customer impact, a runbook link, a tested threshold, and a severity mapping. Alerts failing this are demoted to ticket or deleted.
- Page only on symptoms that affect customers (SLO burn rate, error budget, queue depth against settlement deadlines); cause-based CPU/memory alerts become dashboards or tickets.
- Set a **page budget**: maximum 2 out-of-hours pages per person per week. Breach triggers a mandatory alert-tuning sprint for the owning team and blocks new alert creation.
- Auto-quarantine: any alert that fires more than 5 times a month without action, or has >70% no-action acknowledgements, is silenced automatically and returned to its owner.
- Monthly alert review per team: kill, tune, or keep, with the numbers on screen.
- Target: 3,400 → under 500 pages a month, with actionability above 75% within six months.
11. Detection uplift: SLOs, synthetic journeys and ledger assurance (depends on: 10, 4)
The goal is to stop customers telling you first. Detection must be driven by customer-visible outcomes, not host metrics.
Instrument the money path end to end and alert on it.
- Define SLOs for each Tier 0/1 customer capability: payment initiation success rate, authorisation latency, settlement file timeliness, API availability and reporting freshness. Tie them to the 99.95% contractual SLA with a stricter internal target.
- Deploy synthetic transactions from outside the platform, in both regions, every 60 seconds, covering the full payment lifecycle including a small real-value canary flow where feasible.
- Add ledger assurance checks: continuous double-entry balance reconciliation, replication lag and failover-readiness alarms on the shared PostgreSQL cluster, and settlement-window countdown alerts.
- Build per-customer anomaly detection for the top 100 accounts (volume drop, error spike) so a single-tenant outage is detected before the account manager calls.
- Create an "inbound signal" bridge: any support ticket or account-manager report matching impact keywords auto-creates a triage incident within 5 minutes.
- Track every incident's detection source; make "customer detected first" a reviewed defect with its own follow-up action.
12. Tool consolidation and incident platform implementation (depends on: 5, 7, 8, 10)
Collapse six alerting tools into one paging and incident platform so there is a single queue, a single timeline and a single audit record.
Run a short, time-boxed selection and migrate within the pilot window.
- Select an integrated stack: paging/on-call scheduling plus an incident management layer (e.g. PagerDuty + incident.io/FireHydrant, or a single vendor) and a hosted status page.
- Implement one-command declaration in Slack (`/incident declare`) that creates the channel and bridge, pages the Duty IC, sets severity, opens the timeline and starts the clock.
- Migrate all monitoring sources to route into the one platform; decommission direct paging from the legacy six and block new integrations that bypass it.
- Automate the evidence trail: timestamps, role assignments, severity changes, comms sent, and postmortem linkage exported for SOC 2.
- Integrate with the service catalog for routing, Jira for actions, Salesforce/CS tooling for affected-customer lists, and Zoom/Slack Huddle for the bridge.
- Hard requirement: the platform must work when AWS in one region is down — verify out-of-band paging (SMS/phone) and a printed/offline fallback runbook.
13. Detection-to-escalation path and the five-minute command rule (depends on: 6, 7, 12)
Write the single path from "something looks wrong" to "someone is in charge", and make it impossible to skip.
The design target is detection to commander in under five minutes, any hour.
- Entry points: automated alert, engineer observation, support ticket, account manager, partner bank, customer-facing SEV hotline. All converge on the same declaration command.
- Anyone in the company may declare up to SEV2; nobody is punished for over-declaring. Publish that rule in writing and repeat it.
- Auto-page ladder: Duty IC (5 min) → secondary IC (5 min) → on-call Director (10 min) → CTO. Same ladder for the owning team's responder.
- Cross-team pull: the IC can page any team's on-call directly, with a 10-minute acknowledgement obligation. This is the reciprocal commitment that makes single-team ownership viable.
- Explicit takeover protocol: if no one claims IC within 5 minutes, the platform assigns it and announces it; the assignee cannot decline, only hand over.
- Define standing severity triggers for immediate regional failover, ledger read-only mode and partner-bank notification, with pre-authorised decision rights so the IC does not wait for an executive.
14. Internal communications protocol (depends on: 6, 12)
Standardise the internal channel so responders, executives and support see the same picture without interrupting the IC.
Separate the working channel from the audience channel.
- One incident channel per incident (auto-created), one bridge, and a read-only broadcast channel for executives, Support and Sales.
- Update cadence by severity: SEV1 every 30 minutes even if nothing has changed; SEV2 every 60 minutes; SEV3 at state change.
- Fixed update template: what is happening, customer impact in plain language, what we are doing, ETA or next update time, current IC and Comms Lead.
- Exec briefing rule: executives ask questions only to the Executive Liaison; the IC is not interrupted. Publish this as a behavioural expectation signed by the exec team.
- Support/CS enablement: a live affected-customer list and a holding statement within 15 minutes of SEV1/SEV2 so the front line is never guessing.
- Handover protocol for incidents beyond 4 hours: formal IC handover checklist, fatigue rule, and staffing of a second shift.
15. Customer communications and status page policy (depends on: 5, 14)
Customers currently learn of outages from their own monitoring and hear from whoever happens to be around. Replace that with a timed, owned, pre-approved process.
The Comms Lead is the single author; templates remove the need to write under pressure.
- Timing commitments: status page posted within 15 minutes of SEV1 declaration and 30 minutes for SEV2; updates every 30/60 minutes; resolution notice within 30 minutes of mitigation; customer-facing summary within 5 business days for SEV1.
- Pre-approve 12–15 templates with Legal and Comms (degradation, delay in settlement, API errors, security event, third-party failure) so nothing needs legal review mid-incident.
- Subscription-based status page with per-component granularity mapped to the customer capabilities from S11, plus an email/webhook/RSS feed.
- Tiered outreach: top 100 accounts get a direct named call or email from their account manager within 30 minutes of SEV1, with a briefing pack from the Comms Lead; long tail gets the status page and a proactive email.
- Rules of language: state impact and next update time, never speculate on cause, never assign blame to a vendor before facts are confirmed.
- Run a quarterly customer-perception check with the top accounts on whether comms were timely and useful.
16. Regulatory, partner and legal notification playbook (depends on: 5, 15)
In payments, some incidents are reportable and the clock starts at detection. Build the assessment into the process so it is never an afterthought.
Work with Legal, Compliance and the CISO to produce a decision tree and contact matrix.
- Map obligations: NYDFS Part 500 (72-hour cybersecurity event notification), state breach laws, GLBA/FTC Safeguards, PCI DSS if card data is in scope, sponsor-bank and card-network contractual notice windows, and any FinCEN/OFAC implications.
- Add a mandatory regulatory-assessment checkpoint to every SEV1 and every security-related SEV2, owned by the Executive Liaison, completed within 2 hours of declaration and recorded even when the answer is "not reportable".
- Build the contact matrix: regulators, sponsor banks, card networks, cyber insurer, outside counsel, with 24x7 numbers and named backups.
- Pre-draft notification letters and hold them under legal privilege review.
- Check customer contracts for bespoke notification SLAs (often 1–4 hours for enterprise accounts) and encode them in the customer tiering.
- Test the playbook once per quarter as part of the simulation programme.
17. SLA credit and financial impact workflow (depends on: 5, 15)
Link incidents to money so severity, credits and prioritisation stay consistent — and so Finance stops being surprised.
Make credit calculation an automated output of the incident record, not a negotiation.
- Define the availability measurement method per contract, per component, and agree it with Legal and Finance.
- Auto-compute affected minutes per customer from the incident record and the per-capability telemetry; generate a proposed credit schedule within 5 business days of resolution.
- Decide the posture: proactive credits for the top tier (reputation upside) versus claims-based for the rest; document the approval chain.
- Track credits per incident and per root-cause family; feed a quarterly report showing which reliability investments would have prevented which credits.
- Set a target: reduce credits from $1.3M to under $400k in year one, and use that delta as the ongoing business case.
18. Postmortem standard and blameless review forum (depends on: 5, 6)
Replace "some incidents, various formats" with a mandatory, single-format, blameless process with fixed deadlines.
The discipline is in the deadlines and the forum, not in the template.
- Mandatory for: every SEV1 and SEV2, every incident where a customer detected it first, every incident over 2 hours, every repeat of a known cause, and every near-miss involving the ledger. Optional but templated for SEV3.
- Fixed timeline: draft within 3 business days, peer review within 5, published company-wide within 10. The IC owns delivery; the owning team's manager is accountable.
- One template: timeline, customer and financial impact, detection analysis (why not sooner), response analysis (why mitigation took as long as it did), contributing factors, what went well, action items with owner and due date.
- Blameless rules in writing: describe systems and decisions in the context available at the time; no individual named as a cause; HR and management commit that postmortems are never used in performance reviews.
- Weekly 60-minute Incident Review Board: reviews all postmortems from the prior week, challenges quality, ratifies severity, and approves or rejects action items. Attendance by engineering directors is mandatory.
- Publish a searchable postmortem library and a quarterly "top five recurring causes" analysis.
19. Action item ownership and tracking system (depends on: 18, 12)
11 of 64 actions closed is the single clearest symptom of a process nobody enforces. Give actions the same status as customer commitments.
Track them where engineering work already lives, with visible escalation.
- Every action gets: a named individual owner (not a team), a priority class, a due date and a Jira ticket auto-created from the postmortem.
- Priority classes with hard SLAs: P0 prevents recurrence of a SEV1, due in 30 days; P1 in 60 days; P2 in 90 days. P0s are committed into the next sprint before any roadmap work.
- Capacity rule: teams reserve a standing 15–20% of sprint capacity for reliability and incident actions. Without reserved capacity, the actions will not land.
- Escalation ladder for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue P0s block related feature releases.
- Monthly reporting of closure rate by team in the engineering leadership review; include it in manager performance objectives.
- Target: 90% of P0/P1 actions closed on time within two quarters.
20. Runbooks, major-incident playbooks and the on-call readiness bar (depends on: 4, 8)
Nobody can respond well to unfamiliar systems at 3 a.m. without runbooks — and poor runbooks are a real part of the pager resistance.
Define a minimum readiness bar that a service must meet before it is allowed to page anyone.
- Readiness checklist per Tier 0/1 service: current architecture diagram, dependency map, dashboard link, alert-to-runbook mapping, rollback procedure, feature-flag kill switches, escalation contacts, and a data-loss/latency impact statement.
- Write major-incident playbooks for the top failure modes derived from S2: shared PostgreSQL ledger failure or corruption, cross-region failover, Kubernetes control-plane loss, partner-bank/third-party outage, settlement-window breach, and suspected security compromise.
- Prioritise the shared ledger: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and a clearly stated data-loss tolerance (RPO/RTO) signed off by the exec.
- Runbooks must be tested at least twice a year in a drill; untested runbooks are marked stale in the catalog.
- Enforcement: a service without readiness sign-off cannot create paging alerts, and the gap is reported to its director.
21. Training, certification and the commander academy (depends on: 6, 13, 14, 18)
Command is a skill, not a title. Build a certification path so 24x7 coverage is staffed by people who have practised.
Use a tiered curriculum with real assessment.
- **Scribe** (2 hours): timeline discipline and tooling. The entry point for everyone.
- **Responder** (half a day): severity scale, declaration, escalation, runbook use, comms hygiene. Mandatory for every engineer joining an on-call rotation.
- **Incident Commander** (two days plus shadowing): command presence, delegation, decision-making under uncertainty, severity calls, handover, exec management. Requires two shadowed incidents and one simulated SEV1 before certification.
- **Communications Lead** (one day): status-page writing, customer tiering, legal boundaries, regulator triggers.
- Certification is valid 12 months and renewed via a simulation; the register of certified people is an audit artefact.
- Add on-call onboarding per team: a new joiner shadows two shifts before holding primary, and never holds primary in their first 90 days.
22. Simulation programme: game days, drills and wheel of misfortune (depends on: 21, 12, 20)
The process must be rehearsed before it meets a real SEV1. Simulations also build the commander pool and expose runbook gaps cheaply.
Run a standing calendar rather than one-off exercises.
- Monthly 60-minute tabletop ("wheel of misfortune") per engineering group, using a real past incident from the 31.
- Quarterly full-scale game day in production or a production-like environment: regional failover, ledger replica promotion, dependency failure, with the whole role structure activated and timed.
- Twice-yearly unannounced paging drill to measure real acknowledgement times at night.
- One security-incident and one regulatory-notification exercise per year, with Legal and the CISO in the room.
- Every exercise produces a lightweight postmortem and action items in the same system as real incidents.
- Measure and publish drill metrics: time to IC, time to first status update, time to correct mitigation decision.
23. Pilot with wave 0 teams (depends on: 22, 9, 11)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it.
Run a six-week pilot with tight measurement and a public verdict.
- Select 5–6 teams: core payments, ledger/database, platform/Kubernetes, API gateway, plus two of the 12 teams already on-call.
- Activate the full stack for them: new severity scale, Duty IC rotation, single paging tool, alert budget, status-page policy, mandatory postmortems, paid on-call.
- Hold a weekly pilot retro; expect and document 20–40 process defects, and fix them in the standard before rollout.
- Validate the hard questions: does the 5-minute IC rule hold at 3 a.m.? Do cross-team pulls get answered? Is the severity tree unambiguous?
- Exit criteria: MTTD under 10 minutes for pilot services, IC assigned within 5 minutes in 95% of incidents, page volume down 50%, all postmortems on time, positive on-call sentiment.
- Publish a one-page pilot result to the whole company — this is the main adoption argument for the remaining teams.
24. Metrics, dashboards and the review cadence (depends on: 12, 18, 23)
Instrument the process itself, so improvement is visible and the audit has evidence of monitoring and review.
Define a small set of metrics with owners and a fixed meeting rhythm.
- Response metrics: MTTD, time to declare, time to IC assigned, MTTA, MTTM, MTTR, incidents per month by severity, % detected by customers first.
- Quality metrics: page volume per person per week, alert actionability rate, page budget breaches, postmortem on-time rate, action closure rate and ageing.
- Business metrics: SLA credits paid, availability against 99.95% per capability, error-budget consumption, repeat-incident rate.
- People metrics: on-call load distribution across teams, out-of-hours pages per person, on-call sentiment and attrition among on-call staff.
- Cadence: weekly Incident Review Board (postmortems and actions), monthly Reliability Review (metrics per team, alert hygiene, on-call load), quarterly Executive/Board review (credits, trends, investment asks), annual policy review.
- Every metric gets a target and a named owner; dashboards are self-serve and public inside the company.
25. Wave rollout across all 28 teams with readiness gates (depends on: 23, 24)
Roll out in four waves of six to eight teams, every three weeks, ordered by criticality. Each wave passes an explicit gate rather than a deadline.
Gates keep quality high and make the standard credible.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with 6+ certified responders, one IC candidate nominated, one drill passed.
- Assign each wave a named coach from the program team for three weeks of hands-on support.
- Gate criteria are checked and signed by the director; teams that fail are re-scheduled, not waived.
- Freeze legacy tooling per wave: after onboarding, the old alerting paths are disabled, not left as a fallback.
- Publish a live adoption scoreboard by team so progress is social, not administrative.
26. SOC 2 control mapping, evidence automation and internal dry-run audit (depends on: 18, 19, 25)
Design the process so audit evidence is a by-product of doing the work, then test it before the auditors do.
Engage the auditor early to confirm the interpretation of controls.
- Map the process to the Trust Services Criteria: CC7.3 and CC7.4 (incident identification, response, recovery), CC7.2 (monitoring), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), plus availability criteria A1.2.
- Produce and approve formal policy documents: Incident Response Policy, Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard — versioned, signed, annually reviewed.
- Automate evidence: incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, postmortem library, action-item closure reports, training and certification register, drill records.
- Confirm the observation window with the auditor and ensure the process is operating for a minimum of three months before fieldwork.
- Run an internal dry-run audit at month six: sample 15 incidents and walk the full evidence chain; fix gaps with 8 weeks to spare.
- Keep a remediation log for any incident where the process was not followed, with the corrective action — auditors respond better to documented exceptions than to a claim of perfection.
27. Change management, incentives and communications campaign (depends on: 3, 9, 23)
Run this in parallel from day one. The process will be judged by engineers on fairness, and by executives on visible results.
Communicate the deal explicitly and repeatedly.
- The deal in one sentence: **you are paid for on-call, you are paged only for what you own, a trained commander runs the incident, and your postmortem actions get real sprint capacity**.
- Launch communications: CTO all-hands, per-team roadshows, a one-page process card for laptops, an internal wiki hub and a Slack support channel with a 4-hour answer SLA.
- Recognition: incident-response contribution in promotion criteria and performance frameworks, quarterly awards for best postmortem and biggest alert-noise reduction, public thanks after every SEV1.
- Manager accountability: adoption, alert hygiene, action closure and on-call load in each engineering manager's quarterly objectives.
- Handle the exceptions: a written path for engineers who cannot do nights (caring responsibilities, health), covered by stipended volunteers elsewhere.
- Track sentiment quarterly and publish the results, including bad news, to keep credibility.
28. Program risk register and contingency planning (depends on: 1)
Name the ways this program fails and pre-commit the response. Review it monthly in the steering group.
The main risks are predictable.
- **Volunteer shortfall for the IC pool**: contingency is to make command a rostered duty for engineering managers and senior engineers until the pool reaches 25.
- **Compensation not approved in time**: fall back to time-off-in-lieu plus a phased stipend, but do not launch mandatory night on-call without some compensation.
- **Tool migration slipping**: keep the single-queue requirement and cut scope on the incident-management layer, not on paging consolidation.
- **Alert pruning causing a missed incident**: prune from paging to ticket first, observe for 30 days, then delete; keep a recovery path.
- **Burnout or attrition among the 12 experienced on-call teams**: monitor load weekly and cap individual page counts.
- **A major SEV1 mid-rollout**: pre-agree that the program lead becomes a full-time responder and the wave schedule slips by one wave, with the steering group informed the same day.
29. Continuous improvement, maturity roadmap and post-audit sustainability (depends on: 24, 25, 26)
Protect against the classic failure: the process decays once the audit is passed. Build the second-year plan before the first year ends.
Set a maturity model and a forward roadmap with owners.
- Quarterly process retrospective with the IC pool: what in the process itself slowed us down, what needs simplifying, is the severity scale calibrated?
- Re-baseline targets every six months; a process that hits all targets is under-ambitious.
- Year-two roadmap candidates: follow-the-sun coverage cell, automated mitigation and self-healing for the top three recurring causes, error-budget policy that gates releases, per-customer real-time impact reporting, and blast-radius reduction for the shared ledger cluster (the largest single structural risk).
- Move from lagging metrics (MTTR) to leading ones (error-budget burn, near-miss rate, drill performance).
- Make the annual policy review, certification renewal and drill calendar permanent calendar items owned by the Head of Reliability, independent of the audit cycle.
- Report to the board quarterly on availability, credits and incident trends so the process keeps executive attention after SOC 2 is signed.
--- PROPOSAL 2 (agent gpt5.6-sol_initial_2, openai/gpt-5.6-sol) ---
Estimated complexity: high
Success metrics: - Within 7 days, every suspected SEV0–SEV2 has one incident record, one channel, and a named incident commander.
- After day 30, no SEV0 or SEV1 remains without a named incident commander for more than 10 minutes.
- By day 30, 100% of Tier 0 and Tier 1 services have a named owner, primary escalation, secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and subject-matter coverage.
- By day 120, 100% of production services have sustainable ownership and tested escalation paths.
- At least 95% of SEV0 and SEV1 pages are acknowledged within 5 minutes by month 3.
- At least 95% of SEV2 pages are acknowledged within 10 minutes by month 3.
- Median customer-impact detection time falls from 22 minutes to 10 minutes by day 90 and 5 minutes by month 6.
- The proportion of incidents first detected by customers falls from 40% to below 20% by day 90 and below 10% by month 6.
- Median mitigation time falls from 190 minutes to below 90 minutes by day 120 and below 60 minutes by month 6.
- At least 95% of qualifying incidents meet their initial customer-communication deadline by month 3.
- At least 95% of published incidents meet their required update cadence by month 3.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without loss of Tier 0 or Tier 1 detection coverage.
- Monthly pages fall from 3,400 to no more than 1,500 by day 90, with alert actionability and missed-detection reviews used as countermeasures against unsafe suppression.
- 100% of new paging alerts satisfy the owner, runbook, dashboard, action, severity, and escalation quality rules by day 60.
- 100% of required SEV0 and SEV1 postmortems are drafted within 3 business days and reviewed within 5 business days by month 2.
- At least 90% of postmortem actions are completed by their approved due dates by month 6.
- All 53 currently open historical actions are triaged within 30 days; all unaccepted high-risk items are completed within 90 days.
- Repeat incidents with the same unaddressed contributing factor decline by at least 50% within 6 months.
- Every primary rotation has at least six trained responders or a documented, time-limited executive exception by day 120.
- No responder is routinely scheduled more frequently than one primary week in six by day 120.
- Two end-to-end cross-company exercises, including regional and ledger scenarios, are completed before the audit, with all critical findings assigned and tracked.
- Monthly availability meets or exceeds the 99.95% contractual target by month 6, with exceptions reviewed at the executive reliability meeting.
- SLA credits decline by at least 50% on an annualized trailing basis by month 8.
- The month-7 mock audit finds no unowned high-risk incident-response control gap and at least 95% of sampled incidents have complete operating evidence.
Steps (19):
1. Establish ownership, authority, and funding
Launch the program within 48 hours under an executive sponsor. Give one program owner authority to standardize incident management across all 28 teams.
- Name the CTO or equivalent as executive sponsor and a Head of Incident Management or Reliability as directly accountable owner.
- Form a working group with Engineering, SRE or Platform, Product, Support, Customer Success, Communications, Security, Legal, Compliance, Risk, HR, Finance, and Internal Audit.
- Approve authority for an incident commander to stop deployments, roll back releases, disable features, shift traffic, invoke continuity plans, and pause payment processing when integrity is at risk.
- Preserve financial controls. The incident commander may coordinate ledger recovery but may not bypass dual-control, reconciliation, or privileged-access requirements.
- Fund paging tools, compensation, training, observability work, exercises, and dedicated reliability capacity.
- Reserve engineering capacity for incident remediation. Start with 10% of capacity and adjust through quarterly risk reviews.
- Record the current baselines: 31 customer-impacting incidents, 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, $1.3M in credits, 3,400 monthly alerts, 85% noise, and 11 of 64 actions closed.
- Maintain a risk register for staffing gaps, shared-ledger concentration, regional failover, alert coverage, third parties, and audit readiness.
2. Install immediate minimum controls (depends on: 1)
Put an interim process in place during the first seven days. Do not wait for tool consolidation, policy perfection, or the SOC 2 audit.
- Publish a one-page interim severity guide and incident declaration procedure.
- Establish one continuously monitored incident declaration path through chat, telephone, and the paging system.
- Create a standard incident channel, conference bridge, incident document, and event naming convention.
- Staff an interim primary and backup incident commander at all times. Compensate this duty retroactively under the final compensation policy.
- Give trained duty personnel access to the status page, paging system, dashboards, support queue, service catalog, and emergency contacts.
- Require an incident commander to be named within 10 minutes for every suspected major incident.
- Direct Support to escalate credible customer reports immediately rather than waiting for engineering confirmation.
- Triage all 53 open historical postmortem actions. Complete, re-plan, or formally risk-accept the items affecting ledger integrity, payment duplication, regional resilience, security, and detection first.
- Hold a daily 15-minute operational review until permanent controls are working.
3. Create the service and dependency catalog (depends on: 1)
Build a reliable ownership map for all production services and customer journeys. This is the basis for paging, escalation, impact assessment, and audit evidence.
- Inventory all 180 services, Kubernetes clusters, AWS accounts, data stores, queues, external processors, banking partners, and customer-facing endpoints.
- Assign each component a single accountable team, primary responder group, secondary escalation group, engineering manager, and product owner.
- Classify services as Tier 0, Tier 1, Tier 2, or Tier 3 according to financial integrity, customer impact, dependency centrality, and contractual obligations.
- Treat the ledger, payment orchestration, authentication, settlement, reconciliation, and critical shared infrastructure as Tier 0 or Tier 1.
- Map every important customer journey to its service, database, cloud-region, and third-party dependencies.
- Record SLOs, RTOs, RPOs, data classification, dashboards, runbooks, deployment controls, feature flags, and failover methods.
- Assign separate but coordinated responders for the ledger application and the shared PostgreSQL platform.
- Document whether each service is active-active, active-passive, or region-bound. Identify dependencies that make nominal regional redundancy ineffective.
- Make missing ownership or missing runbooks a release-blocking risk for Tier 0 and Tier 1 services.
4. Adopt severity and incident lifecycle standards (depends on: 1)
Approve one impact-based severity model for operational, security, data, and third-party incidents. When evidence is incomplete, start at the higher credible severity and downgrade later.
- SEV0, crisis: Use for actual or credible unauthorized, lost, duplicated, or corrupted movement of money; ledger integrity loss; material security compromise; material data exposure; both-region failure; or an event likely to require crisis or regulatory management. Page all roles immediately, engage executives, Security, Legal, Compliance, and Risk, and consider pausing payment activity.
- SEV1, critical: Use for widespread inability to initiate, process, settle, or reconcile payments; a core journey failing without a viable workaround; material regional impact; fast SLA-budget exhaustion; or an imminent integrity risk. Staff all incident roles, notify the executive duty officer, and publish customer communications.
- SEV2, major: Use for a material customer subset, one or more critical customers, significant degradation with a workaround, partial transaction failure, or a likely contractual impact. Assign an incident commander and subject-matter responders; add communications and scribe roles whenever customers are affected.
- SEV3, minor: Use for localized, low-impact degradation with no financial-integrity, security, regulatory, or material contractual risk. The owning team leads the response and keeps an internal record; external communication is not normally required.
- Base severity on actual or credible impact, not the seniority of the reporter, number of alerts, or presumed complexity of the fix.
- Permit any employee to declare an incident. Only the incident commander may lower severity after recording the evidence and rationale.
- Use the lifecycle Detected, Declared, Triaged, Mitigated, Monitoring, Resolved, and Reviewed.
- Define mitigation as the end of customer impact. Define resolution only after stability, backlog processing, transaction recovery, and required ledger reconciliation are complete.
- Start measurement from the earliest reliable indication of impact, including telemetry, customer reports, and partner notifications.
5. Define roles and sustainable 24x7 staffing (depends on: 3, 4)
Separate command from technical remediation. This allows trained commanders to coordinate any incident without asking engineers to debug code they do not own.
- Incident commander: Owns severity, priorities, role assignment, escalation, decision cadence, mitigation strategy, handoffs, and final closure. One person has command at a time.
- Communications lead: Owns internal notices, status-page updates, account-manager briefs, approved customer language, and coordination with Legal or regulators.
- Scribe: Maintains a timestamped timeline of observations, decisions, commands, owners, and status changes. Automation may assist but does not replace human validation for SEV0 and SEV1.
- Subject-matter responders: Diagnose and mitigate only services or domains for which they have accepted ownership, training, access, and runbooks.
- Executive duty officer: Removes organizational obstacles and approves exceptional business decisions. This role does not take command unless a formal transfer occurs.
- Security, Legal, Compliance, Support, Vendor Management, and Business Continuity join according to predefined triggers.
- Create a company-wide incident-command rotation with at least eight certified primary commanders and eight qualified backups. Use weekly rotations with explicit handoffs.
- Create similarly sustainable communications and scribe pools using Engineering Operations, Support, Customer Operations, and Communications personnel.
- Group service responders into approximately 8–12 coherent product or platform domains rather than creating 28 fragile rotations. Each domain rotation should normally contain at least six trained responders.
- Do not place an engineer into another team's responder pool without training, access, runbooks, shadow shifts, and explicit acceptance by both teams.
- Maintain dedicated database-platform and ledger-application escalation coverage for the shared PostgreSQL environment.
- Require distinct people for commander, communications, and primary technical lead during SEV0 and SEV1 incidents.
- Require a verbal and written handoff when any incident role changes. Record the exact time and the new role owner.
6. Implement compensation and fatigue safeguards (depends on: 5)
End unpaid on-call before expanding coverage. Treat availability, interrupted personal time, and overnight recovery as compensable work.
- Pay a fixed stipend for each primary on-call week and a secondary stipend equal to a defined percentage of the primary amount.
- Pay a higher holiday stipend. Apply overtime and call-out rules to non-exempt employees as required by law.
- Give exempt employees a minimum call-out credit or equivalent paid recovery time for material after-hours work.
- Provide a paid recovery day after prolonged overnight work, a SEV0, or a qualifying SEV1. Managers must arrange daytime coverage rather than expecting normal output.
- Have HR, Finance, and employment counsel publish dollar amounts, tax treatment, eligibility, and payroll procedures within 14 days. Apply the policy consistently across teams and locations.
- Target rotations no more frequent than one week in six. Exceptions require a time-limited staffing plan and executive risk acceptance.
- Avoid consecutive primary and secondary weeks. A person must not be primary for two simultaneous domain rotations.
- Track after-hours pages, sleep interruptions, swaps, missed acknowledgements, and reported burnout by rotation.
- Trigger a staffing or alert-remediation review when a rotation averages more than two after-hours pages per person per week for four weeks.
- Permit responders to declare themselves temporarily unfit after overnight work without performance penalty.
7. Consolidate incident and paging tooling (depends on: 4)
Create one operational system of record while migrating safely from the six current alerting tools. Consolidation must reduce ambiguity without creating a monitoring gap.
- Select one enterprise paging and escalation platform and one integrated incident record.
- Initially ingest events from all six tools. Deduplicate, correlate, and route them through the new platform before retiring sources.
- Integrate paging with chat, conference bridges, ticket tracking, the service catalog, observability tools, and the customer status page.
- Automatically capture declaration time, acknowledgements, role assignments, severity changes, messages, decisions, mitigated time, and resolved time.
- Use role-based access, multifactor authentication, break-glass controls, immutable audit logs, and periodic access reviews.
- Provide mobile and telephone fallback paths if chat, identity, or the primary paging tool is unavailable.
- Test paging, escalation, status publication, and conference access every week.
- Retire a legacy alert path only after its signals have named owners, successful end-to-end tests, and at least two weeks of verified operation in the new platform.
8. Improve detection and enforce alert quality (depends on: 3, 7)
Shift detection toward customer journeys, payment outcomes, and ledger integrity. Infrastructure metrics alone will not solve the current customer-first detection problem.
- Instrument payment initiation, authorization, orchestration, settlement, reconciliation, refunds, authentication, and reporting with SLOs and business-level success metrics.
- Run external synthetic transactions and API checks from outside the production boundary and from both AWS regions.
- Monitor transaction failure rates, processing latency, queue age, unprocessed volume, reconciliation breaks, unexpected ledger balances, duplicate identifiers, regional asymmetry, and third-party response quality.
- Correlate application telemetry with Kubernetes, AWS, PostgreSQL, network, deployment, and feature-flag events.
- Route high-priority support cases and credible partner notifications into the same incident declaration path within five minutes.
- Define noise as a page that is duplicate, informational, unactionable, non-production, or requires no timely human action.
- Require every paging alert to name an owner, affected service, urgency, customer or SLO risk, dashboard, runbook, expected action, deduplication key, and escalation policy.
- Send non-urgent conditions to a ticket queue rather than a pager.
- Run new alerts in shadow mode for at least seven days unless an emergency risk exception is approved. Test both firing and recovery behavior.
- Review any alert with less than 50% actionability or more than three firings in seven days within two business days.
- Never silently disable a noisy alert. Verify compensating detection, record the decision, and assign a correction owner first.
- Review alert actionability, false positives, missed detection, and page load with every responder group each month.
9. Codify acknowledgement and escalation paths (depends on: 3, 4, 5, 7, 8)
Make escalation automatic, time-bound, and independent of personal contacts. The clock starts when a qualifying signal or customer report enters the system.
- For SEV0 and SEV1, page the owning primary immediately; page the secondary after five unacknowledged minutes; page the domain manager and company incident commander at 10 minutes; and engage the executive duty officer by 15 minutes.
- For SEV2, require primary acknowledgement within 10 minutes and incident-command assignment within 15 minutes. Escalate to the secondary and manager when either target is missed.
- For SEV3, require acknowledgement within 30 minutes when immediate production action is needed. Otherwise create a prioritized work item.
- Automatically page the company incident commander for any credible integrity or security concern, cross-team event, customer-visible Tier 0 failure, regional event, or unresolved ownership question.
- If impact remains unknown after 15 minutes, raise severity rather than waiting for certainty.
- Let the incident commander summon dependency owners, cloud support, database support, payment processors, banking partners, and vendors through maintained escalation contacts.
- Test vendor contacts and premium-support entitlements quarterly.
- Route alerts with no valid owner to the central command rotation, then treat the missing ownership record as a control defect.
- Require human acknowledgement. Delivery to a device or chat channel does not count.
- Record every missed acknowledgement, failed escalation, and manual contact workaround for review.
10. Standardize live incident execution (depends on: 4, 5, 7, 9)
Give responders one concise operating procedure for the first minutes through resolution. Prioritize limiting customer and financial harm before proving a root cause.
- Open a dedicated channel, bridge, incident record, and timeline immediately for SEV0 through SEV2.
- Have the incident commander state severity, known impact, current hypothesis, immediate objective, assigned roles, and next update time.
- Freeze unrelated production changes during SEV0 and SEV1 incidents. Record exceptions approved by the incident commander.
- Prefer reversible mitigation such as rollback, feature disablement, traffic isolation, rate limiting, workload shedding, or partner rerouting.
- Use pre-approved runbooks for region failover, Kubernetes recovery, PostgreSQL failover, credential rotation, queue recovery, and payment suspension.
- Guard against split-brain, replay, duplication, and out-of-order processing during regional or database recovery.
- Require reconciliation and controlled backlog processing before declaring payment or ledger incidents resolved.
- Keep diagnosis and mitigation workstreams separate when enough responders are available.
- State decisions and owners aloud and in the incident record. Avoid unrecorded direct-message command paths.
- Require stability for a severity-specific observation period before closure. Reopen the incident if impact recurs during that period.
- Conduct an explicit operational handback to the owning team, Support, and Customer Success.
11. Standardize internal, customer, and regulatory communications (depends on: 4, 5, 7, 10)
Communicate known impact early without waiting for a root cause. Use approved facts, acknowledge uncertainty, and give the next update time.
- For SEV0, notify executives, Support, Customer Success, Security, Legal, Compliance, and Risk within 10 minutes. Publish an initial customer status within 15 minutes when disclosure is operationally and legally appropriate, then update every 15 minutes.
- For SEV1, notify internal stakeholders within 15 minutes, publish an initial customer status within 15 minutes, and update at least every 30 minutes.
- For customer-visible SEV2, brief Support and account managers and publish or directly send an initial notice within 30 minutes. Update at least every 60 minutes.
- Do not normally publish SEV3 events. Notify specifically affected customers if contracts or material impact require it.
- Use status states such as Investigating, Identified, Monitoring, and Resolved. State affected capabilities, customer symptoms, workarounds, regions, and next update time.
- Do not speculate about root cause, blame, security scope, recovery time, or data integrity.
- Give account managers a single approved briefing and an affected-customer list. Prohibit contradictory or improvised incident explanations.
- Maintain templates for outages, delays, data-integrity investigation, third-party failure, regional failure, security events, and resolution.
- Issue a resolution notice only after operational recovery and required reconciliation. Provide a customer-facing incident summary within five business days for qualifying events.
- Have Legal and Compliance maintain a jurisdiction, regulator, sponsor-bank, network, cyber-insurer, partner, and contract notification matrix.
- Where applicable, explicitly track the current New York cybersecurity-event notification clock, including the 72-hour requirement, without assuming every incident is reportable.
- Have Legal record the reportability decision, decision time, evidence, approver, deadline, and submission confirmation.
- Allow Security or Legal to limit public detail during an active threat, but require the reason and an alternative stakeholder plan to be recorded.
- Coordinate service-credit calculations and contractual notices with Finance and Customer Success from the same incident record.
12. Make postmortems mandatory and actionable (depends on: 4, 7, 10)
Use postmortems to improve systems and controls, not to assign personal blame. Keep performance or misconduct processes separate from the learning review.
- Require a postmortem for every SEV0 and SEV1.
- Require one for a SEV2 that affected customers, incurred credits, breached an SLO or contract, involved financial or data integrity, repeated a prior failure, exposed a control gap, or lasted more than two hours.
- Permit incident command, Security, Compliance, or the service owner to require a review for a near miss.
- Produce a factual draft within three business days and hold the cross-functional review within five business days.
- Use one template covering executive summary, customer and financial impact, detection source, timeline, response analysis, contributing conditions, control performance, what worked, what failed, and lessons.
- Include why monitoring did or did not detect the event before customers.
- Avoid a single-root-cause assumption. Examine technical, organizational, process, dependency, testing, and incentive factors.
- Give every action one owner, due date, priority, expected risk reduction, verification method, and linked engineering item.
- Classify actions as containment due within 7 days, corrective work due within 30 days, or strategic work normally due within 90 days.
- Require director approval and documented residual-risk acceptance for overdue high-risk actions.
- Verify effectiveness after implementation. Closing a ticket without evidence does not close the action.
- Publish broadly useful reviews internally. Maintain access-restricted versions for security, privacy, personnel, or legally privileged details.
13. Measure performance and review it routinely (depends on: 7, 8, 11, 12)
Use outcome, process, quality, and human-sustainability measures together. Do not reward teams for suppressing declarations or hiding incidents.
- Measure detection time from first impact to first internal signal, declaration time, acknowledgement time, role-staffing time, mitigation time, resolution time, and recurrence.
- Report both median and 90th percentile. Break results down by severity, service tier, customer journey, region, detection source, and owning domain.
- Track customer-first detection, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, and role conflicts.
- Track availability, error-budget consumption, failed-payment volume, delayed value, reconciliation breaks, impacted customers, contractual breaches, and service credits.
- Track alert volume, actionability, duplicates, after-hours pages, missed pages, pages per responder, and tool-source distribution.
- Track required postmortems completed on time, actions completed by due date, action age, verified effectiveness, and repeat contributing factors.
- Track rotation size, on-call frequency, swaps, recovery days, attrition signals, and quarterly responder sentiment.
- Hold a weekly operational review for recent incidents, overdue actions, alert problems, and upcoming risk.
- Hold a monthly executive reliability review covering trends, investment decisions, accepted risks, and SLA exposure.
- Hold a quarterly resilience and control review with Security, Compliance, Risk, Internal Audit, and Product leadership.
- Use team scorecards to direct investment and assistance, not individual performance penalties.
- Reconcile dashboard data against a monthly sample of incident records and customer cases to detect metric gaming or missing incidents.
14. Train and certify participants (depends on: 4, 5, 9, 10, 11, 12)
Train people before assigning full independent duty. Use paid working time for training, shadowing, exercises, and certification.
- Train all employees to recognize impact, declare an incident, and find the incident channel and status page.
- Train engineers and Support on severity, escalation, customer-report handling, evidence preservation, and financial-integrity precautions.
- Certify incident commanders through instruction, tabletop exercises, shadow incidents, and observed command performance.
- Train communications leads in status writing, contractual communications, regulator escalation, and avoiding unsupported claims.
- Train scribes in timestamping, decision capture, evidence hygiene, and separating fact from hypothesis.
- Require subject-matter responders to demonstrate dashboard, runbook, rollback, failover, and access competence for their assigned domain.
- Add training to new-hire onboarding and repeat role-specific certification annually.
- Appoint an incident-management champion in each of the 28 teams to collect feedback and support local adoption.
- Conduct listening sessions focused on pager fairness, cross-team boundaries, tooling friction, and psychological safety.
- Publish that command duty is process coordination, not responsibility for understanding or repairing another team's code.
15. Pilot and expand the on-call model (depends on: 3, 5, 6, 8, 9, 14)
Pilot the model on the highest-risk customer journeys before expanding it. Correct staffing, alert, access, and compensation defects at each stage gate.
- Start with the ledger, payment orchestration, Kubernetes platform, PostgreSQL platform, authentication, settlement, and Support intake.
- Run the command, communications, and domain rotations in parallel with existing paths for two weeks.
- Verify primary and secondary coverage, handoffs, access, runbooks, paging, conference access, status publication, and compensation processing.
- Require at least two shadow shifts before independent primary duty.
- Review every pilot page within one business day for routing accuracy, actionability, responder load, and missing context.
- Expand by customer journey and dependency domain, not by arbitrary team order.
- Provide central first-line triage if useful, but keep technical remediation with the accepted service owner.
- Do not use contractors or a managed service as the sole incident commander or sole owner of payment and ledger remediation.
- Permit temporary shared domain rotations only after service owners document training, access, runbooks, and escalation boundaries.
- Set an executive-reviewed deadline and remediation plan for any production service that cannot provide sustainable 24x7 ownership.
16. Exercise regional, ledger, and communication failures (depends on: 10, 11, 14, 15)
Validate the process under realistic conditions before relying on it. Begin in tabletop and staging environments, then use controlled production tests where risk permits.
- Run a company-wide incident-command tabletop within 30 days of policy approval.
- Exercise loss of one AWS region, Kubernetes control-plane degradation, shared PostgreSQL failure, payment-processor failure, queue backlog, credential compromise, and suspected duplicate payments.
- Exercise a simultaneous operational and security event to test command boundaries and disclosure control.
- Exercise status-page failure and loss of the primary chat or paging provider.
- Exercise overnight staffing, role handoff, executive escalation, account-manager messaging, and a potential regulator-notification decision.
- Validate backups, restore procedures, RPO, RTO, failover prerequisites, and post-recovery reconciliation.
- Do not inject uncontrolled changes into the production ledger. Use replicas, staging, simulations, or tightly governed production tests.
- Record exercise observations as tracked actions under the same ownership and due-date rules as incident actions.
- Run at least one domain exercise per quarter and two cross-company exercises before the SOC 2 audit.
17. Execute a time-boxed enterprise rollout (depends on: 2, 6, 7, 11, 12, 13, 15)
Use fixed implementation waves so the audit deadline does not become the start date. Report progress weekly and escalate missed stage gates as business risks.
- Days 0–7: Establish governance, interim command coverage, one declaration path, provisional severity, and daily operational reviews.
- By day 14: Approve the core policy, role definitions, communications timings, compensation design, and historical-action triage.
- By day 30: Complete Tier 0 ownership, certify the first command roster, begin the on-call pilot, enable standard incident records, and run the first tabletop.
- By day 60: Provide 24x7 coverage for all Tier 0 and Tier 1 customer journeys, integrate the six alert sources, and enforce postmortem tracking.
- By day 90: Migrate critical paging, implement customer-journey detection, complete status and regulatory playbooks, and materially reduce alert noise.
- By day 120: Assign sustainable ownership and escalation for every production service and complete the first controlled regional or continuity exercise.
- By day 180: Complete tool retirement decisions, verify action closure, rerun weak scenarios, and demonstrate improving detection and mitigation trends.
- In month 7: Conduct a mock audit and executive readiness review, leaving at least one month to correct evidence or operating defects.
- Use exception records with owners, expiry dates, compensating controls, and executive approval. Do not allow indefinite verbal exceptions.
18. Build SOC 2 evidence as the process operates (depends on: 1)
Design evidence collection at the start rather than reconstructing it before the audit. Demonstrate both control design and sustained operation.
- Map the incident process to applicable SOC 2 criteria with Compliance and the auditor, including detection, response, communication, change management, access, availability, and corrective action.
- Maintain approved, version-controlled policies, procedures, severity definitions, role descriptions, and exception records.
- Preserve rotation schedules, compensation activation, training attendance, certification, paging tests, access reviews, and exercise results.
- Preserve incident declarations, timestamps, role assignments, communications, decisions, status updates, postmortems, and corrective-action evidence.
- Record regulatory and contractual notification assessments, including decisions that no notification was required.
- Define retention, confidentiality, legal-hold, and access requirements for operational and security records.
- Sample evidence monthly and trace incidents from initial signal through action verification.
- Have Internal Audit or an independent control owner test the process in months 4 and 6.
- Correct control failures through tracked actions rather than editing historical records.
- Conduct the formal mock audit in month 7 using the same evidence populations expected for the external audit.
19. Sustain accountability and continuous improvement (depends on: 13, 17, 18)
Make incident management an operating discipline rather than an audit project. Keep policy, staffing, tools, and investment aligned with changing customer and system risk.
- Assign permanent owners for the incident policy, paging platform, status page, service catalog, training program, and metrics.
- Review severity thresholds, communication timings, compensation, and staffing at least annually and after material incidents.
- Use incident trends to prioritize architectural work on the shared ledger, regional independence, deployment safety, dependency isolation, and graceful degradation.
- Review repeat incidents and repeat contributing factors quarterly. Require executive action when remediation repeatedly loses priority.
- Survey responders quarterly and publish actions addressing fatigue, fairness, psychological safety, and tool friction.
- Recognize effective incident leadership, early declaration, useful postmortems, and preventive work.
- Prohibit retaliation for good-faith incident declaration or escalation.
- Provide the board or risk committee a quarterly summary of severe incidents, SLA exposure, regulatory events, overdue high-risk actions, and resilience investment.
- Maintain annual budget and headcount planning for sustainable rotations, reliability engineering, observability, and continuity testing.
--- PROPOSAL 3 (agent qwen3.8-max_initial_3, alibaba/qwen3.8-max) ---
Estimated complexity: high
Success metrics: - **MTTD reduced from 22 minutes to <5 minutes** within 6 months of full rollout.
- **Customer-first detection rate reduced from 40% to <10%** within 6 months.
- **MTTM for SEV1 incidents reduced from 3 h 10 min to <60 minutes** within 9 months.
- **Monthly SLA credit payouts reduced from ~$325K to <$100K per quarter** within 12 months.
- **Alert volume reduced from 3,400/month to <600 actionable alerts/month** within 6 months; signal-to-noise ratio >80%.
- **Postmortem completion rate: 100% of SEV1/SEV2 incidents** have a blameless postmortem within 5 business days.
- **Postmortem action-item completion rate >90% within 30 days** of the postmortem (up from ~17%).
- **Zero incidents with >15 minutes of unowned command** (down from 2 incidents with >1 hour).
- **100% on-call coverage**: all 28 teams staffed with primary + secondary on-call 24×7 within 14 weeks.
- **On-call compensation adopted**: 100% of on-call engineers receiving stipends and page pay; on-call satisfaction score ≥4/5 in quarterly survey.
- **Status-page first-update within 15 minutes for SEV1 and 30 minutes for SEV2**, 100% compliance.
- **SOC 2 Type II audit passed** at month 8 with zero incident-response findings.
- **All 260 engineers trained**; 56+ certified ICs and 28+ certified CLs active within 14 weeks.
- **Six alerting tools consolidated to one** within 6 months; legacy tools decommissioned.
- **Regulator notification process tested**: at least one tabletop exercise includes a NY DFS / FinCEN notification drill, and the legal/compliance playbook is documented and approved.
- **Quarterly IMC reviews held consistently** with published KPI dashboards and action-item tracking.
- **On-call participation resistance resolved**: <10% of engineers report 'unwilling to participate' in the 6-month pulse survey (baseline to be measured in S1).
Steps (13):
1. Assess Current State and Baseline Metrics
Build an evidence-based picture of the current incident management reality before designing anything new.
Collect and catalog the last 12 months of incident data: all 31 customer-impacting incidents, 3,400 monthly alerts, on-call coverage gaps across the 28 teams, and the 11 of 64 closed postmortem action items. Interview one lead from each of the 28 teams to surface pain points, political concerns (the 'carrying a pager for other teams' pushback), and tool sprawl.
Deliverables to produce:
- **Alert inventory**: which of the six alert tools feed which teams, alert volume per team, noise rate per tool, and overlap between tools.
- **Incident timeline analysis**: median detection-to-notification-to-mitigation-to-resolution times, who detected first (internal vs. customer), who mitigated, and where handoff gaps occurred.
- **On-call coverage map**: which 12 teams have on-call, which 16 do not, rotation length, compensation status, and escalation paths (or lack thereof).
- **Postmortem audit**: format variance, action-item tracking gaps, and the two incidents with no clear owner for over an hour.
- **Tooling and integration audit**: Kubernetes observability stack, six alerting tools, status-page provider, communication channels (Slack, email, phone), and any existing runbooks.
- **Compliance gap analysis**: SOC 2 Type II CC7.3/CC7.4 requirements vs. current practice, with a risk register for the eight-month window.
- **Peer benchmarking**: incident management practices at 3–4 comparable B2B fintech platforms (e.g., Plaid, Stripe, Adyen) for severity scales, on-call comp, and MTTR targets.
2. Secure Executive Sponsorship and Form the IM Governance Body (depends on: 1)
Anchor the program with visible, top-down authority so that 28 teams adopt changes they did not individually request.
The CEO email about 'outages we hear about from clients' is a ready-made mandate. Convert it into a formal sponsorship structure.
- Appoint an **executive sponsor** (CTO or VP Engineering) who owns the program end-to-end and reports to the CEO monthly.
- Create an **Incident Management Office (IMO)**: one dedicated senior incident management lead, one tooling/platform engineer, and one part-time data analyst.
- Establish an **Incident Management Council (IMC)**: one engineering manager from each of the 28 teams, plus the VP of Customer Success, a compliance lead, and a security lead. The IMC meets bi-weekly during rollout, monthly thereafter.
- Draft and circulate an **executive mandate memo** that states: incident response is a shared operational obligation, not a per-team favor; participation in on-call rotations is a condition of employment for production-facing roles; and the program is not optional pending the SOC 2 audit.
- Allocate a dedicated budget line for on-call compensation, tooling consolidation, status-page licensing, training, and external facilitation.
3. Define Severity Levels and Automatic Triggers (depends on: 1)
Replace the current ad-hoc triage with a five-level severity taxonomy that every engineer, support agent, and account manager can apply in under 60 seconds.
- **SEV1 – Critical**: Ledger data corruption or loss, complete payment processing halt, confirmed data breach affecting customer PII or funds, or regulatory reporting breach. Triggers: automatic all-hands page to on-call, CTO + CEO paged within 10 minutes, dedicated bridge call within 5 minutes, status-page update within 15 minutes, regulator notification assessment within 1 hour, customer comms within 30 minutes.
- **SEV2 – High**: Payment processing degraded >30% throughput or >5% error rate, single-region failover failure, ledger read-only mode, or any condition likely to breach the 99.95% SLA within the current window. Triggers: primary + secondary on-call paged, incident commander assigned within 10 minutes, bridge call within 15 minutes, status-page update within 30 minutes, VP Engineering notified within 20 minutes.
- **SEV3 – Medium**: Non-critical service degradation affecting <30% of customers, single non-ledger microservice outage with working fallback, or elevated latency above SLA threshold but below halt. Triggers: primary on-call paged, IM notified within 30 minutes, status-page update within 1 hour if customer-visible, daily standup update.
- **SEV4 – Low**: Degraded internal tooling, minor UX bug with workaround, non-customer-facing alert. Triggers: next-business-day response, ticket created, no page unless on-call agrees.
- **SEV5 – Informational / Noise**: Cosmetic issues, planned-maintenance notifications, alert misfires logged for tuning. Triggers: no page, logged for weekly alert-quality review.
Define **escalation rules**: any SEV3 unresolved after 4 hours auto-escalates to SEV2; any SEV2 unresolved after 2 hours auto-escalates to SEV1. Severity can be **downgraded** only by the incident commander with IMC notification.
Publish the taxonomy as a one-page decision tree, a Slack slash-command (`/sev`), and an integration into the alerting tool so that every alert carries a suggested severity.
4. Define Incident Roles and Staffing Model (depends on: 3)
Codify four mandatory roles for every SEV1/SEV2 incident and optional roles for SEV3, then solve the 24×7 staffing problem across 28 teams.
**Roles**
- **Incident Commander (IC)**: owns the incident end-to-end, declares severity, assigns tasks, authorizes mitigations, decides when to escalate or stand down. Never writes code during the incident.
- **Communications Lead (CL)**: owns status-page updates, internal Slack channels, account-manager briefings, and regulator notifications. Separate from the IC so the IC can focus on mitigation.
- **Scribe / Timeline Keeper**: logs every decision, action, and timestamp in the incident channel and the incident-management tool. Produces the raw timeline for the postmortem.
- **Subject-Matter Responders (SMRs)**: 1–3 engineers from the owning team(s) who diagnose and fix. For the shared PostgreSQL ledger, a dedicated DBA responder is always required.
**24×7 Staffing via a Three-Tier Follow-the-Sun Model**
- **Tier 1 – Front-line on-call**: Primary + secondary responder per team, paged first. Covers the team's own services.
- **Tier 2 – Platform / SRE on-call**: A dedicated 6-person SRE rotation covering cross-cutting infrastructure: Kubernetes, the shared PostgreSQL ledger, networking, and the two AWS regions. This tier directly addresses the 'carrying a pager for other teams' concern by absorbing infrastructure incidents.
- **Tier 3 – IMC escalation**: Engineering managers and the IMO on-call for multi-team or SEV1 incidents. Provides the IC and CL when no team-level IC is available.
**Follow-the-Sun**: If any engineering hub exists in a second timezone, use it for overnight Tier 1 coverage. If not, partner with a managed on-call service for overnight first-response triage (severity declaration + paging the correct team), reducing 3 a.m. pages for NY-based engineers.
**IC and CL pools**: Nominate at least 2 ICs and 1 CL per team (56 ICs, 28 CLs minimum). ICs are trained and certified before they rotate. For SEV1 incidents, the IC must be a certified IC from the IMC pool, not just 'whoever is around'.
**Ledger-specific rule**: Because the PostgreSQL ledger is shared, a **Ledger Duty Officer** from the SRE Tier 2 is always on the bridge for any incident touching ledger services, regardless of which team owns the failing microservice.
5. Design On-Call Rotations, Compensation, and Alert-Quality Rules (depends on: 4)
Make on-call sustainable, fairly compensated, and free of alert noise so engineers stop resisting participation.
**Rotation Design**
- 7-day rotations, one primary + one secondary per team per week. No engineer is on-call more than one week in four.
- Minimum 48-hour rest between rotations. No on-call during approved PTO.
- All 28 teams participate. Teams without current on-call get a 90-day ramp with a shadow rotation before going live.
- Tier 2 SRE rotation: 6 engineers, one week on / five weeks off, with a dedicated backup.
**Compensation Package**
- **Base on-call stipend**: $500 per week of primary on-call, $250 for secondary, paid regardless of whether pages fire.
- **Page pay**: $75 per acknowledged page outside business hours; $150 if the page leads to active incident work.
- **Time-off-in-lieu (TOIL)**: Any engineer who works >4 hours overnight (00:00–06:00 local) gets a full TOIL day. >8 hours in a single incident gets 1.5 TOIL days.
- **SEV1 bonus**: $300 flat bonus for every engineer who actively works a SEV1 incident, paid within the next pay cycle.
- **Annual on-call cap**: No engineer exceeds 13 weeks of on-call per year. Exceeding the cap triggers a mandatory team-staffing review.
- Budget estimate: ~$420K/year for stipends and page pay across 28 teams; present this to the CFO as a fraction of the $1.3M annual SLA credit cost.
**Alert-Quality Rules (the '85% noise' problem)**
- Every alert must carry: owning team, suggested severity, runbook link, and a 30-day noise score.
- **Alert budget**: each team gets a maximum of 100 actionable alerts per month. Exceeding the budget triggers a mandatory alert-tuning session with the IMO.
- **Noise threshold**: any alert that fires >10 times in 7 days with no human action is auto-flagged for suppression or tuning within 14 days.
- **Alert review cadence**: weekly 30-minute alert-quality review per team; monthly cross-team alert review in the IMC.
- **Sunset rule**: alerts with no runbook are demoted to SEV5 after 30 days and suppressed after 60 days unless a runbook is written.
- Target: reduce monthly alert volume from 3,400 to <600 actionable alerts within 6 months.
6. Build Detection, Escalation, and Communication Paths (depends on: 3, 4)
Eliminate the 22-minute median detection gap and the 40% customer-first-detection rate with layered monitoring and a single escalation spine.
**Detection Layers**
- **Synthetic transactions**: run a payment end-to-end through the full stack (API → service → ledger → confirmation) every 60 seconds from both AWS regions. Alert if latency >2× baseline or any step fails. This catches what per-service metrics miss.
- **Customer-traffic anomaly detection**: monitor API error rates, payment success rates, and latency percentiles per customer cohort. Alert on >2σ deviation.
- **SLO-based alerting**: define SLIs for the 99.95% SLA (availability, latency p99, ledger consistency). Alert when error budget burn rate exceeds threshold, before the SLA actually breaches.
- **Infrastructure health**: Kubernetes node/pod health, PostgreSQL replication lag, disk I/O, and cross-region latency.
- **Support-ticket spike detection**: if >5 customers open tickets about the same symptom within 10 minutes, auto-create a SEV3 candidate.
**Escalation Path**
- Alert fires → PagerDuty routes to Tier 1 primary → 5-min no-ack → Tier 1 secondary → 10-min no-ack → Tier 2 SRE → 15-min no-ack → IMC on-call manager → 20-min no-ack → VP Engineering auto-page.
- Any SEV1 declaration auto-pages the CTO, opens a dedicated Slack channel + Zoom bridge, and notifies the CL.
- **No incident goes unowned for >15 minutes.** If no IC is assigned by minute 15, the IMC on-call manager assumes IC role by default.
**Internal Communications**
- Dedicated Slack channels: `#inc-sev1`, `#inc-sev2`, `#inc-sev3` (auto-created per incident), plus `#inc-updates` for broadcast.
- IC posts a structured update every 15 minutes (SEV1), 30 minutes (SEV2), 1 hour (SEV3) into the incident channel.
- CL posts a summary to `#inc-updates` and notifies relevant engineering managers.
**Customer Communications**
- **Status page**: auto-updated via API. SEV1: first update within 15 minutes, then every 30 minutes until resolved. SEV2: first update within 30 minutes, then every hour. SEV3: within 1 hour if customer-visible.
- **Account managers**: CL briefs AMs via a dedicated Slack channel within 30 minutes (SEV1) or 1 hour (SEV2). AMs contact their top-20 revenue accounts directly.
- **Customer email/SMS**: for SEV1 and SEV2, automated notification to all 2,100 customers via the status-page subscription system within 30 minutes.
- **Regulator notification**: Legal/Compliance assesses within 1 hour whether NY DFS, FinCEN, or card-network notification is required. If yes, file within the regulatory deadline (typically 72 hours for NY DFS cybersecurity events). Log the decision and filing in the incident record.
**Post-resolution**: CL publishes a 'resolved' update within 30 minutes of mitigation. For SEV1/SEV2, a preliminary customer-facing RCA summary is published within 5 business days.
7. Standardize Postmortems with Tracking and Accountability (depends on: 3, 6)
Fix the '11 of 64 action items closed' problem with a mandatory, uniform, blameless postmortem process backed by engineering-manager accountability.
**When Mandatory**
- All SEV1 and SEV2 incidents: postmortem required within 5 business days.
- SEV3 incidents: postmortem required if >50 customers affected, if the incident lasted >4 hours, or if it was customer-detected.
- SEV4/SEV5: optional, but any recurring SEV4 (≥3 times in 30 days) triggers a mandatory review.
**Format (single template, enforced by tooling)**
- Incident summary (severity, duration, customers affected, revenue impact, SLA credit exposure).
- Timeline (auto-generated from scribe notes + alert timestamps).
- Detection analysis: how was it detected, why did it take X minutes, could it have been faster.
- Root cause analysis using 5-Whys or fault-tree, not blame.
- Contributing factors (process, tooling, staffing, knowledge gaps).
- Impact quantification: customers affected, transactions failed, SLA credits triggered.
- Action items: each with a **named owner**, **due date**, **priority**, and **ticket in Jira**.
- Lessons learned and what went well.
**Blameless Review Meeting**
- Held within 5 business days, facilitated by the IMO or a trained facilitator (never the IC of that incident).
- All responders, the CL, relevant engineering managers, and an IMC representative attend.
- Ground rules: focus on system and process failures, not individual mistakes. The facilitator enforces this.
- Meeting recorded; notes published to the engineering-wide wiki within 48 hours.
**Action-Item Tracking and Accountability**
- Every action item is created as a Jira ticket with a due date and a named owner.
- **Engineering managers are accountable**: action-item completion is a standing agenda item in the bi-weekly IMC meeting. Any action item >7 days overdue is escalated to the VP Engineering.
- **Completion gate**: no team may close its postmortem until 100% of its action items have Jira tickets. Postmortem is 'closed' only when all tickets are resolved.
- **Quarterly audit**: the IMO audits action-item completion rates and reports to the IMC and the executive sponsor. Target: >90% completion within 30 days of the postmortem.
- Link postmortem quality and action-item completion to team health scores and engineering-manager performance reviews.
8. Define KPIs, Dashboards, and Governance Reviews (depends on: 7)
Create a measurable feedback loop so leadership can see whether the program is working and where to intervene.
**Primary KPIs (tracked weekly, reported monthly)**
- **MTTD** (Median Time to Detect): target <5 minutes (from 22).
- **MTTA** (Median Time to Acknowledge): target <5 minutes.
- **MTTM** (Median Time to Mitigate): target <60 minutes for SEV1 (from 3 h 10 min), <4 hours for SEV2.
- **Customer-first detection rate**: target <10% (from 40%).
- **SLA compliance**: maintain 99.95%; track monthly SLA credit payouts, target <$100K/quarter (from ~$325K/quarter).
- **Alert signal-to-noise ratio**: target >80% actionable (from ~15%).
- **Alert volume**: target <600/month (from 3,400).
- **Postmortem completion rate**: 100% for SEV1/SEV2 within 5 business days.
- **Action-item completion rate**: >90% within 30 days (from ~17%).
- **On-call health**: pages per engineer per week (target <5), TOIL usage, on-call satisfaction survey score.
- **Unowned incident duration**: target 0 incidents with >15 minutes without an IC.
**Dashboards**
- Real-time operational dashboard (Grafana): current incidents, active alerts, on-call roster, SLA error-budget burn.
- Weekly leadership dashboard (auto-generated): KPI trends, open action items, alert-noise report, on-call load distribution.
- Quarterly IMC scorecard per team.
**Review Cadence**
- **Weekly**: IMO publishes KPI snapshot to `#inc-updates`.
- **Bi-weekly IMC**: review open incidents, overdue action items, alert-quality exceptions, and on-call load.
- **Monthly executive review**: CTO presents KPI trends, SLA credit cost, and risk register to the CEO.
- **Quarterly incident-management review**: deep-dive into trends, training gaps, tooling needs, and process improvements. Output fed into the next quarter's roadmap.
9. Consolidate Tooling and Build the Incident Management Platform (depends on: 2, 3)
Replace six alerting tools and ad-hoc status-page updates with a single, integrated incident management stack.
**Target Tool Architecture**
- **Single alerting and on-call platform** (e.g., PagerDuty or Opsgenie): ingest all alerts, apply severity routing, manage on-call schedules, handle escalations, and send pages. Retire the other five tools within 6 months.
- **Observability consolidation**: standardize on one APM/metrics stack (e.g., Datadog or Grafana Cloud) for all 180 Kubernetes services across both AWS regions. Ensure the shared PostgreSQL ledger has dedicated dashboards.
- **Status page**: a dedicated, branded status page (e.g., Statuspage.io or Instatus) with API integration for auto-updates. Subscribe all 2,100 customers.
- **Incident coordination tool**: integrate incident-management workflows into Slack (auto-create channels, invite responders, post templates) and a dedicated incident record system (e.g., Jira Service Management, incident.io, or Rootly) for timelines, postmortems, and action-item tracking.
- **Runbook repository**: a central wiki (Confluence or Notion) with mandatory runbooks for every alert. No alert goes live without a linked runbook.
**Implementation Tasks**
- Migrate all 28 teams' alert rules into the single platform in three waves (highest-volume teams first).
- Build the severity-based routing rules and escalation policies per S3 and S6.
- Automate status-page updates triggered by severity declaration.
- Build the synthetic-transaction monitor and SLO-based alerting per S6.
- Integrate Jira for automatic action-item ticket creation from postmortems.
- Decommission legacy tools only after all teams have completed training on the new stack.
- Budget: allocate $150K–$250K/year for licensing, plus engineering time for migration.
10. Prepare for the SOC 2 Type II Audit (depends on: 7, 8, 9)
Ensure the incident management process produces the evidence the auditor will need, well before the audit window opens in eight months.
**SOC 2 Requirements to Address (CC7.3, CC7.4, CC7.5)**
- Documented incident response procedures (the severity taxonomy, role definitions, communication templates).
- Evidence of incident detection, response, and recovery for every SEV1/SEV2 incident during the audit period.
- Postmortem records with action-item tracking.
- On-call schedules, training records, and escalation evidence.
- Status-page update logs and customer notification records.
- Regulator notification logs (if any).
**Preparation Tasks**
- The IMO maintains a **SOC 2 evidence folder**: every incident record, postmortem, action-item ticket, status-page update, and training completion certificate is stored and indexed.
- Conduct a **mock SOC 2 audit** at month 5: an internal or external auditor reviews the incident management process end-to-end and identifies gaps.
- Remediate mock-audit findings before month 7.
- Ensure the incident management tool retains all records for at least 12 months (the SOC 2 Type II observation window).
- Document the **chain of custody** for incident records: who accessed, modified, or closed each record.
- Prepare a **narrative document** describing the incident management process, roles, and controls for the auditor.
- Coordinate with the compliance lead to align incident management evidence with the broader SOC 2 scope (access controls, change management, etc.).
11. Design and Deliver Training, Runbooks, and Change Management (depends on: 4, 5, 9)
Equip all 260 engineers, 28 team leads, account managers, and support staff with the knowledge and muscle memory to execute the new process.
**Training Tracks**
- **All 260 engineers** (2-hour session): severity taxonomy, how to acknowledge a page, how to join an incident bridge, how to hand off to an IC, and how to write a postmortem contribution. Delivered in team-level sessions over 4 weeks.
- **IC pool (56+ engineers)** (8-hour certification): incident command techniques, severity declaration, escalation decision-making, bridge facilitation, and blameless postmortem facilitation. Includes two tabletop exercises. Certification valid for 12 months, renewed annually.
- **CL pool (28+ staff)** (4-hour session): status-page writing, customer communication templates, regulator notification triggers, and AM briefing protocol.
- **Account managers and support staff** (1-hour session): how to read the status page, how to escalate a customer report into an incident, and what information to collect.
- **SRE Tier 2** (16-hour onboarding): Kubernetes and PostgreSQL ledger deep-dive, cross-region failover runbooks, and escalation authority.
**Runbooks**
- Every alert must have a runbook before it is routed to on-call. The IMO provides a runbook template and audits compliance weekly.
- Priority runbooks to write first: shared PostgreSQL ledger failover, Kubernetes cluster degradation, payment-processing pipeline failure, cross-region failover, and ledger data-integrity check.
- Runbooks are peer-reviewed and version-controlled.
**Change Management for Adoption**
- Address the 'carrying a pager for other teams' concern directly: publish an FAQ explaining the three-tier model, the SRE Tier 2 absorbing cross-team infrastructure, the compensation package, and the TOIL policy.
- Run **office hours** weekly for the first 8 weeks where any engineer can ask questions or raise concerns.
- Identify **team champions**: one engineer per team who volunteers as an early adopter and peer mentor.
- Publish a **weekly 'incident management newsletter'** during rollout: what changed, what improved, KPI trends, and success stories.
- Make on-call participation a documented expectation in job descriptions and performance reviews for production-facing roles.
12. Execute Phased Rollout, Tabletop Exercises, and Continuous Improvement (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11)
Introduce the process in three waves so teams are not overwhelmed, then validate with exercises and iterate continuously.
**Phase 1 – Weeks 1–6: Foundation**
- Publish the severity taxonomy, role definitions, and communication protocols (S3, S4, S6).
- Launch the single alerting platform for the 12 teams already on-call; begin migration for the other 16.
- Activate the SRE Tier 2 rotation for the shared PostgreSQL ledger and cross-cutting infrastructure.
- Deploy the status page and test the API integration.
- Begin IC and CL training (first cohort of 20 ICs, 10 CLs).
- Publish the on-call compensation package; HR integrates stipends into payroll.
- Write the top 10 priority runbooks.
**Phase 2 – Weeks 7–14: Expansion**
- All 28 teams live on the single alerting platform; legacy tools in read-only mode.
- All 28 teams on the on-call rotation schedule (the 16 new teams in shadow mode for the first 4 weeks).
- Second and third IC/CL training cohorts completed.
- First **tabletop exercise**: simulate a SEV1 ledger corruption scenario with all roles, test the escalation path, status-page updates, and AM briefings. Debrief and fix gaps.
- Postmortem template and Jira integration live; all new incidents use the standard process.
- Alert-tuning sprint: each team reduces its alert volume by 50%.
**Phase 3 – Weeks 15–24: Optimization**
- All teams fully live; legacy alerting tools decommissioned.
- Second **tabletop exercise**: simulate a SEV1 cross-region failure with regulator notification.
- First quarterly IMC review with full KPI dashboard.
- Mock SOC 2 audit (month 5) and remediation.
- Retrospective on the rollout: survey all 260 engineers for feedback, adjust compensation or rotation rules if needed.
- Establish the **continuous improvement cadence**: quarterly process review, annual severity-taxonomy review, and annual on-call compensation benchmarking.
**Ongoing Governance**
- The IMC owns the process document and approves changes.
- The IMO tracks all KPIs and reports to the CTO monthly.
- Any process change requires IMC approval and a 2-week notice period before enforcement.
- Annual external benchmarking against peer B2B payments platforms.
13. Establish Ongoing Governance, Annual Review, and Audit Readiness Cycle (depends on: 12)
Embed incident management as a permanent organizational capability, not a one-time project.
- **Annual process review**: the IMC reviews the severity taxonomy, role definitions, on-call structure, and compensation against industry benchmarks and internal KPIs. Update as needed.
- **Bi-annual tabletop exercises**: one SEV1 infrastructure scenario, one SEV1 data-breach/regulator scenario. Rotate the IC and CL assignments so everyone gets practice.
- **Quarterly alert-quality audit**: the IMO reviews alert volumes, noise rates, and runbook coverage across all 28 teams.
- **On-call health survey**: quarterly anonymous survey measuring burnout, fairness, and compensation satisfaction. Results reviewed by the IMC.
- **SOC 2 readiness cycle**: begin evidence collection immediately after each audit ends. The IMO maintains a rolling evidence folder. Mock audit at month 5 of every 12-month audit cycle.
- **Postmortem maturity tracking**: track the action-item completion rate monthly. If it drops below 80%, the VP Engineering intervenes.
- **Incident management maturity model**: adopt a 5-level maturity model (ad-hoc → defined → managed → optimized → predictive). Assess annually. Target: Level 3 within 12 months, Level 4 within 24 months.
- **Budget review**: annually review on-call compensation, tooling costs, and training budget against the reduction in SLA credits and incident frequency.
--- PROPOSAL 4 (agent grok4.6_initial_4, xai/grok-4.6) ---
Estimated complexity: high
Success metrics: - Median time to detect customer-impacting incidents ≤ 5 minutes within 6 months of go-live.
- Share of SEV-1/SEV-2 incidents first detected by customers ≤ 5% (from 40%).
- Median time to mitigate SEV-1/SEV-2 ≤ 45 minutes (from 3 h 10 min).
- Named Incident Commander assigned within 5 minutes for ≥ 95% of SEV-1/SEV-2.
- First status-page update within policy time for ≥ 95% of SEV-1/SEV-2.
- SLA credits down ≥ 80% versus the trailing $1.3M within 12 months.
- Paging volume ≤ 500 per month and noise ≤ 15% (from 3,400 and 85%).
- 100% of production services have a named owning team and a paging policy.
- Postmortems filed within 5 business days for 100% of SEV-1/SEV-2; action-item close rate ≥ 80% within 30 days.
- 24×7 IC and critical-path coverage with zero unfilled shifts per quarter.
- Paid on-call live for every rotation before that rotation pages humans.
- SOC 2 Type II incident-response controls evidenced for ≥ 5 months before the auditor's report.
- On-call pulse: ≥ 70% of engineers agree rotations are fair and limited to their services.
Steps (30):
1. Secure executive mandate and budget
Get a written CEO/CTO mandate that incident command is a company process, not a team hobby.
The mandate must state that **paid on-call** is required for production ownership. "No pager for other teams' code" is solved by named ownership, not by refusing coverage.
- Approve budget for tooling, stipends, training, and a dedicated program lead for six months.
- Name an executive sponsor (CTO or VP Engineering) who will chair the weekly incident review.
- Tie the clock to SOC 2 Type II: the process must be live in about 10 weeks so ~6 months of evidence remain.
- Commit that the CEO will hear about outages from this process, not from customers.
2. Form the working group and decision rights (depends on: 1)
Stand up a small group that can decide. Do not form a 28-team committee.
**Core seats:** SRE/platform lead, payments/ledger engineering manager, support lead, legal/compliance, HR, one rotating team EM, and a program manager.
- Meet twice a week for 10 weeks, then weekly.
- RACI: the group proposes; the sponsor decides in 48 hours; teams implement.
- Publish one Slack channel and one source-of-truth doc on day one.
- Time-box design to four weeks. Ship v1 rather than wait for consensus.
3. Inventory services, owners, and on-call gaps (depends on: 1)
Build a living catalog of all ~180 services: owning team, criticality, current on-call, alert sources, and runbook link.
Walk the last 31 customer-impacting incidents and the two events where **nobody was in charge**. Record who detected, who led, time to mitigate, and which alerts fired.
- Tag each service as critical-path, customer-visible, or internal.
- List the 16 teams with no on-call and every orphan service with no owner.
- Map all six alerting tools and the 3,400 monthly alerts onto services.
- Flag the shared PostgreSQL ledger and two-region failover as named special cases.
4. Map regulatory and contractual notification duties (depends on: 1)
Legal and compliance list every duty an incident can trigger. Do not invent clocks that violate a contract.
Cover SOC 2 CC7, customer MSA/SLA credit terms, money-transmitter rules, NYDFS 23 NYCRR 500 if applicable, PCI if in scope, and breach clocks.
- Extract **notification timings** from the largest customer contracts (status page, named AM, written notice).
- Define when legal, regulators, insurers, or the board must be told.
- Feed these clocks into severity triggers and the communications playbook.
5. Approve paid on-call and incident pay (depends on: 1, 4)
Unpaid on-call is why 16 teams refuse the pager and why nights are uncovered. Fix the money before asking for coverage.
HR, legal, and finance design a New York–compliant package: weekly stipend for primary and secondary, extra stipend for company IC and comms, and **after-hours incident pay** or comp time.
- Treat exempt vs non-exempt staff explicitly under NY wage-hour rules.
- Put stipend in the next pay cycle after policy publish, not "later."
- Cap consecutive night weeks. Fund hiring if a team cannot rotate fairly (minimum six people for 24x7 primary plus secondary).
- Publish the package before any new rotation starts. This is the main answer to pager pushback.
6. Ratify four severity levels and their triggers (depends on: 3, 4)
Adopt a business-impact scale. Engineers do not invent severity in the moment.
**SEV-1:** material payments failure; ledger down or inconsistent; security or customer-data incident; both regions impaired; or many customers already in SLA-credit territory.
**SEV-2:** degraded payments or a contracted feature down for multiple customers; SLA at risk.
**SEV-3:** narrow or single-customer impact with a workaround; no fleet-wide SLA risk.
**SEV-4:** no customer impact; ticket only.
- SEV-1 pages IC, comms, scribe, owning SMEs, and an exec; war room in 5 minutes; status page in 10; AM outreach in 20.
- SEV-2 pages IC and owning SMEs; comms may be the IC; status page in 15 minutes; updates every 30 minutes.
- SEV-3 pages the owning team only; customer notice only if that customer is affected.
- Anyone may declare. Only the IC may downgrade. When unsure, start high.
7. Define incident roles and their authority (depends on: 6)
Four roles. Separate coordination from debugging so "who is in charge" cannot stall for an hour again.
- **Incident Commander:** owns severity, the room, the clock, and the next action. Does not write code. May page anyone, freeze deploys, and invoke failover. Staffed from a company-wide trained pool, not from the failing team.
- **Communications Lead:** status page, customers, AMs, execs, regulators. Speaks only from IC-approved facts.
- **Scribe:** timeline in the incident tool. Required for SEV-1 and SEV-2.
- **SME responders:** the owning team's on-call. They mitigate. They do not run the room.
Publish a one-page authority card. The IC stays in charge even if a VP joins.
8. Design 24x7 coverage without 28 night rotations (depends on: 3, 5, 7)
Do not put 28 teams on 24x7. That is what engineers are rejecting.
Use **three layers** so people page for their own code, plus a trained commander.
- Layer A — company IC and SEV-1 comms: 24x7; about 24 trained people; week-long primary and secondary.
- Layer B — critical-path team on-call (ledger, payments processing, auth, API edge, platform/Kubernetes, data stores): 24x7 primary plus secondary.
- Layer C — all other teams: business-hours on-call; after hours the IC pages the team EM, who has a written escalation list.
Platform on-call is the safety net for unknown-owner pages, never the permanent owner. Every service must have a named team within 60 days or be scheduled to shut off.
9. Set rotation, handoff, and load rules (depends on: 8)
Write mechanical rules so rotations are fair and load is visible.
Primary week, then secondary week, then at least two weeks off. No one holds primary on two rotations at once.
- Handoff is a 30-minute overlap covering open incidents, silenced alerts, and upcoming changes.
- Page-load SLO: p50 ≤ 4 pages per 12-hour night shift; p95 ≤ 10. A breach opens an alert-quality action.
- Require a shadow week before a first IC shift or a first critical-path rotation.
- Swaps live in the paging tool. Managers own coverage gaps, not the last person on the roster.
10. Write detection and escalation paths (depends on: 6, 9)
Customers currently detect 40% of incidents and median time to detect is 22 minutes. That is the first failure mode.
Detection path: synthetic full-payment probes in both regions, SLO burn-rate alerts, support-to-incident intake, and one customer callback path that can create a SEV.
- A page must be acked in **5 minutes** or it auto-escalates to secondary, then IC, then the EM, then the VP.
- Support may declare SEV-2 or higher without engineering permission.
- If ownership is unclear for 10 minutes, the IC keeps the incident and assigns a temporary owner. Never wait.
- An exec bridge auto-opens for every SEV-1 at T+15 minutes.
11. Write internal, customer, and regulator communications (depends on: 4, 6, 7)
Stop "whoever is around" from writing the status page. Comms follow the clock, not convenience.
Timings from declaration:
- Internal war room: immediate. Exec summary for SEV-1/2 at 15 minutes, then every 30 minutes.
- **Public status page:** SEV-1 in 10 minutes, SEV-2 in 15. Updates at least every 30 minutes until resolve. Templates only. No speculation.
- Account managers get an affected-customer list and a script at T+20 minutes for SEV-1/2.
- Resolve notice and credit assessment within one business day.
- Comms pages legal on SEV-1 security, ledger integrity, or any outage that will breach contractual notice. Legal owns outbound regulatory letters; the IC owns facts.
12. Standardize blameless postmortems and action tracking (depends on: 6)
A written postmortem is mandatory for every SEV-1 and SEV-2 within 5 business days. SEV-3 if the IC or EM requests it.
Use one template: timeline, customer impact (volume, duration, credits), detection gap, what went well, what did not, process-focused five whys, and numbered actions with owner and due date.
- Review is **blameless** and scheduled. The IC attends. The exec sponsor reads every SEV-1.
- Actions live in one tracker, not in the doc. No action without an owner and a date. Default due date 14 days; 30 days max unless architecture work with a milestone.
- Close rate is a published metric. The old 11-of-64 pattern is a process failure.
13. Set alert quality rules that make paging acceptable (depends on: 3, 6)
3,400 alerts a month and 85% noise is why on-call feels like punishment. Pages are a product with a quality bar.
A page (not a ticket) must map to a customer-facing SLO or a hard dependency of one. It must have an owner team, a runbook link, and a default severity. It must be actionable at 3 a.m. by the person who is paged.
- Ban parallel paging from six tools. One paging policy: symptom-based; burn-rate preferred over raw thresholds.
- Every team gets a monthly noise budget. Exceeding it is a sprint task, not heroics.
- A human may silence a flapping alert only with a linked ticket.
14. Publish Incident Management Policy v1 (depends on: 5, 6, 7, 8, 9, 10, 11, 12, 13)
Collapse the design into a short policy people will open during an outage.
Ten pages or fewer, plus one-page cards for severity, roles, and comms timings. Host it where the incident tool can link it.
- Include the compensation summary and the rule: you are **not on-call for other teams' services**.
- Version it. v1 is mandatory from the pilot start date.
- Legal, HR, and the exec sponsor sign. Announce in all-hands, not only in Slack.
15. Implement a single incident command tool (depends on: 7, 10, 11)
Put one tool in the path that creates the room, pages roles from severity, records the timeline, and prompts status-page updates.
Requirements: Slack (or equivalent) incident bot, severity in one click, role assignment, stakeholder groups, and timeline export for postmortems and auditors.
- Integrate with the pager so IC, comms, and SME pages are automatic.
- Retain artifacts at least one year for SOC 2.
- Ad-hoc Zoom/Slack threads are no longer the system of record.
16. Consolidate six alerting tools onto one pager (depends on: 9, 13)
Pick one paging product. Connect existing monitors to it. Migrate **pages** first, tickets second.
- Inventory every page-producing rule. Delete or downgrade the noisy majority in S23.
- Route by service label → owning team schedule → the escalation policy from S10.
- IC and comms schedules live in the same product.
- Set a hard date after which pages outside the chosen tool are not valid on-call obligations.
17. Operationalize the status page and AM path (depends on: 11, 15)
Put the status page behind the Comms Lead role. Use templates for investigating, identified, mitigating, and resolved.
Subscribe AMs and customers whose contracts require it. Generate the affected-customer list from the incident (tenant, region, payment method).
- Dry-run a SEV-2 update before the pilot goes live.
- Record every public update in the incident timeline for the audit.
- Define partial vs full-outage wording so impact cannot be understated.
18. Stand up one postmortem repo and action board (depends on: 12)
Create the template, the filing location, and one Jira/Linear board with states: open, in progress, blocked, done, and won't-do (with exec reason).
Wire the incident tool so a SEV-1/2 automatically opens a draft postmortem and action tickets.
- Each week the program manager reports actions past due to the exec sponsor.
- If it is not in the board, it does not exist.
19. Assign owners and write critical-path runbooks (depends on: 3, 9)
Close the ownership gaps that cause "pager for other people's code."
Every production service gets a team in the catalog. Unowned services get an owner in 30 days or a decommission date.
- Write runbooks for ledger Postgres, regional failover, payments API, auth, and the Kubernetes control plane: symptoms, dashboards, mitigate vs escalate, customer impact.
- Link runbooks from alerts. If there is no runbook, the alert cannot page at night unless the EM accepts the gap in writing.
20. Detect payments failures before customers do (depends on: 3, 13)
Build or finish end-to-end synthetics: create payment, ledger write, webhook, both regions, both critical payment methods.
Alert on SLO burn, not on a single 500. Page SEV-2 or SEV-1 from these probes. This is the fastest lever on the 22-minute MTTD and the 40% customer-detected rate.
- Add ledger lag, replication, disk, and failover-readiness as first-class pages to the ledger team.
- Every postmortem asks first: **why did a customer see this first?**
21. Train the first cadre of ICs, comms leads, and scribes (depends on: 7, 14, 15)
Train about 24 ICs and 12 comms leads before the pilot. Classroom plus a recorded shadow of a simulated SEV-1.
Curriculum: severity, authority card, tool, comms timings, when to call legal, how to run a room of 20, how to hand off at 2 a.m.
- Certification: pass a tabletop. No certificate, no rotation.
- Recertify yearly and after any SEV-1 where process failed.
- Managers of ICs protect calendar time. This is part of the job.
22. Pilot on the payments critical path for six weeks (depends on: 5, 14, 15, 16, 17, 18, 19, 21)
Go live with policy, tool, paid rotations, and IC coverage for ledger, payments, API, platform, and support intake.
Keep old paths as backup for one week, then cut over. Real incidents use the new process only.
- Staff the program lead in every SEV-2+ as coach, not as secret IC.
- Collect friction daily. Fix tooling and wording in 48 hours.
- Expansion gate: named IC in under 5 minutes, first status update on time, no unpaid pages, postmortem filed.
23. Cut alert noise with a forced burn-down (depends on: 13, 16)
Give every team a numbered list of their noisiest alerts. Move noise from 85% to **under 15%**, and monthly pages from 3,400 toward 500.
Each sprint, critical-path teams must delete, debounce, or convert to ticket a fixed quota. Platform provides burn-rate and grouping libraries.
- Publish a weekly noise leaderboard. Shame systems, not people.
- After eight weeks, any page without a runbook or with >30% false pages in 14 days is auto-downgraded until fixed.
24. Roll remaining teams onto the model by risk (depends on: 22)
After the pilot gate, add teams in waves of four to five every two weeks. Highest customer-impact first.
Layer C teams get business-hours schedules and the EM night list. Do not surprise anyone with a pager.
- Each wave: ownership confirmed, alerts routed, runbooks for paging alerts, paid rotation in HR, one tabletop.
- Finish all 28 teams at least five months before the SOC 2 report date so the observation window covers the company.
- Orphan services still unowned at wave end are escalated to the sponsor for shutdown or reassignment.
25. Run tabletops and multi-region game days (depends on: 21, 22)
Schedule a monthly tabletop: SEV-1 ledger, SEV-1 region loss, SEV-2 degraded payments, customer-detected incident, and a "who is in charge" chaos drill.
Quarterly game day: fail a region or a ledger replica in staging or a controlled production drill.
- Include support, AMs, legal, and an exec. Process fails if only engineers show up.
- Capture actions on the same board as real postmortems.
- Use results as SOC 2 evidence that IR is tested.
26. Resolve ownership fights and pager culture (depends on: 5, 8, 22)
Treat pushback as design input, not defiance. Repeat the contract in office hours: you carry a pager for **your** services; IC is coordination; nights are paid; noise is a defect.
- EMs who cannot staff a fair rotation get headcount or have services reassigned. Do not run two-person 24x7.
- Publicly close the two historical "nobody in charge" incidents with what would be different now.
- Pulse-survey on-call at 60 and 120 days. If load or fairness is red, stop expansion until fixed.
27. Launch metrics, weekly review, and error budgets (depends on: 15, 22)
The CEO email exists because there was no operating rhythm. Stand up a dashboard and use it.
Track incident count by SEV, MTTD, MTTA, MTTM, customer-detected %, time-to-IC, time-to-status, pages per shift, noise %, SLA credits, postmortem SLA, and action close rate.
- **Weekly 30-minute review** chaired by the exec sponsor. Last week's incidents, open actions, alert SLO breaches. Mandatory for EMs of teams that had SEV-2+.
- Error budgets on critical SLOs: burn too fast and the team stops features and pays down reliability.
- Monthly report to the CEO so customer-first news does not recur.
28. Collect SOC 2 Type II evidence as a side effect (depends on: 4, 14, 22, 27)
Do not create a parallel "audit process." The real process is the evidence.
From pilot start, retain policy versions, on-call schedules, incident records, status-page history, postmortems, action tickets, training attendance, tabletop notes, and access reviews of the incident and pager tools.
- Map each artifact to SOC 2 CC7.2–CC7.5. Compliance owns the matrix; engineering owns completeness.
- Internal audit or a consultant runs readiness reviews at week 12 and week 20, not the week before the auditor.
29. Inspect and adapt the process after 90 days (depends on: 22, 27)
After 90 days live, revise v1 using data, not opinions.
Look at MTTD/MTTM, uncovered shifts, IC burnout, missed status updates, action close rate, and survey results.
- Change severity definitions if teams inflate or deflate them.
- Adjust Layer B vs C membership from real pages.
- Cut steps nobody follows. Add only what the last 90 days proved missing.
- Freeze v2 as the SOC 2 described process unless a defect is found.
30. Rehearse the audit and freeze the observation window (depends on: 24, 28, 29)
Six to eight weeks before the auditor, run a mock interview with an IC, a random engineer, support, and compliance.
Show one SEV-1 and one SEV-2 end-to-end: detect, page, command, comms, postmortem, closed actions.
- Fix evidence gaps immediately.
- Freeze process wording for the rest of the observation window; log exceptions.
- Brief the CEO and customer success with the metrics so the "outages we hear about from clients" story is retired.
--- PROPOSAL 5 (agent deepseek-v4-pro_initial_5, deepseek/deepseek-v4-pro) ---
Estimated complexity: high
Success metrics: - Customer-detected incidents decrease from 40% to less than 15% within six months.
- Median time to detect (MTTD) is under 5 minutes for SEV1 and SEV2 incidents.
- Median time to mitigate (MTTM) is under 60 minutes for SEV1 and under 2 hours for SEV2.
- Alert noise decreases from 85% to below 10% within three months.
- 100% of SEV1 and SEV2 incidents have a completed blameless postmortem within 5 business days.
- 100% of postmortem action items are tracked with an owner and due date; 90% are completed on time.
- 24x7 on-call coverage achieved across all 28 teams with no unpaid on-call.
- 99% of pages are acknowledged within 5 minutes.
- SLA credits paid reduce by at least 50% over the next 12 months.
- SOC 2 readiness: all incident response controls are documented, tested, and evidence is produced by month 7.
Steps (14):
1. Baseline current incident response and align stakeholders
Collect data from the last 12 months of incidents, all six alert tools, on-call practices, and team interviews. Identify gaps against the target incident management process and secure executive sponsorship.
- Gather incident timeline, detection source, mitigation time, customer impact, SLA credit, and postmortem status for all 31 incidents.
- Survey 28 teams on on-call burden, alert quality, and operational pain.
- Map current tools, escalation paths, and communication workflows.
- Create baseline metrics and a stakeholder map with executive sponsor and audit owner.
2. Define severity levels and response triggers (depends on: 1)
Define a four-level severity scale with objective business-impact criteria so any engineer can classify an incident consistently.
- SEV1: widespread transaction processing outage, data breach, security incident, or severe SLA breach; triggers full incident command, executive notification, and a 5-minute status page update.
- SEV2: major feature outage, significant degradation without workaround, or customer financial risk; triggers incident commander, full communications role, and status page updates.
- SEV3: partial impairment with workaround or limited customer impact; triggers on-call response, internal communication, and optional status page update.
- SEV4: minor or internal issue, no customer impact; handled during business hours through ticketing.
- Include an escalation matrix showing who can declare, downgrade, and invoke regulatory or legal involvement.
3. Define incident roles and decision authority (depends on: 1, 2)
Define incident roles, responsibilities, and decision authority using RACI to remove ambiguity about who is in charge.
- Incident Commander: owns the incident, declares severity, and coordinates resolution.
- Communications Lead: owns internal and external messaging, status page updates, and account manager notifications.
- Scribe: maintains timeline, incident log, and postmortem notes.
- Subject-matter responders: diagnose and fix the incident; may come from multiple teams.
- Executive sponsor: optional for SEV1; customer liaison: handles account managers.
- Define decision rights for severity declaration, escalation, rollback, customer communications, and incident closure.
4. Design 24x7 staffing model across 28 teams (depends on: 3)
Design 24x7 coverage across 28 teams without overloading engineers. Use service-based on-call plus a central incident command pool.
- Each service or domain team assigns primary and secondary on-call for its own services.
- Create central incident commander, communications, and scribe rotations staffed from a trained incident response guild across all teams; use follow-the-sun between the two AWS regions and time zones.
- Define escalation layers: service on-call to team lead or manager to service owner to executive.
- Define handoff times, shadow shifts, and load balancing; target at most one week of on-call per engineer per month.
- Bridge the current 12-team paid on-call to 28-team paid coverage; no team remains uncovered.
5. Define on-call rotations, compensation, and alert quality rules (depends on: 4)
Define sustainable rotations, pay, and rules that eliminate noisy pages.
- Rotations: weekly or biweekly, at least one primary and one secondary, with 12-hour shifts where possible or 24-hour for low-volume services.
- Compensation: monthly on-call stipend for all on-call engineers, additional incident response bonus for after-hours work, and time off in lieu; align with market rates.
- Alert quality rules: every page must be actionable, have a runbook link, specify a service owner, include severity, and be based on SLO burn or known failure signals; no dashboard-only alerts.
- Noise budget: reject or downgrade non-actionable alerts; all pages must go to on-call only after suppression and deduplication.
- Weekly alert review removes the top noisy alerts.
6. Design detection, escalation, and alert routing (depends on: 2, 4, 5)
Define how incidents are detected, routed, and escalated so nothing waits on a human to notice.
- Consolidate the six alert tools into one alerting and paging platform with routing by service, severity, and tags.
- Detection sources: infrastructure metrics, application synthetic transactions, log-based anomalies, business transaction SLI monitoring, and customer-reported issues through support or account managers.
- Routing: alert is paged to service on-call within 30 seconds; primary must acknowledge within 5 minutes; if no ack, page secondary then on-call manager.
- Escalation timeouts: unresolved SEV1 escalates to service owner at 15 minutes and to leadership at 30 minutes; any engineer can escalate to the incident commander.
- Define customer-reported incident intake and classification in the same tool.
7. Define internal and external communication protocols (depends on: 2, 3)
Define communication channels, templates, and timing for internal, customer, and regulator audiences.
- Internal: dedicated incident Slack channel, internal status page mirror, and war room bridge for SEV1; incident commander and communications lead own these channels.
- Status page: SEV1 post within 5 minutes, updates every 30 minutes or on material change, resolution within 60 minutes of mitigation; SEV2 post within 15 minutes, updates hourly; SEV3 optional.
- Account managers: SEV1 and SEV2 notify account managers within 15 minutes with an approved customer-facing description and expected impact.
- Regulators: legal or compliance determines notification for data breaches, security incidents, funds availability issues, or regulatory reportable events; criteria and timing follow legal and regulatory requirements; communications lead coordinates.
- Use pre-approved message templates and an approval chain; no ad-hoc wording.
8. Define postmortem policy and action tracking (depends on: 3)
Define mandatory blameless postmortems and action tracking.
- Mandatory for all SEV1 and SEV2 incidents, and any SEV3 that breaches SLA or is customer-detected.
- Format: impact, timeline, root causes, contributing factors, detection and response gaps, what worked well, and action items.
- Blameless: focus on system and process causes, not individual blame; use trained facilitators.
- Ownership: each action has an owner, due date, and tracking ID in a single backlog.
- Review postmortems at the weekly incident review; track action closure; expect 100% completion.
- Complete postmortems within 5 business days for SEV1 and SEV2 incidents.
9. Define metrics, dashboards, and review cadence (depends on: 2, 3, 8)
Define metrics and review cadence to measure process health.
- Metrics: MTTD, MTTM, customer detected percentage, alert noise percentage, on-call response time, on-call load, SLA credits paid, and postmortem action completion.
- Dashboards: real-time operational dashboard for on-call engineers and management.
- Weekly incident review: review all SEV1 and SEV2 incidents, action items, and noisy alerts.
- Monthly trends with leadership; quarterly review against SLOs and audit controls.
- Success thresholds: MTTD under 5 minutes, MTTM under 60 minutes for SEV1, customer detected under 15%, and alert noise under 10%.
10. Configure incident tooling and integrations (depends on: 5, 6, 7, 8, 9)
Implement and integrate the tools that automate the defined process.
- Aggregate alerts from the existing six tools into PagerDuty, Opsgenie, or a similar platform.
- Configure on-call schedules, escalation policies, and paging targeted at service owners.
- Integrate status page API for automated or one-click updates.
- Add Slack commands to declare incidents, start war rooms, assign roles, and post status updates.
- Integrate runbook and service catalog access; create postmortem templates in Jira or Notion with action item tracking.
- Ensure audit trails and role assignments are logged for SOC 2.
11. Pilot with 2-3 volunteer teams and iterate (depends on: 10)
Run a controlled pilot before full rollout to validate and refine the process.
- Select 2-3 volunteer teams with representative services and on-call patterns.
- Run the new severity, roles, on-call, alerting, and communication process for 2 weeks.
- Track metrics and gather feedback from on-call engineers, incident commanders, and communications leads.
- Iterate severity thresholds, alert rules, templates, and runbooks based on findings.
- Exit criteria: no SEV1 without a declared incident commander, alert noise below target, and positive on-call survey results.
12. Train and certify all 28 teams (depends on: 11)
Prepare all 28 teams so the process is not a paper exercise.
- Create role-specific training: all engineers, incident commanders, communications leads, scribes, subject-matter responders, and managers.
- Deliver interactive workshops and online modules.
- Run at least two game-day drills per quarter simulating SEV1 and SEV2 incidents, testing escalation, communication, and postmortem.
- Certify incident commanders and communications leads; only certified staff can hold those roles in rotations.
- Include scenario-specific training for account managers and legal or regulatory handling.
13. Full rollout and enforce process (depends on: 12)
Move every team onto the new process and retire old habits.
- Migrate all 28 teams into consolidated on-call schedules and alert routing.
- Retire old alerting tools or methods that bypass the process.
- Activate status page templates and postmortem action tracking globally.
- Monitor the first 30 days with a daily operational review and quick fixes.
- Communicate rollout status to leadership and all engineers; provide office hours and support.
14. SOC 2 readiness and continuous improvement (depends on: 13)
Map the process to SOC 2 and keep improving after go-live.
- Map each process component, including severity, roles, detection, communications, postmortems, and metrics, to SOC 2 control requirements.
- Produce runbooks, evidence of training, incident logs, postmortems, and action item tracking.
- Conduct an internal readiness assessment at month 6 and remediate gaps before the audit at month 8.
- Maintain weekly and monthly metrics reviews to drive continuous improvement.
- Refresh on-call compensation and alert quality rules based on operational data.
HOW THE ROUNDS WENT (from the round analyses):
Round 0: All five agents converge on the same skeleton — exec mandate, baseline of the 31 incidents, a 4–5 level severity scale, separated IC/comms/scribe/SME roles, paid on-call, one paging tool, synthetic payment probes, timed status-page updates, mandatory blameless postmortems with tracked actions, pilot-then-waves, SOC 2 evidence as a by-product — so they differ mainly in staffing design, pace, and operational depth. P1 and P4 are the most granular (29–30 steps), P2 is the most payments-literate and the only one with day-1 interim controls, P3 is mid-weight with some internal inconsistencies, and P5 is the thinnest and least dated.
shared: **Command separated from debugging**: every plan defines IC / comms lead / scribe / SMEs, forbids the IC from typing, allows anyone to declare, and lets only the IC downgrade (P1 S6, P2 S5, P3 S4, P4 S7, P5 S3).
shared: **Pay for on-call and page only for owned code** as the explicit answer to the pushback, with a service-ownership catalog and criticality tiers as the routing foundation (P1 S4/S9, P2 S3/S6, P3 S4/S5, P4 S5/S18, P5 S4/S5).
shared: **Alert-quality gate plus customer-journey detection**: every page needs owner, runbook, severity and SLO linkage; synthetic end-to-end payment probes from both regions; consolidation of the six tools into one pager (P1 S10/S11/S12, P2 S7/S8, P3 S6/S9, P4 S13/S16/S20, P5 S5/S6/S10).
shared: **Mandatory postmortems with enforced action tracking**: SEV1/SEV2 always, 3–5 day drafts, single template, named individual owners, due dates in one tracker, escalation of overdue items, and SOC 2 evidence produced from the live process rather than reconstructed.
differences: **24x7 staffing design** is the biggest split: P1 S7/S8 uses a 25–35-person Duty IC pool plus tiered team obligations; P2 S5 collapses 28 teams into 8–12 domain rotations of ≥6; P3 S4 adds an SRE Tier 2 that absorbs infra pages and floats a managed overnight triage vendor; P4 S8 puts Layer C teams on business hours only with an EM night list; P5 S4 invokes "follow-the-sun between the two AWS regions", which does not create time-zone coverage.
differences: **Pace and interim cover**: P2 S2 installs interim IC staffing, one declaration path and a provisional severity guide within 7 days and triages the 53 open actions immediately; P4 S1 targets live in ~10 weeks and S30 freezes the process before the observation window; P1 runs design 1–6, pilot 7–12, rollout 13–24; P3 ends rollout at week 24 with a month-5 mock audit; P5 gives almost no dates and pilots only at step 11 of 14.
differences: **Payments/ledger rigour**: P2 S10 guards against split-brain, duplication and replay, requires reconciliation and backlog processing before "resolved", and separates mitigation from resolution; P1 adds ledger double-entry assurance (S11), a regulator matrix incl. NYDFS/FinCEN/sponsor banks (S16) and an automated SLA-credit workflow (S17); P4 S4 extracts notification clocks from contracts before setting severity triggers; P5 S7 defers regulator timing entirely to legal with no clock.
differences: **How failure and resistance are handled**: P1 S28 is a standing risk register with pre-committed fallbacks (volunteer shortfall, comp not approved, pruning a real alert); P4 S26 stops wave expansion if the 60/120-day pulse is red; P2 S17 requires exception records with owners and expiry dates; P3 and P5 have no comparable contingency mechanism.
Proposal 1: 29-step program run by a Head of Reliability with exec steering, starting from a forensic re-coding of all 31 incidents and the six-tool alert estate. Builds a service catalog with Tier 0–3, a five-level severity scale with auto-escalation on any ledger-touching incident, a 25–35-person Duty IC rotation plus tiered team on-call, and an on-call readiness bar that blocks a service from paging anyone until runbooks exist. Unusually complete on money and law: stipend ranges, FLSA/NY review, a regulator playbook, and an automated SLA-credit workflow targeting $1.3M → <$400k. Ends with a month-6 dry-run audit over 15 sampled incidents and a year-two maturity roadmap.
Proposal 2: 19 dense steps that put interim controls live in the first 7 days — one declaration path, staffed interim commanders, and triage of the 53 open historical actions — before any tool consolidation. Uses SEV0–SEV3, domain rotations of ≥6 instead of 28 fragile ones, compensation and fatigue rules (one week in six, paid recovery day, self-declared unfitness) and shadow-mode alerts for 7 days before they can page. Strongest on payments correctness: dual-control preserved during ledger recovery, split-brain/duplication guards, reconciliation required before "resolved". Milestones are day-anchored (7/14/30/60/90/120/180) with a month-7 mock audit and metric-gaming checks.
Proposal 3: 13 steps run by a new Incident Management Office and a 28-manager Council, starting with a baseline plus peer benchmarking against Stripe/Adyen. Distinctive is the three-tier staffing where a 6-person SRE Tier 2 absorbs Kubernetes/PostgreSQL pages and a Ledger Duty Officer joins any ledger-touching incident, explicitly to defuse the cross-team pager objection. Compensation is fully costed ($500/$250 weekly, $75–150 per page, $300 SEV1 bonus, ~$420K/yr). Weak spots: a per-team budget of 100 alerts/month contradicts the <600/month target, "56 certified ICs and all 260 engineers trained in 14 weeks" is optimistic, and step 12 funnels all eleven prior steps into one rollout.
Proposal 4: 30 short, opinionated steps that time-box design to four weeks and aim to be live in ~10 weeks so five-plus months of SOC 2 evidence accumulate. Organises coverage as Layer A (24 trained company ICs, 24x7), Layer B (critical path 24x7), Layer C (business hours plus an EM escalation list), with mechanical rotation rules and a page-load SLO of p50 ≤4 per night shift. Maps contractual and regulatory notification clocks (S4) before setting severity triggers, runs a forced noise burn-down with a weekly leaderboard, and gates wave expansion on a red on-call pulse survey. Closes with a mock auditor interview and a deliberate wording freeze for the observation window.
Proposal 5: 14 conventional steps covering the full brief: SEV1–SEV4, RACI roles, service on-call plus a central IC guild, stipends and TOIL, one paging platform, comms timings, mandatory postmortems, metrics and a month-6 readiness assessment. Correct in outline but the thinnest on specifics — no dates until months 6–8, no wave plan, no page budget, no risk register, and no targeted answer to the "other teams' code" objection beyond a survey. Two realism problems: a 5-minute SEV1 status-page post, and a 24x7 model built on "follow-the-sun between the two AWS regions" for a single New York organisation. The pilot lands at step 11 with training and full rollout after it, leaving little room to iterate before the audit.
Round 1: All five plans converged on a common architecture — 7-day interim command, forensic baseline, service catalog with tiers, 4–5 level severity, IC/Comms/Scribe/SME roles, a central command pool plus 8–12 domain rotations, paid on-call before mandatory nights, page budgets, money-path detection, one pager, 3/5/10-day postmortems with tracked actions, critical-path pilot, gated waves, month-6/7 audit rehearsal. The real remaining spread is in severity thresholds, compensation concreteness, and how much program scaffolding (policy v1, risk register, 90-day revision) each keeps.
differences: **Severity design.** P2 (S5) alone gives quantitative declaration guardrails (>10% payment failures for 5 min = SEV1; 1–10% = SEV2) plus a SEV0 crisis tier and no time-based auto-escalation; P5 (S5) uses five levels with aggressive clocks (SEV2 unresolved >1h becomes SEV1); P1/P3/P4 (S6) use SEV1–SEV4 with ledger-touch and 2-hour auto-escalation.
differences: **Compensation and load caps.** P1 (S9) and P5 (S9) publish dollar figures (~$800–1,200/primary week, $150/night page); P4 (S9) budgets $600k–1M with structure only; P2 (S8) and P3 (S7) defer amounts to HR in 14 days. Caps diverge: one primary week in six (P1, P2, P4) versus one in four (P5 S8; P3 S7 says "four or six depending on staffing").
differences: **Program scaffolding.** P1 keeps a standalone Policy v1 with exception register (S24), risk register (S31) and a 90-day inspect-and-adapt (S32); P3 keeps a risk register (S26). P4 dropped its own Policy v1 and 90-day revision; P2 and P5 have no contingency/risk step at all.
differences: **Interim protection and density.** P1 (S2), P2 (S2), P4 (S2) and P5 (S1) stand up interim command in 48h–7 days; P3 only lists "interim controls week 1" in S1 with no step behind it. On density, P4 is tightest at 24 steps, P1 heaviest at 32 with overlapping alert work (S10 standard vs S28 burn-down) and an 11-dependency join at S24.
influences: P2's 7-day interim minimum controls, retroactive pay for interim duty, triage of the 53 open actions and the daily 15-minute ops review (R0 S2) were taken by P1 (S2), P4 (S2) and P5 (S1); P3 left it as a single timeline bullet.
influences: P4's three-layer coverage and the "deal" fairness contract (R0 S8, S26) were taken by P1 (S8, S29), P3 (S8, S24) and P5 (S7–S8); P2 reached the same result via 8–12 responder domains (S7) without the layer naming.
influences: P1's forensic baseline, page budget, SLA-credit workflow, Incident Review Board and gated waves (R0 S2, S10, S17, S18, S25) were adopted almost wholesale by P3 (S2, S9, S14, S16, S22) and selectively by P4 (S3, S10, S16, S17, S22) and P2 (S3, S11, S15, S18, S21).
influences: P2's safety rails travelled widely: shadow-mode alerts and "never silently disable", two weeks of verified operation before retiring a legacy path, human acknowledgement required, mitigation-vs-resolution semantics and ledger dual-control (R0 S1, S4, S7, S8, S9) now appear in P1 (S6, S7, S10, S12), P4 (S7, S10, S12, S14), P3 (S9, S11) and P5 (S13).
influences: Nobody took P3's outsourced overnight triage vendor or its per-page pay table (R0 S4, S5) — P3 dropped both itself — and nobody kept P5's 5-minute SEV1 status-page deadline (R0 S7); all five settled on 15 minutes for the top availability tier.
Proposal 1 (improved): Closed its biggest R0 gap — six weeks of design with no protection — by adding a 7-day interim command bridge (S2). Replaced "14–18 teams at 24x7" with a Layer A/B/C model over 10–12 domains (S8), added live-execution doctrine (S14), Policy v1 (S24), a noise burn-down campaign (S28) and staged day-90/month-6 metric targets. Now the most complete plan, but at 32 steps it is also the heaviest.
Proposal 2 (improved): Grew from 19 to 24 steps by adding the pieces it was missing — readiness/runbook standard (S9), detection uplift (S12), pilot (S20), waves (S21), exercises (S22) — while keeping its distinctive strengths: compliance designed on day one (S4, before severity in S5), a balanced anti-gaming scorecard (S18), and now numeric severity guardrails.
Proposal 3 (improved): Doubled from 13 to 27 steps and adopted most of P1's R0 architecture: baseline, listening tour, layered staffing, escalation ladder, regulatory playbook, credits, runbooks, exercises, gated pilot and waves, risk register, SOC 2 dry-run. Far more complete and better sequenced, but largely derivative and it lost its own concrete compensation numbers.
Proposal 4 (improved): Consolidated 30 loose steps into 24 denser ones (postmortems and actions merged into S17, runbooks and readiness folded into S14, all comms into S15) and imported the financial and training machinery it was missing. Remains the most readable plan, but it dropped two of its own best R0 steps.
Proposal 5 (improved): Expanded from 14 thin steps to 26, closing nearly every R0 gap: interim duty officer, forensic baseline, listening tour, escalation ladder, regulatory playbook, credits, runbooks, exercises, pilot, waves and SOC 2 dry-run. The gains come almost entirely from copying P1's R0 structure and P2's SEV0, so it is now complete but the least differentiated plan.
Round 2: All five plans converged onto a near-common template: four-level severity, a central command corps plus 10–12 domain rotations plus business-hours Layer C, paid on-call gated before any mandatory night rotation, a two-page-per-week budget, one paging platform, 3/5/10-day blameless postmortems, wave rollout with gates and a month-7 mock audit. Real differentiation narrowed to severity anchoring (P4's integrity flag, P1's payment anchors), plan weight and audit sequencing, while P3 essentially re-issued P1's round-1 plan.
differences: **Severity anchoring.** P4 step 6 attaches a financial-integrity/security flag to any level (SEV0's forcing function without a fifth level); P1 step 6 adds payments anchors (value at risk per hour, settlement-deadline proximity, named accounts); P2 step 6 demotes percentages to guardrails; P3 step 7 and P5 step 6 stay purely qualitative. P1/P4 also tighten the SEV3 postmortem trigger (customer-detected, >2h, repeat), while P3/P5 keep "postmortem on request".
differences: **Audit sequencing.** P4 confirms the observation window in week 1 (step 1), P2 in step 3, P3 in step 5; P1 leaves it in step 30, which depends on steps 24/26/28, and cuts independent testing from two rounds to one at month 5. P2/P3/P5 keep tests in months 4 and 6.
differences: **Coverage economics.** P1 alone funds an overnight Duty Triage Desk owning the first ten minutes (step 8) but never sizes or budgets it; P4 alone shows the arithmetic (260 engineers sustain 10–12 rotations, not 28, step 8); P2 alone refuses to merge small teams for schedule convenience, scopes the six-responder rule to direct 24x7 rotations (steps 8, 22) and keeps a surge roster for concurrent incidents.
differences: **Weight and pace.** P1/P3 run 32 steps with standalone noise campaign, Policy v1 and risk register; P4 compresses to 26 with an 11-bullet comms mega-step (15) and the risk register demoted to last (26); P5 merges regulatory+credits (16) and postmortems+actions (17) and has no policy-publication step; P2 has 26 steps, no risk register at all, and the fastest schedule (Tier 0/1 by week 10, all teams by week 18 versus week 24 elsewhere).
influences: P2 and P5 abandoned their SEV0 tiers for the SEV1–SEV4 scale of P1/P3/P4 (P2 step 6, P5 step 6); P4 instead converted SEV0 into a financial-integrity/security flag (step 6) and P1 into payments anchors (step 6) — the round's one genuine design argument, resolved three different ways.
influences: P3 imported P1's round-1 plan almost wholesale (21 matching steps: Layer A/B/C 9, noise burn-down 26, Policy v1 24, risk register 31, 90-day inspect 32) and grafted on P2's day-one compliance step (P3 step 5 from P2 step 4) and P2's fake-redundancy check (P3 step 12 from P2 step 12).
influences: P5 took P1's Layer A/B/C coverage (step 8), readiness bar and ledger blast-radius workstream (step 10), risk register (step 25) and culture step (23), and tightened its own 1-in-4 rotation cap to P1's 1-in-6 (step 9).
influences: P2 took the standalone listening tour/fairness contract from P1 step 4, P3 step 3 and P4 step 4 (now P2 step 4) and concrete pay bands from P1/P5 step 9; in the other direction P4 took P2's week-1 observation-window confirmation (P4 step 1) and P1 added P2's 99.95% availability metric and an Internal Audit seat (step 1).
influences: Nobody took P5's named vendor shortlist (PagerDuty/incident.io/FireHydrant, round-1 step 12) — P5 dropped it itself; nobody adopted P2's numeric declaration thresholds (10% / 1–10% failure rates), which P2 also demoted; and only P2 (step 8) still plans for simultaneous incidents with a surge roster.
Proposal 1 (improved): Structure is unchanged at 32 steps; the gains are inside steps. Step 8 adds a paid overnight Duty Triage Desk, step 6 adds payments-specific severity anchors, step 11 adds an incident-replay test with "replay coverage" as a leading metric. Independent control testing shrank from two rounds to one.
Proposal 2 (improved): Renamed the scale to SEV1–SEV4, ending the SEV0/SEV1 ambiguity, and split seven previously buried topics (fairness contract, alert contract, customer comms, actions, policy, metrics, sustainability) into their own steps. Compensation moved from "fixed weekly stipends" to actual dollar bands and an annual budget.
Proposal 3 (improved): Effectively re-issues P1's round-1 plan: 21 of its 32 steps match P1 and only 5 match its own previous version. It gains the machinery it lacked (Policy v1, noise campaign, risk register, 90-day inspect) plus a day-one compliance step from P2, but contributes nothing new and leaves two internal inconsistencies.
Proposal 4 (improved): Compressed to 26 steps while adding the round's neatest severity idea and the only staffing arithmetic. Gains: integrity/security flag (step 6), 260-engineer rotation math (step 8), week-1 audit clock (step 1), postmortems split from action tracking. Losses: the change-management step disappeared and the risk register is demoted to the end.
Proposal 5 (improved): Dropped SEV0 for a four-level scale, adopted P1's Layer A/B/C model, metric set and risk register, and added the readiness bar, simulations and a culture step. Still 26 steps, but two of them are now overloaded and the policy-publication step is missing.
Round 3: Round 3 converged hard: all five plans now run the same skeleton (charter → 7-day floor → forensic baseline → fairness contract → catalog/tiers → severity → roles → three-layer 24x7 → paid on-call → alert budget → detection → single paging platform → escalation → comms → regulatory/credits → postmortems/actions → training → exercises → policy → pilot → gated waves → metrics → SOC 2 → 90-day inspect), with near-identical numbers ($1,000 primary week, 2 out-of-hours pages/responder/week, 15% reserved capacity, 30 ICs / 18 comms leads). Differentiation now comes from a few genuine additions — P1's change intelligence and Support-as-detection tier, P4's Duty Triage Desk carve-out and vendor-incident class, P2's defined noise/actionable metrics and severity modifiers — while P3 and P5 mostly merge material authored by P1 and P4.
differences: **Unique content**: only P1 has change intelligence/deployment safety (step 13), Support and account managers as an instrumented detection tier (step 20) and a concurrency doctrine with a Multi-Incident Coordinator (step 8); only P4 has a vendor-incident class for processor/bank/cloud failures (step 13) and a rule for who writes the 3 a.m. status page (step 8); only P2 defines "actionable page" and "noise" and uses five modifiers instead of flags (steps 12, 6). P3 and P5 contribute nothing the others do not have.
differences: **Overnight first line**: P1 (step 9), P4 (step 8) and P5 (step 8) fund a paid Duty Triage Desk; P4 adds the safety carve-out that it never holds suspected ledger, payment-halt or security pages. P2 (step 8) and P3 (step 9) refuse a triage desk and route unknown-owner pages to Duty Command plus platform, P2 additionally reclassifying lower-tier services that can cause severe overnight harm.
differences: **Schedule**: P2 (step 1) compresses design to weeks 1–4, pilot 5–10, rollout by week 18, control tests in months 4 and 6; P1, P3, P4 and P5 keep design weeks 1–6, pilot 7–12, rollout 13–24 and a month-5 internal test. P2 buys a longer observation window at the cost of a 4-week design for 180 services.
differences: **Failure handling and metric honesty**: P1 (step 34), P3 (step 28) and P5 (step 29) keep a standing monthly risk register; P4 compresses contingencies into bullets in step 26; P2 has none. On claims, P2 targets "no unresolved material exception" and P1 warns credits may rise before they fall (step 22), while P3 and P5 still promise "SOC 2 passes with zero exceptions".
influences: P1's replay test (R2 step 11) went everywhere: P2 step 13 plus a replay-coverage metric, P3 step 12, P4 steps 3 and 11 (as a "replay catalog" with a gap owner), P5 step 11. It is now the field's standard proof that detection actually improved.
influences: P4's severity flag (R2 step 6) was taken by P1 (step 7, financial-integrity/security flag forcing dual control and the reportability checkpoint) and generalised by P2 into five modifiers — FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, VENDOR (step 6). P3 and P5 declined it and kept four levels plus auto-escalation.
influences: P1's Duty Triage Desk (R2 step 8) was adopted by P4 (step 8) and P5 (step 8); P4 improved it with the carve-out that ledger, payment-halt and security events page commander and domains in parallel, plus a half-day triage-desk training module (step 18). P2 and P3 explicitly did not.
influences: P4's "confirm the SOC 2 observation window in week 1" (R2 step 1) was taken by P3 (step 1) and P5 (step 3), and expanded by P1 into a full week-one audit, legal and evidence scoping step (step 4) covering privilege, legal hold and the obligation matrix.
influences: P4's staffing arithmetic (R2 step 8 — 260 engineers cannot sustain 28 night rotas, never a two-person rotation) was taken verbatim by P3 (step 9) and restated by P1 (step 9, ~100 of 260 engineers carry a night obligation at one week in six). Nobody took P2's R2 caution against merging small teams for schedule convenience — all five consolidate into 10–12 domains.
Proposal 1 (improved): Four substantive new steps and three hardening additions, all addressing gaps the field had left open. It is now the most complete plan, at the cost of being the longest (35 steps).
Proposal 2 (improved): It decomposed its round-2 mega-steps into discrete, individually gated steps and added the sharpest measurement definitions in the field. It remains the only plan without a risk register.
Proposal 3 (improved): A competent merge: it absorbed P4's audit clock and staffing math, P1's replay test and P2's accommodations, and split the readiness bar out of the execution doctrine. It adds nothing the other plans do not already have, and it created two new mega-steps.
Proposal 4 (improved): It imported P1's Duty Triage Desk and replay test and then improved both, and added two things nobody else has: a vendor-incident class and a rule for who writes the status page overnight. It stays the most compact serious plan at 26 steps.
Proposal 5 (improved): It closed two real holes from its round-2 version — there was no policy-publication step and no separated postmortem/action or internal/customer comms steps — but it did so by copying P1's round-2 plan almost step for step. It contributes no idea of its own and picks up none of P1's round-3 advances.
THE PROPOSALS OF THE FINAL ROUND (3):
--- PROPOSAL 1 (agent opus5_refine_1, anthropic/claude-opus-5) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6 and under 5% by month 12.
- Replay coverage: by day 90 a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95% internally and externally.
- Top-100 account outreach completed within 30 minutes of SEV1 in 95% of qualifying cases, using the approved briefing pack.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5 and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- Change-correlated incidents are identified within 10 minutes of declaration in 90% of cases, and the change-correlated share of incidents declines quarter over quarter.
- SLA credits fall to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, one concurrency drill and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and a tested read-only mode by month 6, with milestones reported to the board quarterly.
Steps (35):
1. Charter the programme: one owner, one mandate, funded, dated against the audit
Turn the CEO email into a chartered company programme with a single accountable owner and authority over all 28 teams. Incident response stops being a per-team preference and becomes a **company operating process**.
- Name the CTO as executive sponsor and a full-time Head of Reliability & Incident Management as accountable owner, supported by a programme office of three: programme lead, incident-platform engineer, reliability analyst.
- Form a decision group (Engineering, Platform/SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables on day one: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, named service ownership, and **paid on-call**.
- Publish the timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal control test month 5, mock audit month 7, SOC 2 fieldwork month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.3M/yr, training and exercises, plus 15% of engineering capacity reserved for reliability work.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence immediately.
- Publish a one-page interim severity card and a single declaration path: one chat command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim 24x7 Duty Incident Commander roster (primary plus backup) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document and one naming convention for every major incident, starting now.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is the design input, the executive narrative and the frozen "before" picture for the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the **detection backlog**.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Correlate incidents with deployments, config changes and feature flags to quantify how many started with a change we made.
- Quantify true cost beyond credits: failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost, churn risk on named accounts.
- Freeze the baselines in a signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
4. Audit, legal and evidence scoping in week one (depends on: 1)
Design evidence as a by-product of operating, never as a reconstruction before fieldwork. Settle the scope and the legal handling of incident records now, not in month seven.
- Confirm with the external auditor the Type II observation window, the incident population definition, and the evidence they will sample. Everything from the day-7 floor onwards must count.
- Map incident response to the Trust Services Criteria with Compliance: CC7.2–CC7.5 (monitoring, identification, response, recovery), CC2.2/CC2.3 (internal and external communication), CC5 (control activities), A1.2 (availability).
- Define the evidence set and where it is produced automatically: incident records, paging and acknowledgement logs, role assignments, status-page history, reportability decisions, postmortems, action closure with verification, training register, drill records, access reviews.
- Agree retention, confidentiality, legal hold and access rules. Decide with counsel which postmortem content is privileged and how privileged material is segregated **without making the ordinary postmortem secret**.
- Start the obligation matrix with Legal: NYDFS 23 NYCRR 500, state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network contractual windows, cyber-insurer notice.
- Begin monthly evidence sampling from month 1 so control operation is visible long before the audit.
5. Listening tour, resistance map and the written on-call deal (depends on: 1, 3)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written deal signed by the sponsor.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
- State the hard gate publicly: no new mandatory night rotation starts before compensation, training, runbooks and staffing rules are live.
6. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- Record for every service: one accountable team, engineering manager, chat channel, escalation policy, dashboards, runbook link, dependency list, regions and data stores.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record SLO, RTO, RPO, active-active or region-bound, failover method, and the dependencies that make nominal two-region redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
7. Severity standard, declaration rights and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning responders and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and responders paged, comms and scribe if customer-visible, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, credit-generating or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Attach a financial-integrity flag or a security flag to any severity. The flag forces dual control, Legal engagement and the reportability checkpoint without inventing a fifth level.
- Anchor on payments reality alongside error rates: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including a customer report.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
8. Roles, authority, concurrency and handover discipline (depends on: 7)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice to the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1 and SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, Finance and Vendor Management join on defined triggers rather than by invitation.
- Rules: command claimed within 5 minutes and stated in channel; distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- **Concurrency doctrine:** two simultaneous SEV1/SEV2 incidents activate the secondary commander and a surge roster; a designated Multi-Incident Coordinator arbitrates shared resources such as the ledger, the database platform and the deploy freeze.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
9. Three-layer 24x7 coverage: command corps, domain rotations, overnight triage desk (depends on: 5, 6, 8)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving roughly one primary week per person every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement and reconciliation, auth, API edge, Kubernetes platform, data and reporting, partner integrations). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Publish the staffing arithmetic: 10–12 domain rotations of six to eight people plus a 30-person command corps means roughly 100 of 260 engineers carry a night obligation, at about one week in six. Twenty-eight independent rotations would be unstaffable and is therefore rejected on the numbers.
- Teams too small for a fair rotation get headcount, service reassignment, or a time-limited executive exception. **Never a two-person 24x7 rotation.**
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: a human acknowledgement within 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
10. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 5, 9)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour, and a premium for the overnight triage desk.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per responder per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty since day 7 is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons without career penalty, and on-call load is published quarterly by team.
11. Alert quality contract and page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression can never hide an incident.
12. Detection uplift on the money path, validated by incident replay (depends on: 6, 11)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action and a missed-detection review.
13. Change intelligence and deployment safety on the critical path (depends on: 6, 12)
Most of these incidents start with something we changed. Make change the first hypothesis the tooling answers, and make changes safer to reverse.
- Stream every deployment, configuration change, feature-flag flip, schema migration and infrastructure change into the incident timeline with service labels and owners.
- Give the commander an automatic "what changed in the last 60 minutes on the affected journey" panel at declaration time.
- Require Tier 0/1 changes to be progressively delivered with a documented rollback that is tested, time-bounded and executable by the on-call responder without the author.
- Treat ledger schema migrations and settlement-affecting changes as a separate class: dual approval, rehearsed rollback, no deploys inside the settlement window.
- Enforce a change freeze during SEV1 and SEV2, lifted only by the commander and logged.
- Report change-correlated incidents monthly; a rising ratio is a signal to strengthen release safety, not to blame a team.
14. One pager, one incident record, one status page — migrated without a detection gap (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy paging path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access, MFA and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page delivered outside this platform creates no on-call obligation.
15. Escalation ladder and the five-minute command rule (depends on: 8, 9, 14)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts and support entitlements: AWS and database premium support, processors, sponsor banks, card networks.
16. Live execution doctrine and payments safety rules (depends on: 8, 14, 15)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed in advance rather than improvised.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and automatic reopening if impact recurs.
17. Service readiness bar, major-incident playbooks and ledger blast-radius reduction (depends on: 6, 9, 16)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human, and the shared ledger is the largest structural risk in the estate.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as a concentration risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers, stronger regional independence — with board-visible quarterly milestones.
- Enforcement: no readiness sign-off, no night paging alerts unless the manager accepts the gap in writing with an expiry date and a compensating control. Never respond to a gap by turning detection off.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
18. Internal communications protocol (depends on: 8, 14)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support, Sales and Security.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change. First internal brief within 10 minutes of SEV1 and 15 of SEV2.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
19. Customer communications, status page and account-manager outreach (depends on: 7, 18)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; monitoring notice after mitigation; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, monitoring, resolution.
- Component-level status page mapped to customer journeys rather than internal service names, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1 and 60 of SEV2, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state observed impact, affected capabilities, any workaround and the next update time; **never speculate on cause, recovery time, data integrity or blame**.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
- Record any legally required restriction or delay of public detail, its approver, and the alternative stakeholder plan.
20. Support and account management as a detection and intake tier (depends on: 7, 12, 19)
Customers detected 40% of incidents first, which means the front line already holds the signal. Turn Support and account managers into an instrumented detection channel rather than a bystander.
- Give Support explicit declaration rights, a one-page trigger card, and a macro that opens an incident candidate directly in the incident platform.
- Automate the clustering rule: five similar tickets or calls in ten minutes auto-creates a triage incident assigned to the duty commander.
- Route credible partner, processor and sponsor-bank notifications into the same declaration path within five minutes.
- Generate the affected-customer list automatically from journey telemetry and the incident record, and push it to Support, CS and the account-manager briefing.
- Train Support and account managers on approved language and prohibit independent technical explanations to customers.
- Measure and publish "signal was in Support before it was in monitoring" as a detection defect, and feed each instance into the detection backlog.
21. Regulatory, partner and legal notification playbook (depends on: 4, 7, 19)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Complete the obligation matrix started in scoping and have counsel validate the triggers, deadlines, channels and submitting authority for each obligation.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with facts considered, decision-maker, timestamp and reassessment trigger.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, card networks, outside counsel, cyber-insurer, critical vendors.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the operational facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly inside the exercise programme, including one full regulator-notification simulation per year.
22. SLA credit workflow, true cost model and incentive guardrails (depends on: 7, 19)
Tie incidents to money so severity, credits and investment decisions stay honest, and Finance stops learning about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component, including how partial degradation counts.
- Compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report shows **which reliability investments would have prevented which payouts**.
- Track failed payment volume, delayed value, reconciliation breaks and impacted customers alongside credits as the true incident cost.
- Guardrail against the perverse incentive: better detection will surface incidents that previously went unbilled, so credits may rise before they fall. Publish this expectation to the executive team in advance, and make it a written rule that no finance or commercial pressure may influence severity or declaration.
23. Blameless postmortem standard and Incident Review Board (depends on: 4, 7, 8)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger; any control failure.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning engineering director is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions where security, privacy or privileged content requires it.
24. Action ownership, reserved capacity and enforcement (depends on: 14, 23)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, accountable manager, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness through tests, telemetry, exercises or production evidence before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
25. Training and certification academy (depends on: 8, 15, 18, 23)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, the 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, contractual clocks, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days of joining a domain.
- Certification is valid 12 months and renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
26. Exercise programme: tabletops, game days, night drills and vendor rehearsals (depends on: 14, 17, 25)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander. First company-wide tabletop within 30 days.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times, plus a concurrency drill with two simultaneous incidents.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down. Rehearse joint escalation with AWS, a processor and a sponsor bank.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
27. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 15, 18, 19, 21, 23, 24)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, escalation and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix, Readiness Bar.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
28. Pilot on the payments critical path (depends on: 10, 12, 13, 14, 17, 25, 27)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement and reconciliation, PostgreSQL platform, Kubernetes platform, API edge, auth and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, action tracking, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The programme lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time in 95% of cases, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
29. Alert noise burn-down campaign (depends on: 11, 14, 28)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → under 1,500 pages a month by day 90, under 500 with noise below 15% by month 6.
30. Metrics, dashboards, review cadence and anti-gaming (depends on: 14, 23, 28)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time from first impact to detect, declare, commander, acknowledge, mitigate, resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors, change-correlated incident ratio.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
31. Wave rollout to all 28 teams with readiness gates (depends on: 28, 30)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the programme office for three weeks; the director signs the gate.
- **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard by team, tier and control gap.
32. Change management, fairness and pager culture (depends on: 5, 10, 28)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
33. SOC 2 evidence by design, internal control test and mock audit (depends on: 4, 27, 30, 31)
The real process is the evidence. Never build a parallel audit process, and never reconstruct records after the fact.
- Maintain the evidence set defined in scoping, produced automatically and indexed: versioned policies and exceptions, catalog records, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Sample incidents monthly from first signal through verified corrective action, deliberately including customer-reported events, downgraded incidents and missed timelines.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling end to end and reporting gaps to the sponsor.
- Run a formal mock audit in month 7 using the populations, evidence requests and interviews the auditor will use: a commander, a random engineer, Support, Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
34. Programme risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps and immediate staffing intervention.
- **A major SEV1 mid-rollout:** the programme lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
35. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 30, 31, 33)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, exercise calendar and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and quarterly board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 2 (agent gpt5.6-sol_refine_2, openai/gpt-5.6-sol) ---
Estimated complexity: high
Success metrics: - By day 7, every suspected major incident uses one record, one coordination channel, and a named Incident Commander within 10 minutes.
- After day 30, no SEV1 or SEV2 remains without named command for more than 10 minutes.
- By month 3, at least 95% of SEV1 and SEV2 incidents have a named commander within 5 minutes.
- By month 3, at least 95% of SEV1 and SEV2 technical pages are acknowledged by a human within 5 minutes.
- By day 30, 100% of Tier 0 services have an owner, primary and secondary escalation, dashboard, and runbook.
- By day 60, 100% of Tier 0 and Tier 1 customer journeys have compensated 24x7 command and appropriate technical coverage.
- By week 18, all 180 services have an owner, tier, escalation path, and tested coverage model.
- No new mandatory night rotation begins before compensation, training, access, runbooks, and minimum staffing are active.
- Every direct 24x7 technical rotation has at least six qualified responders or an approved, expiring executive exception.
- No responder is routinely primary more often than one week in six or assigned to two simultaneous primary rotations.
- Median impact-to-detection time falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to below 20% by day 90 and below 10% by month 6.
- By day 90, incident replay identifies a current detector for at least 90% of the 31 historical incidents, including its expected detection minute.
- Median time to mitigation falls from 190 minutes to under 90 minutes by month 4 and under 60 minutes by month 8.
- At least 95% of applicable SEV1 status notices are published within 15 minutes and SEV2 notices within 30 minutes by month 3.
- At least 95% of incidents meet the required internal and customer update cadence by month 3.
- Monthly human notification episodes fall from the current 3,400 alert events to no more than 1,500 by day 90 and 500 by month 6.
- Alert noise falls from 85% to below 30% by day 90 and below 15% by month 6, without reducing Tier 0 or Tier 1 replay coverage.
- Average after-hours load remains at or below two notification episodes per responder per week; every sustained breach receives a dated remediation plan.
- All six monitoring sources route human pages through the controlled paging platform by week 12, with direct legacy routes retired by week 18.
- All required postmortems are drafted within 3 business days, reviewed within 5, and published within 10 from month 3 onward.
- All 53 historical open actions are triaged within 30 days.
- At least 90% of high-priority corrective actions are completed by their approved due dates with effectiveness evidence by month 6.
- Repeat incidents involving an unaddressed known contributing factor decline by at least 50% within 8 months.
- Reportability is assessed and recorded within 1 hour for 100% of SEV1 and qualifying SEV2 incidents, including not-reportable decisions.
- At least two end-to-end cross-company exercises, including regional impairment and ledger recovery, are completed before the audit.
- Approved ledger RTO and RPO, controlled failover, read-only mode, and post-recovery reconciliation are exercised before month 6.
- Quarterly responder surveys reach at least 75% favorable responses for fairness, ownership boundaries, compensation, and sustainability by month 6.
- Monthly contracted availability meets or exceeds 99.95% by month 6 using the contractually authoritative measurement method.
- Annualized SLA credits decline by at least 50% within 12 months, with a stretch target below $400,000.
- The month-7 mock audit finds no unowned high-risk control gap and at least 95% of sampled incidents contain complete operating evidence.
- The SOC 2 Type II incident-response controls complete external testing without an unresolved material exception.
Steps (28):
1. Charter the program and fund immediate action
Make incident management a **company operating process** within 48 hours. Give one accountable leader authority to set standards across all 28 teams.
- Name the CTO as executive sponsor and a Head of Reliability or Incident Management as accountable owner.
- Assign a small permanent team: program lead, incident-platform engineer, and reliability analyst.
- Include Engineering, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit in a decision group. The sponsor resolves blocked decisions within 48 hours.
- Approve non-negotiables: one severity model, one human-paging path, one incident record, paid on-call, named service ownership, mandatory postmortems, and tracked corrective actions.
- Reserve 10%–15% of engineering capacity for detection, runbooks, resilience, and incident actions.
- Fund tooling, compensation, training, exercises, and program staff. Compare the cost with the existing $1.3M in annual SLA credits.
- Set the schedule: operating floor by day 7, design during weeks 1–4, pilot during weeks 5–10, rollout during weeks 11–18, control tests in months 4 and 6, and mock audit in month 7.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for the final policy or new tooling. Put a minimum viable process into operation immediately and retain evidence from the first day.
- Publish a one-page interim severity guide, declaration procedure, role card, and communications clock.
- Provide one monitored declaration route through chat, telephone, and the current paging environment.
- Create one channel, bridge, timeline, and incident identifier for every suspected major incident.
- Staff temporary primary and backup Duty Incident Commanders from experienced managers and engineers.
- Require a named commander within 10 minutes. The duty engineering director assumes command if the command page is unclaimed.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Stop uncompensated mandatory after-hours expansion. Pay interim duty under a temporary stipend, retroactive to program launch.
- Hold a 15-minute daily operational review until the permanent process is active.
3. Create the factual, contractual, and control baseline (depends on: 1)
Build one defensible baseline for process design, investment decisions, and SOC 2 testing. Preserve the original data so progress cannot be created by changing definitions.
- Reconstruct all 31 incidents from impact start through detection, declaration, command, mitigation, recovery, communications, credits, and corrective actions.
- Reconstruct the two incidents with unclear command minute by minute.
- Identify the missing signal for every customer-first detection.
- Inventory all six alert sources, 3,400 monthly alert events, duplicates, noisy rules, missing owners, and missing runbooks.
- Record current rotations, unpaid duty, overnight activations, schedule size, and uncovered services.
- Freeze the baseline: 22-minute median detection, 40% customer-first detection, 190-minute median mitigation, 85% noise, $1.3M in credits, and 11 of 64 actions closed.
- Inventory customer-specific availability definitions, notice periods, credit terms, sponsor-bank obligations, and other incident-related contracts.
- Confirm the required SOC 2 Type II observation period and evidence expectations with the auditor during week 1.
4. Publish the on-call fairness contract (depends on: 1)
Treat pager resistance as a legitimate design constraint. Adoption depends on a written agreement that separates command from technical ownership.
- Interview representatives from all 28 teams and from Support, Customer Success, Security, and Operations.
- Separate concerns about unpaid work, sleep loss, unfamiliar systems, noisy alerts, inadequate runbooks, and blame.
- Recruit respected engineers from the 12 existing rotations to co-design and pilot the process.
- Publish the core promise: responders are paid, paged only for systems they own or have formally accepted and trained to support, and assisted by a separate Incident Commander.
- State that Platform may temporarily triage unknown ownership but does not inherit another team's service.
- Provide confidential accommodations for health, disability, pregnancy, or caregiving constraints without career penalty.
- Measure baseline trust, fairness, fatigue, and alert confidence. Repeat at days 60 and 120, then quarterly.
5. Build the service catalog and customer-journey map (depends on: 3)
Make a machine-readable catalog the source of truth for routing, impact analysis, status-page components, and control evidence. Every production service must have one accountable owner.
- Record the owning team, manager, business capability, repository, channel, dashboard, runbook, escalation policy, dependencies, regions, and data stores for all 180 services.
- Map initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, webhooks, reporting, and onboarding to their dependencies.
- Classify Tier 0 as financial-integrity and essential money-movement systems, including the ledger and PostgreSQL platform.
- Classify Tier 1 as customer-facing systems that can create material contractual impact.
- Classify Tier 2 as deferrable internal or batch systems and Tier 3 as non-critical systems.
- Record SLO, RTO, RPO, regional recovery mode, contractual commitments, and critical vendors for Tier 0 and Tier 1.
- Keep separate but coordinated ownership for the ledger application and PostgreSQL platform.
- Give orphan services an owner or approved decommission date within 30 days. Treat unowned Tier 0 services as release-blocking executive risks.
6. Adopt the severity model and incident modifiers (depends on: 3, 5)
Classify incidents by credible customer, financial, security, regulatory, and contractual harm. Start at the higher plausible severity while scope or integrity remains unknown.
- **SEV1 — crisis:** incorrect, duplicated, unauthorized, or lost money movement; ledger integrity in doubt; material security or data exposure; broad loss of a core payment journey; both regions impaired; or processing must be stopped. Immediately page all response roles and the Executive Duty Officer. Open the bridge within 5 minutes, freeze unrelated changes, issue internal notice within 10 minutes, publish applicable customer status within 15 minutes, start legal assessment within 1 hour, and require a postmortem.
- **SEV2 — major:** material payment degradation; settlement deadline at risk; regional impairment with reduced resilience; a critical customer or material cohort unavailable; or an SLA breach is likely. Page command and technical roles immediately. Issue internal notice within 15 minutes, applicable customer status within 30 minutes, and require a postmortem.
- **SEV3 — limited:** narrow customer impact with a safe workaround and no credible integrity, security, regulatory, or material contractual risk. The owning team leads. Page only when immediate action can reduce harm.
- **SEV4 — operational event:** no current customer impact and no credible imminent harm. Create a ticket and handle during normal hours.
- Add FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, and VENDOR modifiers. These invoke specialist controls without distorting customer-impact severity.
- Use at least SEV2 posture for unknown impact lasting 15 minutes, credible ledger-integrity risk, or a cross-domain incident without clear ownership.
- Escalate a customer-visible SEV3 that remains unmitigated for two hours.
- Permit anyone to declare. Only the Incident Commander may downgrade, with evidence and rationale recorded.
- Publish a decision tree and examples based on the 31 historical incidents.
7. Define roles, authority, and handoffs (depends on: 6)
Separate coordination, communications, recordkeeping, and technical repair. One named person holds command continuously throughout every SEV1 and SEV2.
- **Incident Commander:** owns severity, objectives, priorities, role assignment, escalation, mitigation coordination, handoffs, and closure. The commander does not act as the primary technical operator.
- **Communications Lead:** owns internal broadcasts, status-page updates, account-manager briefs, executive updates, and coordination with Legal.
- **Scribe:** maintains the timestamped record of facts, hypotheses, decisions, commands, owners, role changes, and communications.
- **Subject-Matter Responders:** diagnose and mitigate only systems for which they have ownership, access, training, or a formally accepted support agreement.
- **Executive Duty Officer:** removes organizational obstacles and makes exceptional business decisions without displacing the commander.
- Add Security, Legal, Compliance, Finance, and Vendor Management on modifier-specific triggers.
- Require distinct commander, Communications Lead, scribe, and technical lead for SEV1. Communications and scribe may combine for the first 10 minutes of a bounded SEV2.
- Pre-authorize commanders to freeze changes, order rollback, disable features, shed load, reroute traffic, and invoke approved continuity procedures.
- Preserve dual control, privileged-access restrictions, reconciliation, and change evidence for all ledger operations.
- Announce every role assignment and handoff verbally and in writing, including the exact transfer time and unresolved risks.
8. Create sustainable 24x7 coverage (depends on: 4, 5, 7)
Use central command coverage and risk-based technical coverage instead of creating 28 fragile night rotations. Responders do not carry pagers for unfamiliar code.
- Create a 24x7 Incident Command corps of approximately 30 certified people with primary and backup schedules. Two schedules create about 104 weekly assignments a year, or roughly three to four weeks per person annually.
- Create a 24x7 Communications pool of 18–24 trained people from Support, Customer Operations, Engineering Operations, and management.
- Create a similarly sized scribe pool. The backup commander temporarily records the first minutes if a scribe has not joined.
- Maintain a 24x7 Executive Duty Officer schedule and specialist contact paths for Security and Legal.
- Group Tier 0 and Tier 1 systems into roughly 8–12 coherent response domains only where responders share training, access, runbooks, and explicit support acceptance.
- Staff every critical domain with primary and secondary responders and at least six qualified people. Target eight where overnight activation is frequent.
- Give Tier 2 and Tier 3 services business-hours ownership plus tested manager and director escalation.
- Reclassify any lower-tier service capable of causing severe overnight harm rather than hiding the risk behind manager callback.
- Route unknown-owner incidents to Duty Command and Platform temporarily. Record each occurrence as a catalog control defect.
- Prohibit simultaneous primary assignments and two-person 24x7 rotations.
9. Implement compensation and fatigue protections (depends on: 4, 8)
End unpaid on-call before expanding mandatory coverage. Use fixed duty compensation so responders are not rewarded for alert volume.
- Use planning bands of $900–$1,200 per Tier 0/1 primary week and $300–$500 per secondary week.
- Use planning bands of $1,000–$1,300 per Duty Commander week and $400–$700 for Communications or scribe primary duty.
- Pay holiday premiums. Compensate all legally compensable active and waiting time for non-exempt staff, including overtime where required.
- Have HR, Finance, Payroll, and employment counsel approve final bands, tax handling, FLSA classification, New York wage-hour treatment, and schedule constraints within 14 days.
- Provide a protected recovery day after a SEV1, more than two hours of overnight work, or another defined fatigue threshold.
- Reduce normal delivery commitments by about 15% during a primary week.
- Prohibit consecutive primary weeks, on-call during leave, hidden schedule swaps, and primary duty more often than one week in six.
- Allow responders to declare temporary fatigue-related unfitness without career penalty.
- Trigger staffing and alert-remediation review when a rotation averages more than two after-hours pages per responder per week for four weeks.
- Make active payroll setup, training, access, and readiness hard gates before any new mandatory night rotation starts.
10. Codify the live incident lifecycle (depends on: 6, 7)
Give every incident the same operational flow from first signal through verified recovery. The first objective is limiting customer and financial harm, not proving root cause.
- Use the states Detected, Declared, Triaged, Mitigating, Mitigated, Monitoring, Resolved, and Reviewed.
- Record impact start separately from detection. Use the earliest defensible evidence and revise it transparently when later facts emerge.
- Open with a standard command message: severity, known impact, assigned roles, immediate objective, workstreams, and next update time.
- Freeze unrelated production changes during SEV1 and normally during SEV2. Record every exception.
- Prefer reversible mitigation: rollback, feature disablement, isolation, load shedding, rate limiting, traffic shift, partner rerouting, or controlled processing suspension.
- Separate mitigation and diagnosis workstreams when staffing allows.
- Keep decisions in the shared incident record rather than direct messages.
- Require formal command and technical handoffs for shift changes, fatigue, or incidents exceeding four hours.
- For payment incidents, verify backlog handling, duplicate protection, customer state, settlement exposure, and ledger reconciliation before resolution.
- Require a severity-specific stability period and explicit handback to the owning team, Support, and Customer Success.
11. Establish one paging and incident system of record (depends on: 5, 7, 8)
Monitoring tools may remain specialized, but every human page and major-incident record must enter one controlled platform. This provides consistent routing and an audit trail.
- Select one enterprise paging and scheduling platform, one integrated incident record, and one hosted status page within two weeks.
- Ingest events from all six monitoring tools before disabling their direct human-paging routes.
- Route pages using the service catalog and deduplicate events belonging to the same symptom.
- Provide one declaration command that creates the record, channel, bridge, severity prompt, roles, and response clocks.
- Capture declarations, acknowledgements, escalations, role assignments, decisions, severity changes, communications, mitigation, resolution, and postmortem linkage automatically.
- Integrate ticketing, customer-success data, status publishing, and conference facilities.
- Apply MFA, role-based access, periodic access reviews, and tamper-evident history.
- Keep privileged legal or security material in restricted linked records rather than exposing it in the general timeline.
- Provide telephone, SMS, and offline procedures for loss of chat, identity, paging vendors, the status provider, or an AWS region.
- Retire each legacy paging route only after ownership review, end-to-end testing, and two weeks of verified operation.
12. Enforce the alert-quality contract (depends on: 3, 5)
Treat every human page as a production interface with an owner and a required action. Measure human notification episodes rather than raw monitoring events.
- Require every paging rule to identify the service, owner, customer or SLO risk, urgency, expected action, dashboard, runbook, deduplication key, and escalation policy.
- Define an actionable page as one that causes or materially informs a timely intervention or risk decision.
- Define noise as duplicate, non-urgent, unactionable, stale, test-generated, or incorrectly routed notification.
- Page on payment outcomes, error-budget burn, queue age against deadlines, and financial-integrity risk rather than raw CPU, memory, pod, or log thresholds.
- Run new rules in shadow mode for seven days unless a documented emergency exception applies.
- Review repeated no-action pages within two business days.
- Set a page budget of no more than two after-hours notification episodes per responder per week, measured over four weeks.
- Make a sustained budget breach trigger a tuning sprint and block additional non-emergency paging rules.
- Require compensating detection and central approval before suppressing a Tier 0 or Tier 1 rule.
- Never disable an existing critical detector solely because its metadata or runbook is incomplete. Track the gap with a dated remediation owner.
13. Detect payment failures before customers (depends on: 5, 12)
Move detection from infrastructure health to customer journeys and ledger truth. Validate coverage against actual historical failures.
- Define SLIs and internal SLOs for initiation, authorization, orchestration, ledger posting, settlement, reconciliation, refunds, APIs, webhooks, and reporting freshness.
- Set internal objectives with enough headroom to protect the contractual 99.95% availability commitment.
- Run external synthetic transactions through critical journeys at least once per minute from paths independent of the production platform.
- Test each region and expose dependencies that defeat nominal regional redundancy.
- Add continuous controls for ledger imbalance, duplicate identifiers, unexpected balances, replication lag, backup health, and failover readiness.
- Alert on queue age and delayed payment value relative to settlement deadlines.
- Add tenant and cohort anomaly detection for high-value customers and critical payment methods.
- Convert credible Support, account-manager, processor, bank, and network reports into incident candidates within five minutes.
- Replay all 31 historical incidents. Record which current detector would fire and at what minute.
- Treat customer-first detection as a mandatory missed-detection review with a tracked action.
14. Implement the detection and escalation ladder (depends on: 7, 8, 11)
Create one time-bound path from the first credible signal to named command and the correct technical owner. Device delivery does not count as acknowledgement.
- Converge automated alerts, engineer observations, support cases, account-manager reports, partner notices, and customer calls on the same declaration path.
- Page the Duty Incident Commander and owning critical-domain primary immediately for suspected SEV1 or SEV2.
- Escalate an unacknowledged technical page to secondary at 5 minutes, manager at 10 minutes, and director at 15 minutes.
- Escalate unclaimed command to the backup commander at 5 minutes. The Executive Duty Officer assumes temporary command at 10 minutes until a certified transfer occurs.
- Let the commander page any owning domain directly, with a 10-minute acknowledgement obligation.
- Keep command with the current commander when service ownership remains unclear. Assign a temporary technical lead and record the ownership gap.
- Maintain tested external escalation routes for AWS, database support, processors, sponsor banks, networks, and critical vendors.
- Test the full declaration, acknowledgement, fallback, conference, and status-publishing path weekly.
- Treat missed acknowledgement, failed routing, and unowned incidents as control failures requiring review.
15. Standardize internal communications (depends on: 7, 10, 11)
Give responders one working room and stakeholders one controlled source of truth. Executives must not interrupt the technical command path.
- Maintain one working channel and bridge plus a read-only broadcast channel for executives, Support, Customer Success, Sales, Security, and Operations.
- Issue the initial internal notice within 10 minutes for SEV1 and 15 minutes for SEV2.
- Update internal stakeholders every 30 minutes for SEV1 and every 60 minutes for SEV2, including when there is no material change.
- Use a fixed format: affected capability, customer symptoms, known scope, current mitigation, open risk, next update time, commander, and Communications Lead.
- Give Support and Customer Success an approved holding statement and preliminary affected-customer information within the same initial-notice window.
- Route executive questions through the Executive Duty Officer or Communications Lead.
- Record every material decision and outbound message in the incident timeline.
- Require the executive team to sign the communication behavior rules.
16. Standardize customer, account, and regulatory communications (depends on: 3, 6, 15)
Replace ad hoc writing with timed, role-owned, pre-approved communication. Communicate observed impact before root cause is known.
- Publish customer status within 15 minutes of a customer-visible SEV1 and within 30 minutes of a customer-visible SEV2.
- Update at least every 30 minutes for SEV1 and every 60 minutes for SEV2 until monitoring or resolution.
- Publish a monitoring notice after mitigation and a resolution notice within 30 minutes of verified recovery.
- Map public status components to customer capabilities rather than internal service names.
- Pre-approve templates for payment failure, degradation, settlement delay, regional impairment, third-party failure, integrity investigation, security restriction, monitoring, and resolution.
- State observed impact, affected capabilities, available workaround, and next update time. Do not speculate about cause, blame, integrity, scope, or recovery time.
- Have account managers contact affected strategic accounts within 30 minutes for SEV1 and 60 minutes for SEV2 using the approved briefing.
- Offer status subscriptions to all customers. Auto-enroll only where contracts, consent, and applicable communication rules permit.
- Encode customer-specific notice deadlines and channels in the customer record.
- Have Legal and Compliance maintain a counsel-validated matrix covering applicable NYDFS, breach, GLBA or FTC, PCI, money-transmitter, sponsor-bank, network, insurance, and customer obligations.
- Complete and record a reportability assessment within one hour for every SEV1 and every security, privacy, or integrity-related SEV2, including decisions of not reportable.
- Let Legal own regulatory text and submission while the Incident Commander owns operational facts. Record any legally required restriction of public detail and its alternative stakeholder plan.
17. Tie incidents to SLA credits and financial exposure (depends on: 3, 16)
Link the incident record to contractual and financial outcomes. Finance should not discover outages through later credit claims.
- Define the authoritative availability calculation for each contract and customer journey with Legal and Finance.
- Calculate affected customers and minutes from incident scope and journey telemetry.
- Produce a preliminary credit and contractual-exposure estimate within five business days of resolution.
- Record failed payment count, failed or delayed value, settlement exposure, reconciliation breaks, support effort, and engineering effort.
- Establish a documented approval path for proactive credits and claims-based credits.
- Attribute credits and financial harm to recurring failure families.
- Use the quarterly credit analysis to prioritize detection, resilience, and architectural investment.
18. Establish mandatory blameless postmortems (depends on: 6, 7)
Use one learning standard with fixed deadlines. Keep learning separate from disciplinary and misconduct processes.
- Require a postmortem for every SEV1 and SEV2.
- Also require one for customer-first detection, impact lasting more than two hours, SLA credits, contractual breach, repeated contributing factors, major control failures, and ledger-integrity near misses.
- Produce a factual draft within three business days, hold the review within five, and publish the approved version within 10.
- Make the owning engineering director accountable for completion. The commander owns response analysis, and the scribe supplies the timeline.
- Use one template covering impact, financial exposure, detection source, timeline, response performance, contributing conditions, control performance, recovery, and actions.
- Require explicit answers to why detection was not earlier and why mitigation took as long as it did.
- Use a trained facilitator who was not the commander or primary technical responder.
- Describe decisions using the context and information available at the time. Do not name an individual as the root cause.
- Keep HR, misconduct, and personnel matters in separate processes.
- Publish broadly useful findings internally while restricting security, privacy, personnel, and privileged content appropriately.
19. Make corrective actions enforceable risk commitments (depends on: 18)
An action is not complete when its ticket is closed. It is complete when the intended risk reduction is verified.
- Give every action one individual owner, accountable manager, priority, due date, expected outcome, verification method, and linked incident.
- Use default classes: containment within 7 days, corrective work within 30 days, and strategic work within 90 days with milestones.
- Prefer actions that remove hazards or reduce blast radius over vague actions such as retraining or adding monitoring.
- Put SEV1 recurrence-prevention work ahead of discretionary roadmap work unless an executive accepts the residual risk.
- Reserve 10%–15% of engineering capacity for approved reliability work.
- Escalate overdue high-risk actions to the manager after 7 days, director after 14, and CTO after 30.
- Require written residual-risk acceptance, compensating controls, and a new review date for deferrals.
- Permit unrepaired severe conditions to block related releases.
- Verify effectiveness using tests, telemetry, exercises, or production evidence before closure.
- Triage all 53 historical open actions within 30 days as complete, re-planned, superseded with evidence, or formally risk-accepted.
20. Set the service-readiness bar and payments playbooks (depends on: 5, 10, 12, 13)
A critical service must be supportable at 3 a.m. before it enters direct overnight coverage. Existing critical detection remains active while readiness gaps are repaired.
- Require current architecture and dependency diagrams, dashboards, SLOs, alert-to-runbook links, access, rollback, feature controls, escalation contacts, RTO, RPO, and integrity constraints.
- Require responders to demonstrate access and safe execution before independent primary duty.
- Create playbooks for PostgreSQL failure, suspected ledger corruption, duplicate payments, regional impairment, Kubernetes degradation, processor or bank outage, settlement risk, queue backlog, and credential compromise.
- Define stop-processing and read-only modes for credible financial-integrity risk.
- Document split-brain prevention, replay protection, failover, controlled backlog recovery, and post-recovery reconciliation.
- Require executive-approved RTO and RPO for the shared ledger cluster.
- Exercise critical runbooks at least twice a year and after material changes.
- Block new Tier 0 or Tier 1 releases and paging rules when readiness requirements are missing.
- Handle existing gaps using named owners, compensating controls, executive-approved expiry dates, and remediation plans.
- Run a parallel architecture workstream to reduce ledger blast radius, isolate non-critical readers, and strengthen regional independence.
21. Train and certify every response role (depends on: 7, 14, 15, 18, 20)
Command and communication are learned skills. Use paid working time and certify people before independent duty.
- Give all employees a 30-minute module on recognizing impact, declaring incidents, and locating the status page.
- Give responders a half-day module on severity, acknowledgement, escalation, evidence, runbooks, and financial-integrity precautions.
- Train scribes for two hours on timeline quality and fact-versus-hypothesis labeling.
- Train Incident Commanders for two days on delegation, uncertainty, severity, mitigation strategy, fatigue, handoffs, and executive management.
- Require commander candidates to complete a simulated SEV1 and two shadowed incidents or exercises.
- Train Communications Leads for one day on status writing, customer segmentation, legal boundaries, and contractual clocks.
- Require domain responders to demonstrate dashboards, access, rollback, failover, escalation, and relevant playbooks.
- Require two shadow shifts before independent primary duty.
- Renew certification annually through simulation.
- Maintain the training, assessment, and certification register as operational and audit evidence.
- Nominate an incident-management champion in each of the 28 teams.
22. Exercise command, recovery, and tool failure (depends on: 11, 20, 21)
Test the process before relying on it during a crisis. Exercise findings use the same action tracker and escalation rules as real incidents.
- Run the first cross-company command and communications tabletop within 30 days.
- Run monthly domain tabletops using scenarios from the 31 historical incidents.
- Run quarterly cross-company exercises for regional impairment, PostgreSQL recovery, settlement risk, partner failure, and simultaneous operational and security events.
- Test loss of chat, identity, paging, conference, and status-page providers.
- Include Support, Customer Success, account managers, executives, Security, Legal, Compliance, and critical vendors when relevant.
- Conduct at least one unannounced after-hours paging test before the audit and two annually thereafter.
- Validate backup restoration, RPO, RTO, financial controls, and reconciliation without uncontrolled experiments on the production ledger.
- Measure acknowledgement, command, customer notice, mitigation decision, handoff, and recovery times.
- Create tracked corrective actions for every material exercise finding.
23. Publish the signed policy and control set (depends on: 6, 7, 8, 9, 10, 12, 14, 15, 16, 18, 19)
Convert the design into concise documents that people can use during an incident. The actual operating process must also be the documented and audited process.
- Publish a short Incident Management Policy plus separate severity, on-call, communications, alert-quality, postmortem, evidence, and exception standards.
- Include one-page cards for severity, roles, authority, escalation, and communication timings inside the incident tool.
- State explicitly that responders support only owned or formally accepted and trained service portfolios.
- Include the compensation structure, fatigue rules, declaration rights, and non-retaliation commitment.
- Obtain approval from the CTO, HR, Legal, Security, Compliance, and Internal Audit.
- Announce the policy at an all-hands and through team briefings.
- Create an exception register with owner, rationale, compensating control, approver, review date, and expiry.
- Version every policy change. Do not rewrite historical records when the process changes.
24. Instrument the scorecard and review forums (depends on: 3, 11, 18, 19)
Measure the process before the pilot so failures become visible immediately. Report medians and 90th percentiles rather than averages alone.
- Measure impact-to-detection, detection-to-declaration, declaration-to-command, acknowledgement, mitigation, recovery, and resolution.
- Split results by severity, service tier, customer journey, region, detection source, and business-hours status.
- Track customer-first detection, missed escalations, unclear command, communication timeliness, and cadence compliance.
- Track notification episodes, actionability, duplicates, after-hours load, missed detection, routing errors, and page-budget breaches.
- Track postmortem timeliness, action age, due-date performance, verified effectiveness, and repeated contributing factors.
- Track journey availability, error-budget burn, failed or delayed value, reconciliation breaks, and SLA credits.
- Track rotation size, duty frequency, overnight work, recovery days, exceptions, sentiment, and responder attrition.
- Hold a weekly Incident Review Board chaired by the Head of Reliability with relevant directors.
- Hold a monthly executive reliability review and a quarterly control and resilience review with Internal Audit.
- Reconcile incident records monthly against support cases, customer complaints, status history, credits, and major operational anomalies to detect under-reporting.
- Use team-level scorecards to direct help and investment. Never penalize an individual for good-faith declaration.
25. Pilot the complete process on the payment path (depends on: 9, 11, 13, 20, 21, 23, 24)
Run a four-to-six-week pilot across the highest-risk journey before expanding. The interim operating floor remains active for the rest of the company.
- Include payment orchestration, ledger application, PostgreSQL platform, API edge, authentication, settlement, reconciliation, Kubernetes platform, and Support intake.
- Include teams with existing on-call experience and teams new to the model.
- Activate paid primary and secondary rotations, Duty Command, Communications, the incident record, status templates, alert standards, postmortems, and action tracking together.
- Run legacy and new paging paths in parallel for no more than one week, then make the new platform authoritative.
- Review every pilot page within one business day for actionability, context, routing, escalation, and responder load.
- Have the program team coach incidents without silently taking command.
- Correct critical process or tooling defects within 48 hours.
- Exit only after 95% timely command assignment, 95% communications compliance, no unpaid pages, complete required postmortems, tested fallbacks, and at least 50% lower pilot noise.
- Publish the pilot results, defects, and policy changes company-wide.
26. Roll out by risk with readiness gates (depends on: 25)
Expand in controlled waves and finish early enough to accumulate operating evidence before the audit. A calendar date does not override a failed readiness gate.
- Roll out remaining Tier 0 domains first, followed by Tier 1, Tier 2, and Tier 3.
- Use four waves of six to eight teams, each lasting two to three weeks.
- Gate each service on catalog ownership, appropriate coverage, active compensation, trained responders, access, tested escalation, alert quality, runbooks, and a passed tabletop.
- Require at least six responders only for direct 24x7 technical rotations. Use business-hours coverage for lower tiers.
- Give each wave a named coach and director sign-off.
- Reschedule failed gates or use a time-limited executive exception with compensating controls. Do not create silent waivers.
- Disable legacy human-paging routes after verified cutover for each wave.
- Run a quota-based noise reduction sprint in every wave, starting with the highest-volume rules.
- Pair every suppression with a compensating-detection check.
- Publish an internal adoption dashboard by team, service tier, coverage, and control gap.
- Complete critical coverage by approximately week 12 and all 28 teams by week 18.
27. Prove SOC 2 operating effectiveness (depends on: 22, 23, 24, 26)
Generate evidence through normal operation rather than reconstructing it before fieldwork. Test both control design and consistent execution.
- Map controls to the applicable Trust Services Criteria with Compliance and the auditor, including monitoring, incident identification, response, recovery, communications, and availability.
- Retain approved policies, exceptions, service ownership, schedules, compensation activation, access reviews, training, incidents, communications, reportability decisions, postmortems, actions, and exercises.
- Sample incidents monthly from first signal through verified corrective action.
- Include customer-reported events, downgraded incidents, missed timelines, non-reportable decisions, and exercises in the testing population.
- Have Internal Audit or an independent control owner test design and operation in months 4 and 6.
- Run a formal mock audit in month 7 using the evidence populations and interviews expected from the external auditor.
- Correct deviations through tracked actions with owners and dates. Never edit history to create apparent compliance.
- Verify that evidence retention covers the full auditor-defined observation period.
- Brief commanders, engineers, Support, and Compliance on the actual process without scripting inaccurate answers.
28. Inspect, adapt, and institutionalize ownership (depends on: 26, 27)
Prevent the process from decaying after rollout or the audit. Change it using measured operating evidence rather than opinion.
- Review the policy after 90 days of live operation using severity calibration, page load, missed detection, communication compliance, action closure, fatigue, and survey results.
- Remove steps that create work without reducing risk. Add controls only where incidents, exercises, or evidence show a gap.
- Reassess Tier 0 and Tier 1 classification and domain boundaries every six months.
- Assign permanent owners for policy, catalog, paging, status page, training, metrics, evidence, and exercise scheduling.
- Review compensation bands, rotation burden, accommodations, and staffing annually.
- Report severe incidents, credits, overdue high-risk actions, and ledger concentration risk to the board or risk committee quarterly.
- Maintain the ledger blast-radius program as an executive risk until failover, degraded mode, reconciliation, and regional independence meet approved objectives.
- Evaluate follow-the-sun coverage using one year of actual activation and staffing data.
- Build year-two plans for automated mitigation, safer deployments, graceful degradation, and error-budget release controls.
--- PROPOSAL 3 (agent qwen3.8-max_refine_3, alibaba/qwen3.8-max) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform with legacy paging paths disabled by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of 'not reportable'.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (29):
1. Charter the program, fund it, and start the audit clock
Convert the CEO email into a company operating process with one accountable owner, a budget, and a dated timeline that starts this week.
- Name the CTO as executive sponsor and appoint a Head of Reliability & Incident Management as the single accountable owner with full-time authority over all 28 teams.
- Stand up a three-person program office: program lead, platform engineer, reliability analyst.
- Form an eight-person steering group spanning Engineering, SRE, Support, Customer Success, Security, Legal/Compliance, Finance, and HR. It proposes; the sponsor decides within 48 hours. Never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one human-paging path, one incident record, one postmortem format, paid on-call, mandatory action tracking, and named service ownership.
- Approve budget anchored against the $1.3M in SLA credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training and exercises, 2–3 program FTEs, and 15% reserved engineering capacity.
- Confirm the SOC 2 Type II observation period with the auditor in week 1 so the interim process counts as evidence.
- Publish a one-page charter company-wide on day 2. Incident response is a company process, not a per-team preference.
- Timeline: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day interim command floor (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place this week so the next outage has a named commander and the evidence clock starts immediately.
- Publish a one-page interim severity card and one declaration path: a Slack command, a phone number, and the existing pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and senior SREs from the 12 teams already on-call. Compensate retroactively under the final policy.
- Rule from day one: a named commander within 10 minutes of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline document, and one naming convention for every major incident.
- Triage all 53 open historical action items into complete, re-plan, or formally risk-accept within 30 days, prioritising ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
3. Rebuild the forensic baseline of incidents, alerts, and money lost (depends on: 1)
Rebuild the facts before locking any design. This is both the design input and the frozen 'before' picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. That list becomes the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule, and outcome. Identify the top 50 noisiest rules and every rule with no owner or runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines in one signed document: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11/64 actions closed, on-call in 12 of 28 teams.
4. Run the listening tour and publish the written fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem, and answer it with a written, signed deal.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Publish the deal in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- State explicitly that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive accommodation process for health, disability, or caregiving constraints.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness, and burnout. Re-measure at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Map compliance, evidence, and audit requirements from day one (depends on: 1)
Design evidence as a by-product of operations, not a reconstruction before the auditor arrives. The interim process in week 1 is already evidence.
- Map the process to SOC 2 Trust Services Criteria: CC7.2–CC7.5 for monitoring, identification, response, and recovery; CC2.2/CC2.3 for internal and external communication; CC5 for control activities; A1.2 for availability.
- Define the evidence set: incident records with timestamps, paging and acknowledgement logs, role assignments, status-page history, postmortems, action closure with verification, training register, drill records, reportability decisions including 'not reportable'.
- Establish retention, confidentiality, legal-hold, and access rules for operational, security, and privileged records. Require role-based access, MFA, and periodic access review.
- Version and approve all policy documents from day one: Incident Response Policy, Severity Standard, On-Call Policy, Communications Standard, Postmortem Standard, Alert Quality Standard.
- Record every control exception with an owner, compensating control, approval, and expiry date.
- Start monthly evidence sampling immediately rather than reconstructing before the audit.
6. Build the service ownership catalog and criticality tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has a named owner. A wrong owner recreates the pager objection. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Assign one accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link, and dependency list.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. **Tier 1**: customer-facing but degradable. **Tier 2**: internal or batch. **Tier 3**: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, and third parties including sponsor banks and processors.
- Record SLO, RTO, RPO, active-active vs region-bound, failover method, and dependencies that defeat nominal redundancy for every Tier 0/1 service.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform. They fail differently and need different hands.
- Orphan services get an owner within 30 days or a decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Missing ownership or runbooks blocks Tier 0/1 releases.
7. Adopt the severity scale, declaration rights, and incident lifecycle (depends on: 3, 6)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix. Anyone may declare. Nobody is penalised for over-declaring.
- **SEV1 (crisis):** money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; deploy freeze; status page in 15; regulatory assessment within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms and scribe if customer-visible; status page in 30 minutes; mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads; business-hours comms; postmortem if customer-detected, over two hours, a repeat, or credit-generating.
- **SEV4:** no customer impact. Ticket only. Never pages.
- Payments-specific anchors: value of payments at risk per hour, settlement-deadline proximity, number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: any incident touching the ledger cluster, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident become at least SEV2.
- Only the Incident Commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
8. Define incident roles, authority, dual-control, and handover discipline (depends on: 7)
Solve 'nobody in charge for an hour' by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- **Communications Lead:** the single voice for status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners, and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, shields the commander from executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory submissions; the commander owns operational facts.
- Command claimed within 5 minutes and stated in channel. Distinct people for command, comms, and technical lead at SEV1/SEV2. Every handover announced verbally and in writing with the exact time.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
9. Staff 24x7 with a three-layer coverage model, not 28 night rotations (depends on: 6, 8)
Do not create 28 night rotations. That is precisely what engineers are rejecting. Centralise coordination in a trained command corps and keep technical ownership local.
- **Layer A — Incident Command corps:** approximately 30 certified volunteers drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to seven months, with primary and secondary at all times. Paired with a Communications Lead pool of approximately 18 from Support, CS, and engineering management, and a scribe pool used as the training entry point.
- **Layer B — critical-path domain rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Each carries 24x7 primary plus secondary with a minimum of six trained responders.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- Staffing math: 260 engineers can sustain 10–12 domain rotations and one command corps. They cannot sustain 28 independent night rotas. Teams too small get headcount, service reassignment, or a time-limited executive exception. Never a two-person 24x7 rotation.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Evaluate a US-only paid night rotation now. Treat a Lisbon or APAC follow-the-sun cell as a 12-month option, not a year-one dependency.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
10. Approve paid on-call, New York labor compliance, and fatigue safeguards (depends on: 4, 9)
Unpaid on-call in New York is both a retention problem and a wage-hour exposure. Pay must be live in payroll before any new mandatory night rotation starts. Publish the actual numbers, then ask people to sign up.
- Indicative scheme locked by HR, Finance, and employment counsel within 14 days: approximately $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, holiday premiums, and approximately $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as approximately 15% of delivery load. Count incident leadership in promotion criteria. Publish a documented exemption path for caring or health reasons with no career penalty.
- Publish on-call load by team quarterly.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
11. Set the alert-quality standard and a hard page budget (depends on: 3, 6)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to page a human. Pair every noise cut with a missed-detection review so suppression cannot hide incidents.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a page budget of two out-of-hours pages per person per week. Breach blocks new alert creation for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target: 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, noise below 15%, with no loss of Tier 0/1 detection coverage.
12. Detect payment and ledger failures before customers do (depends on: 6, 11)
The goal is blunt: stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. Every customer-first incident becomes a named defect with a tracked action.
- Define SLOs and business SLIs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts so a single-tenant outage is caught before the account manager calls.
- Five similar support tickets in ten minutes auto-creates a triage incident. Credible partner and processor notifications enter the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a reviewed control failure, not a footnote.
- Validate that nominal two-region services do not depend on single-region databases, identities, queues, or third parties.
- Run the replay test: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish replay coverage as a leading metric.
13. Consolidate to one paging platform, one incident record, one status page (depends on: 7, 9, 11)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test, and two weeks of verified operation in the new platform.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration time, acknowledgements, role assignments, severity changes, decisions, comms sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test paging, escalation, status publication, and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
14. Codify the five-minute escalation path and live execution doctrine (depends on: 8, 9, 13)
Write one unskippable path from 'something looks wrong' to 'someone is in charge'. The default action is never waiting. If nobody claims command within 5 minutes, the platform assigns it and announces it.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If no one claims command within 5 minutes, the platform assigns it and announces it. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute.
- Separate diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Controlled backlog drain. Reconciliation completed before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Closure requires a stability observation window and explicit handback.
15. Set the service-readiness bar and write major-incident playbooks (depends on: 6, 9, 12)
Nobody responds well to unfamiliar systems at 3 a.m. without runbooks. A service must earn the right to page a human at night. The shared ledger cluster is the single largest structural risk.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, RPO/RTO, and a stated data-loss and latency impact.
- Write major-incident playbooks for: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO.
- Open a parallel architecture workstream on blast-radius reduction: partitioning, read replicas, isolation of non-critical readers, with board-visible milestones.
- Enforcement: no readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Runbooks are peer-reviewed, version-controlled, and marked stale if not exercised twice a year.
16. Run one communications clock for internals, customers, and regulators (depends on: 7, 8, 13)
Replace 'whoever is around' with one timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure. State impact and the next update time. Never speculate on cause, recovery time, data integrity, or blame.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer tier list.
- Regulatory checkpoint on every SEV1 and every security-related SEV2, completed within one hour and recorded even when the answer is 'not reportable'.
- Obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
17. Tie incidents to SLA credits and true financial cost (depends on: 7, 16)
Link incidents to money so severity, credits, and investment decisions stay honest. Finance should not learn about outages from invoices.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer automatically from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families so the quarterly report can show which reliability investments would have prevented which payouts.
- Track failed payment volume, delayed value, reconciliation breaks, and impacted customers alongside credits as the true incident cost.
- Target: credits from $1.3M to under $400k on a trailing annualised basis within 12 months. Use the delta as the standing business case for on-call pay and reliability capacity.
18. Make blameless postmortems mandatory with one format and fixed deadlines (depends on: 7, 8)
Replace 'some incidents, various formats' with one mandatory format, fixed deadlines, and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident detected by a customer first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did?
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and are never the commander of the incident under review.
- Publish a searchable library and a quarterly top-five recurring causes analysis, with restricted versions for security or privileged content.
19. Enforce action ownership, reserved capacity, and tracking (depends on: 13, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity. Closing a ticket without evidence does not close the action.
- Every action gets a single named owner, priority class, due date, expected risk reduction, verification method, and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment within 7 days, corrective within 30, strategic within 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing 15% of engineering capacity for reliability and incident actions, protected by the sponsor. Without it the actions will not land.
- Escalation for overdue items: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items require director approval and written residual-risk acceptance. They can block related feature releases.
- Verify effectiveness before closing. Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 currently open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
- Target: 90% of high-priority actions closed by due date within two quarters.
20. Train and certify every role before independent duty (depends on: 8, 14, 16, 18)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it. The certification register is an audit artefact.
- All employees (30 min): how to recognise impact, declare, and find the incident channel and status page.
- Responder (half day): severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe (2 hours): timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander (2 days plus shadowing): command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead (1 day): status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification valid 12 months, renewed by simulation. Nominate an incident-management champion in each of the 28 teams.
21. Rehearse with tabletops, game days, and unannounced drills (depends on: 13, 15, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced paging drill to measure real overnight acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents. Publish time-to-commander, time-to-first-update, and time-to-mitigation.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
22. Publish Incident Management Policy v1 and the exception register (depends on: 7, 8, 9, 10, 11, 14, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: you are not on-call for other teams' services.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 10, 12, 13, 15, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once: severity scale, command corps, single paging tool, page budget, status-page policy, mandatory postmortems, paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes for pilot services, pages halved, no unpaid page, every required postmortem filed on time, on-call sentiment not worse.
- Publish a one-page result to the whole company.
24. Instrument metrics, dashboards, review cadence, and anti-gaming (depends on: 13, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone. Do not reward hiding incidents or silencing pages.
- Response: time to detect from first impact, time to declare, time to commander, acknowledgement, mitigate, resolve. Split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, page-budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Team scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
25. Roll out in risk-ordered waves with readiness gates (depends on: 23, 24)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Finish all 28 teams at least five months before the audit report date so the observation window covers the whole company.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, never waived.
- Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Run noise burn-down as a quota inside each wave. Give every team its ranked list of noisiest rules from the baseline. After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check. Review missed detections monthly with the same seriousness as noise.
- Publish a live adoption scoreboard.
- Repeat the fairness contract in every wave. If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
26. Run change management, fairness, and pager culture from day one (depends on: 4, 10)
Run this in parallel from day one. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for your services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
27. Produce SOC 2 evidence by operating, test internally, and mock-audit (depends on: 22, 24, 25)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence. Documented exceptions beat a claim of perfection.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5, CC2.2/CC2.3, CC5, and A1.2. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and role assignments, paging and acknowledgement logs, status-page history, notification decisions including 'not reportable', postmortems, action closure with verification, training and certification register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in months 4 and 6, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the same evidence populations the auditor will request, including interviews with a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
28. Maintain the program risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and pre-commit the response. Review monthly with the sponsor.
- Too few commander volunteers: command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- Compensation not approved in time: fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- Tool migration slips: cut scope on the incident-record layer, never on paging consolidation.
- Noise pruning hides a real failure: demote to ticket first, observe 30 days, then delete. Keep a recovery path and a monthly missed-detection review.
- Burnout in the 12 experienced teams during transition: weekly load monitoring with hard per-person page caps.
- A major SEV1 mid-rollout: the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- Shared ledger concentration: if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
29. Inspect at 90 days, lock year-two ownership, and prevent decay (depends on: 24, 25, 27)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging metrics to leading ones: error-budget burn, near-miss rate, drill performance, detection coverage.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 4 (agent grok4.6_refine_4, xai/grok-4.6) ---
Estimated complexity: high
Success metrics: - A named incident commander is announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- By day 7, every suspected major incident uses one record, one channel, and a named commander; Support can declare without engineering confirmation.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Replay coverage: by day 90, a named detector exists that would have caught at least 90% of the 31 historical incidents, with the expected detection minute documented.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of customer-visible SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk. No two-person 24x7 rotation.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any new mandatory night rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of not reportable.
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (26):
1. Charter the program, fund it, and start the audit clock
Turn the CEO email into a chartered company program within 48 hours. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability as full-time owner, with a three-person office: program lead, platform engineer, and analyst.
- Form a small decision group of Engineering, SRE, Support, Customer Success, Security, Legal, Compliance, Finance, HR, and Internal Audit. It proposes. The sponsor decides within 48 hours.
- Lock the non-negotiables now: one severity scale, one human-paging platform, one incident record, one postmortem format, mandatory action tracking, paid on-call, and named service ownership.
- Confirm the SOC 2 Type II observation window with the auditor in week 1 so the interim process counts as evidence.
- Fund tooling, compensation, training, exercises, and reserved engineering capacity against the $1.3M in SLA credits.
- Reserve 15% of engineering capacity for detection, runbooks, and incident actions, protected by the sponsor.
- Publish a one-page charter on day 2. Clock: floor by day 7, design weeks 1–6, pilot weeks 7–12, rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
2. Install a seven-day operating floor (depends on: 1)
Do not wait for policy, tooling, or the audit. Put a crude but real process in place this week so the next outage already has an owner.
- Publish a one-page interim severity card and one declaration path: chat command, phone number, and current pagers, all reaching the same duty person.
- Staff interim primary and backup Duty Incident Commanders 24x7 from engineering managers and the 12 teams already on-call. Pay them retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Authorize Support and Customer Success to declare from credible customer reports without waiting for engineering confirmation.
- Use one channel, one bridge, one timeline, and one naming convention for every suspected major incident.
- Triage the 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept, starting with ledger integrity, duplicate payments, regional failover, and detection gaps.
- Hold a 15-minute daily operations review until the permanent process is live.
- Replay one of the two nobody-in-charge incidents as a tabletop within 14 days using this floor process.
3. Rebuild the forensic baseline and incident-replay catalog (depends on: 1)
Rebuild the facts before locking design. This is the design input, the frozen before-picture for the CEO, and the test set for detection work.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, and resolve, who led, credits paid, and contributing factors.
- For each of the 40% customer-first detections, name the exact missing signal. That list is the detection backlog.
- Reconstruct the two nobody-in-charge incidents minute by minute. They are the test case for every design decision.
- Profile 3,400 monthly alerts by tool, team, rule, and outcome. Name the top 50 noisy rules and every rule with no owner or runbook.
- Quantify true cost: credits, failed payment volume, delayed value, reconciliation breaks, and engineering hours lost.
- Freeze baselines: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 85% noise, 11 of 64 actions closed, on-call in 12 of 28 teams.
- Build a **replay catalog**: for each historical incident, the detector that should now fire, at which minute, and the owner of the gap.
4. Publish the on-call fairness contract (depends on: 1)
Engineer pushback against carrying a pager for other teams' code is the largest delivery risk. Treat it as a design constraint, not an attitude problem.
- Interview all 28 teams plus Support, Customer Success, and Sales within two weeks. Separate unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, and fear of blame. Each needs a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commanders. Do not punish them by making them the company's permanent night watch.
- Recruit 12–15 credible engineers as co-authors so the process is not done to the teams.
- Publish the **fairness contract** in writing: you are paged only for services your team owns, a trained commander runs the room, nights are paid, noise is a defect we fix, and postmortem actions get protected sprint capacity.
- State that platform on-call may triage unknown ownership but does not permanently inherit another team's service.
- Provide a non-punitive exemption path for health, disability, or caregiving.
- Run a baseline sentiment survey on fairness, alert trust, willingness, and burnout. Repeat at 60, 120, and 365 days.
- No new mandatory night rotation starts until this contract, compensation, and training are live.
5. Build the service catalog, ownership, and journey tiers (depends on: 3)
You cannot route a page correctly across 180 services until every service has one owner. The catalog is the single source of truth for paging, impact, status-page components, and audit evidence.
- Record owning team, manager, chat channel, escalation policy, dashboards, runbook, dependencies, regions, and data stores for all 180 services.
- Tier by business impact, not technology. **Tier 0**: money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation. Tier 1: customer-facing but degradable. Tier 2: internal or batch. Tier 3: non-critical.
- Map every customer journey — initiate, authorise, settle, reconcile, refund, report, onboard — to services, data stores, regions, sponsor banks, and processors.
- Record SLO, RTO, RPO, regional mode, failover method, and fake-redundancy dependencies for every Tier 0/1 service.
- Keep coordinated but separate responders for the ledger application and the PostgreSQL platform. They fail differently.
- Orphan services get an owner in 30 days or an approved decommission date. Unowned Tier 0 is an executive escalation and a release blocker.
- Coverage follows the journey, not the org chart. Small teams that own Tier 0 pieces get headcount, service reassignment, or membership in a domain rotation. Never a two-person 24x7 rota.
6. Lock severity levels, integrity flags, and the incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Classify on actual or credible customer, financial, security, and contractual harm, never on who reported it or how hard the fix looks.
- **SEV1 crisis**: money moved wrongly, duplicated, or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs, and Executive Duty Officer; bridge in 5 minutes; status page in 15; deploy freeze; regulatory assessment within 1 hour; mandatory postmortem.
- SEV2 critical: material degradation of a payment journey; settlement window at risk; a large or strategic customer fully down; SLA breach likely. Triggers: commander and SMEs paged; comms if customer-visible; status page in 30 minutes; mandatory postmortem.
- SEV3 contained: narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads. Business-hours comms. Postmortem if customer-detected, longer than 2 hours, a repeat, or credit-generating.
- SEV4: no customer impact. Ticket only. Never pages a human.
- Attach an **integrity, security, settlement, or regulatory flag** to any severity. The flag forces dual-control, Legal, and the reportability checkpoint without inventing a fifth level. A one-customer ledger corruption is still a flagged crisis.
- Auto-escalate to at least SEV2: any ledger-cluster event, any SEV3 open beyond 2 hours, unknown impact after 15 minutes, and any cross-team incident.
- Anyone may declare. Nobody is punished for over-declaring. Only the commander may downgrade, with evidence recorded.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed.
- Publish a decision tree with 12 worked examples taken from the real 31 incidents.
7. Define roles, authority, dual-control, and handover (depends on: 6)
Solve nobody-in-charge-for-an-hour by making command explicit, single-holder, transferable, and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander**: owns severity, priorities, roles, cadence, and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode, and commit emergency spend. Retains command when a VP joins.
- Communications Lead: single voice for the status page, account managers, executives, and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- Scribe: timestamped record of observations, decisions, owners, and state changes. Tooling assists. A human validates at SEV1 and SEV2.
- Subject-matter responders: engineers of the owning team. They mitigate. They do not run the room.
- Executive Duty Officer on SEV1: removes obstacles, shields the commander, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Security, Legal, Compliance, and Vendor Management join on defined triggers. Legal owns regulatory text. The commander owns the facts.
- Distinct people for command, comms, and technical lead at SEV1. Comms and scribe may combine only for bounded SEV2.
- Financial controls survive the incident. The commander may coordinate ledger recovery but may not bypass dual control, reconciliation, or privileged-access rules.
- Every role assignment and handover is announced verbally and in writing with the exact time and open risks.
8. Staff 24x7 with command, domains, and a Duty Triage Desk (depends on: 4, 5, 7)
Do not create 28 night rotations. That is what engineers are rejecting. Centralise coordination. Keep technical ownership local. Page people only for services they own.
- Layer A, Incident Command corps: about 30 certified people from across the 28 teams plus managers. Primary and secondary at all times. Each person roughly one primary week every six to eight months. Pair with a Comms pool of about 18 from Support, CS, and engineering management, and a scribe pool as the training entry point.
- Layer B, critical-path domains: consolidate ownership into 10–12 coherent domains such as ledger app, PostgreSQL platform, payments orchestration, settlement, auth, API edge, Kubernetes platform, data/reporting, and partner integrations. Each carries 24x7 primary plus secondary with at least six trained responders.
- Layer C, everyone else: business-hours on-call plus a written after-hours list held by the engineering manager. No night pager.
- **Duty Triage Desk**: a small paid overnight first-line rotation that owns the first ten minutes of ambiguous, unowned, or low-confidence pages. It verifies, enriches, and applies only the runbook's safe steps, then wakes the owning domain. It never sits on a suspected ledger, payment-halt, or security event. Those page commander and likely domains immediately, in parallel.
- Overnight comms: SEV1 pages a Communications Lead 24x7. Customer-visible SEV2 lets the commander publish the first status from a template; Comms is paged if the incident is still open at 30 minutes or a top-100 account is affected.
- Staffing math: 260 engineers can sustain 10–12 domain rotations, one command corps, and one triage desk. They cannot sustain 28 night rotas. Role exclusivity: nobody is primary on two rotations in the same week. Commanders may also be domain responders in different weeks.
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement is a human action within 5 minutes at SEV1/SEV2. Delivery to a device does not count. Misses fail over to secondary, then manager, then director.
- Evaluate follow-the-sun as a 12-month option, not a year-one dependency. Seed Layer B from the 12 teams already on-call.
9. Pay on-call, meet New York labor rules, and cap fatigue (depends on: 4, 8)
Unpaid on-call in New York is a retention problem and a wage-hour exposure. Pay must be in payroll before any new mandatory night rotation starts.
- Indicative scheme, locked by HR, Finance, and employment counsel within 14 days: about $1,000 per primary 24x7 week, $400 secondary, $250 business-hours, a separate $1,200 Duty Commander stipend, $400–$800 for Comms or triage-desk duty, holiday premiums, and about $150 per out-of-hours page plus hourly beyond the first hour.
- Document FLSA exempt/non-exempt treatment and New York wage-hour rules explicitly. Get the mechanics into payroll before the first new page.
- Mandatory paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- Count on-call as about 15% of delivery load. Count incident leadership in promotion criteria.
- **Hard gate: no mandatory night rotation starts before compensation is live in payroll.** Interim duty is paid retroactively. Budget roughly $0.9M–$1.2M a year, then refine with actual rotation count.
10. Enforce an alert-quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes. Nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping, and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm: SLO burn, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days. Test both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements. Return it to its owner with a correction deadline.
- Never silently disable a rule. Verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise.
- Target 3,400 to under 1,500 pages a month by day 90, under 500 by month 6, actionability above 75%, with no loss of Tier 0/1 coverage.
11. Detect payment and ledger failures before customers, then replay the past (depends on: 5, 10)
Stop customers telling you first. Detection must be driven by payment outcomes and ledger truth, not host metrics. A detector is not done until it would have caught the last 12 months.
- Define SLIs and SLOs per customer journey: initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end every 60 seconds from outside the platform and from both regions, covering critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts. Five similar support tickets in ten minutes auto-creates a triage incident.
- Route credible partner, processor, and sponsor-bank notifications into the same declaration path within five minutes.
- Record detection source on every incident. Customer-detected-first is a named defect class with a mandatory tracked action.
- Run the **replay test** against the S3 catalog. For each of the 31 incidents, name the detector that would now fire and at which minute. Close gaps the replay exposes before calling detection improved.
12. Consolidate to one pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into one operational system of record so there is one queue, one timeline, and one audit trail. Migrate without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after named owners, a successful end-to-end test, and two weeks of verified operation.
- One chat command creates the channel and bridge, pages the duty commander, sets severity, opens the timeline, and starts the clock.
- Route every page through the service catalog: service label to owning domain schedule to escalation policy. Unowned pages go to the Duty Triage Desk and Duty Command, and log a catalog defect.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- Survivability: out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat, or the identity provider is down. Test weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Codify escalation, the five-minute command rule, and vendor incidents (depends on: 7, 8, 12)
Write one unskippable path from something looks wrong to someone is in charge. The default action is never waiting. Payments also fail at processors and banks, which you cannot patch.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately; secondary at 5 minutes; domain manager and Duty Commander at 10; Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner. The missing catalog entry is logged as a control defect.
- Pre-authorise regional failover, ledger read-only mode, payment suspension, and partner-bank notification so the commander never waits for an executive. Dual-control still applies to ledger writes.
- Vendor-incident class: processor, sponsor-bank, card-network, or cloud-control-plane failure. Command is still required. Success is time-to-customer-notice, time-to-failover-decision, and queue management, not root cause at the vendor.
- Maintain and test quarterly the external contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Write live execution doctrine, readiness bars, and ledger playbooks (depends on: 5, 7, 13)
Give responders one short operating procedure from the first minute through closure. Priority is limiting customer and financial harm, not proving root cause. A service must earn the right to page a human at night.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2. Exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute, before deep diagnosis. Split diagnosis and mitigation workstreams once staffing allows.
- Decisions are stated aloud and written down. No command by direct message.
- Payments guards: protect against split-brain, replay, duplication, and out-of-order processing during regional or database recovery. Drain backlogs under control. Complete reconciliation before any payment or ledger incident is declared resolved.
- Formal commander handover beyond four hours, with a second shift staffed. Reopen if impact recurs during the stability window.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback, kill switches, escalation contacts, RPO/RTO. No sign-off, no night paging unless the manager accepts the gap in writing with an expiry date. Existing detection stays on while gaps are repaired.
- Write playbooks first for shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, and suspected duplicate payments.
- Treat the shared ledger cluster as the largest structural risk. Document and rehearse failover, read-only degraded mode, and post-recovery reconciliation. Open a parallel blast-radius workstream for partitioning and isolation of non-critical readers, with board-visible milestones and executive-signed RPO/RTO.
15. Run one communications clock for internals, customers, and account managers (depends on: 6, 7, 12)
Replace whoever is around with a timed, owned, pre-approved clock. The Communications Lead is the single author and never writes from scratch under pressure.
- Internal: one working channel and bridge per incident, plus one read-only broadcast for executives, Support, and Sales. SEV1 updates every 15–30 minutes even when nothing has changed. SEV2 every 60 minutes. SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a signed executive behaviour rule.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Status page from declaration: within 15 minutes for customer-visible SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal. Map the status page to customer journeys, not internal service names. Subscribe all 2,100 customers by default.
- Top 100 accounts: named account-manager call or email within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page and a proactive email. Honour bespoke contractual notice windows encoded in the customer record.
- Language rules: state impact and the next update time. **Never speculate** on cause, recovery time, data integrity, or blame.
16. Operationalize regulatory notice, partner clocks, and SLA credits (depends on: 6, 15)
In payments, some incidents start a legal clock at detection. Tie incidents to money so Finance does not learn about outages from invoices.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 72-hour cybersecurity event, state breach laws, GLBA/FTC Safeguards, PCI DSS if in scope, money-transmitter rules, FinCEN/OFAC, sponsor-bank and card-network windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when not reportable**, with facts, decision-maker, and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Legal owns outbound regulatory text. The commander owns the facts. Security or Legal may restrict public detail during an active threat, with the reason and an alternative stakeholder plan recorded.
- Agree with Legal and Finance the availability measurement method per contract and per component.
- Compute affected minutes per customer from the incident record and journey telemetry. Produce a proposed credit schedule within five business days of resolution.
- Decide posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain.
- Attribute credits to root-cause families. Track failed payment volume, delayed value, and reconciliation breaks as the true incident cost.
17. Make postmortems mandatory and actions enforceable (depends on: 6, 7)
Replace some incidents, various formats with one mandatory format, fixed deadlines, and a forum with teeth. Eleven of 64 closed is a process nobody enforces.
- Mandatory for every SEV1 and SEV2, any incident a customer detected first, any incident over two hours, any repeat of a known cause, any credit-generating or contract-breaching event, and any near-miss touching the ledger.
- Deadlines: factual draft within three business days, review meeting within five, published company-wide within ten. The commander owns delivery. The owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and gap, timeline, response analysis, contributing conditions, control performance, what worked, actions.
- Two mandatory questions in every review: why did a customer see this first, and why did mitigation take as long as it did.
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately.
- Weekly 60-minute Incident Review Board chaired by the sponsor, with mandatory director attendance. It ratifies severity, challenges quality, approves or rejects actions, and reviews overdue items. Facilitators are trained and are never the commander of the incident under review.
- Every action gets one named owner, priority, due date, expected risk reduction, verification method, and a ticket created automatically. If it is not on the board, it does not exist.
- Classes: containment in 7 days, corrective in 30, strategic in 90. SEV1 recurrence-prevention enters the next sprint ahead of roadmap work. Reserve 15% capacity.
- Overdue ladder: manager at +7 days, director at +14, CTO at +30. Overdue high-risk items need written residual-risk acceptance and can block related releases. Verify effectiveness before closing.
18. Train and certify every response role before independent duty (depends on: 7, 13, 15, 17)
Command is a skill, not a title. Certify before assigning duty. Use paid working time for all of it.
- All employees, 30 minutes: how to recognise impact, declare, and find the incident channel and status page.
- Responder, half day: severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- Scribe, 2 hours: timeline discipline, separating fact from hypothesis. The entry point for everyone.
- Incident Commander, 2 days plus shadowing: command presence, delegation, decisions under uncertainty, severity calls, running a room of 20, handover at 2 a.m., executive management. Certification requires two shadowed incidents and one simulated SEV1.
- Communications Lead, 1 day: status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Duty Triage Desk, half day: enrichment, safe-step limits, when to wake a domain immediately, when not to delay.
- Domain responders demonstrate dashboard, runbook, rollback, failover, and access competence for their own domain before independent primary duty. Two shadow shifts minimum. Never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
19. Publish Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17)
Collapse the design into a document people will actually open mid-outage, and make it official. Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal, and Compliance. Announced at an all-hands, not only in chat.
- Create an exception register. Every waiver has an owner, a compensating control, an expiry date, and executive approval. No indefinite verbal exceptions.
- Link the policy directly from the incident tool.
20. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 18, 19)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict. That verdict is the adoption argument.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge, plus Support intake, including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, Duty Triage Desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability, and load.
- Expect and log 20–40 process defects. Fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- Exit gate: commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
21. Instrument metrics, reviews, and anti-gaming (depends on: 12, 17, 20)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- Response: time from first impact to detect, declare, commander, acknowledge, mitigate, and resolve, split by severity, tier, journey, region, and detection source.
- Quality: customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- Alerts: volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- Learning: postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- Business: availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- People: rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk, and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits, and status-page history. Scorecards direct investment and help, never individual penalties. Declaring an incident is never held against anyone.
22. Rehearse with tabletops, game days, and night drills (depends on: 12, 14, 18, 19)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply. Never inject uncontrolled change into the production ledger.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly unannounced overnight paging drill to measure real acknowledgement times.
- One combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO, before the audit.
- Also exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Validate backups, restore, RPO/RTO, and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents.
- Complete two cross-company exercises before the audit, including regional failure and ledger recovery.
23. Burn down alert noise and roll out by risk with gates (depends on: 20, 21)
Expand in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date. Noise reduction is a quota inside each wave, not a background hope.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks. Gates are signed by the director. Failing teams are rescheduled, not waived.
- Legacy paging paths are disabled per wave, not kept as a comfort fallback. Layer C teams get business-hours schedules plus the manager escalation list. Nobody is surprised with a pager.
- Give every team its noisiest rules from the baseline. Each sprint, critical-path teams must delete, debounce, group, or convert to ticket a fixed quota.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is downgraded from paging until fixed. Pair every suppression with a compensating-detection check.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
- If the 60- or 120-day pulse shows fairness or load red, pause expansion until it is fixed.
24. Operate the fairness and culture program in parallel (depends on: 4, 9)
Run this from the moment the deal is published. Engineers judge the process on fairness. Executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical nobody-in-charge incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120, and 365 days. If fairness or load is red, pause expansion until it is fixed.
25. Prove SOC 2 operating effectiveness before fieldwork (depends on: 19, 21, 22, 23)
Design the evidence as a by-product of doing the work. Never build a parallel audit process. The real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria, notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for communication, and A1.2 for availability. Confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including not reportable, postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support, and Compliance.
- Correct control failures through tracked actions, never by editing history. A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes. Log any later change as an exception.
26. Inspect at 90 days and lock year-two ownership (depends on: 21, 23, 25)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then assign permanent owners before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows. Add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme, metrics, and the ledger blast-radius workstream, independent of the audit cycle.
- Pre-committed contingencies: too few commander volunteers makes command a rostered duty for managers until the pool reaches 30. Compensation delay means time-off-in-lieu plus a phased stipend, but never mandatory nights with nothing. A major SEV1 mid-rollout slips the wave schedule by one wave and is told to the sponsor the same day. Shared-ledger concentration that slips is escalated to the board as an accepted risk with a dated plan.
- Year-two candidates: follow-the-sun coverage, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar, and board reporting as permanent commitments owned by the Head of Reliability.
--- PROPOSAL 5 (agent deepseek-v4-pro_refine_5, deepseek/deepseek-v4-pro) ---
Estimated complexity: high
Success metrics: - Named incident commander announced within 5 minutes for 95% of SEV1/SEV2 incidents from month 3; zero incidents with unclear ownership beyond 10 minutes after day 30.
- Median time to detect falls from 22 minutes to under 10 minutes by day 90 and under 5 minutes by month 6.
- Customer-first detection falls from 40% to under 20% by day 90, under 10% by month 6, and under 5% by month 12.
- Median time to mitigate falls from 3 h 10 min to under 90 minutes by day 120 and under 60 minutes by month 6; SEV1 90th percentile under 3 hours.
- 95% of SEV1/SEV2 pages acknowledged by a human within 5 minutes by month 3; device delivery never counted as acknowledgement.
- Status page posted within 15 minutes of SEV1 and 30 minutes of SEV2 in 95% of cases; update-cadence compliance above 95%.
- Monthly pages fall from 3,400 to under 1,500 by day 90 and under 500 by month 6, with actionability above 75% and noise below 15%, with no loss of Tier 0/1 detection coverage.
- Out-of-hours pages stay at or below 2 per responder per week; page-budget breaches trend to zero.
- All six legacy alerting tools consolidated into one paging platform, with legacy paging paths disabled per wave and fully retired by month 6.
- 100% of Tier 0/1 services have a named owning team, tier, escalation policy, dashboard, and runbook by day 30; 100% of all 180 services by day 120.
- 24x7 primary and secondary coverage for every Tier 0/1 journey by day 60, staffed from 10–12 domain rotations of at least six trained responders each, plus a staffed overnight Duty Triage Desk.
- 30 certified incident commanders and 18 certified communications leads active, giving continuous primary and secondary command cover with no rotation more often than one week in six.
- Paid on-call live in payroll before any rotation pages a human; zero unpaid night pages after policy publication; interim duty paid retroactively.
- 100% of required postmortems drafted in 3 business days, reviewed in 5, and published in 10, in the single approved format, from month 2.
- Postmortem action closure rises from 17% (11/64) to 90% of high-priority actions closed by due date within two quarters; all 53 legacy open actions triaged within 30 days.
- Repeat incidents sharing an unaddressed contributing factor fall by at least 50% within 6 months.
- SLA credits fall from $1.3M to under $400k on a trailing annualised basis within 12 months; customer-impacting incidents fall from 31 to under 15 per year.
- Monthly availability meets or exceeds 99.95% per customer journey by month 6, with exceptions reviewed at the executive reliability meeting.
- Regulatory reportability assessed and recorded within 1 hour for 100% of SEV1 and security-related SEV2 events, including decisions of "not reportable".
- At least one tabletop per group per month, one game day per quarter, two unannounced night paging drills, and two full cross-company exercises completed before the audit, each producing tracked actions.
- Month-5 internal control test and month-7 mock audit find no unowned high-risk gap, with complete evidence for at least 95% of sampled incidents; SOC 2 Type II incident-response controls pass with zero exceptions.
- On-call fairness sentiment above 70% agreement at 120 days, improving quarter over quarter, with no increase in attrition among responders.
- Ledger blast-radius reduction workstream has executive-signed RPO/RTO, a rehearsed failover and read-only mode by month 6, with milestones reported to the board quarterly.
Steps (30):
1. Charter the program and start the audit clock
Turn the CEO email into a chartered company program with one accountable owner and authority across all 28 teams. Incident response becomes a **company operating process**, not a per-team preference.
- Name the CTO as executive sponsor and a Head of Reliability & Incident Management as full-time accountable owner, with a small office: 1 program lead, 1 platform engineer, 1 analyst.
- Form a decision group (Engineering, SRE/Platform, Support, Customer Success, Security, Legal, Compliance, Finance, HR, Internal Audit). It proposes; the sponsor decides within 48 hours. It is never a 28-person committee.
- Lock the non-negotiables now: one severity scale, one paging platform, one incident record, one postmortem format, mandatory action tracking, and paid on-call.
- Set the clock ahead of the audit: operating floor day 7, design weeks 1–6, pilot weeks 7–12, wave rollout weeks 13–24, internal test month 5, mock audit month 7, audit month 8.
- Fund it against the $1.3M of credits: tooling $150–250k/yr, on-call compensation $0.9–1.2M/yr, training, exercises, and 15% reserved engineering capacity.
- Publish a one-page charter company-wide on day 2 so nobody learns about this from rumour.
2. Seven-day operating floor so the next outage already has an owner (depends on: 1)
Do not leave the company unprotected for six weeks of design. Put a crude but real process in place within seven days, and start retaining evidence from day one.
- Publish a one-page interim severity card and a single declaration path: one Slack command, one phone number, one pager service, all reaching the same duty person.
- Stand up an interim Duty Incident Commander roster (primary plus backup, 24x7) from engineering managers and senior SREs. Pay it retroactively under the final policy.
- Rule from day one: a **named commander within 10 minutes** of any suspected major incident, announced in writing in the incident channel.
- Instruct Support and Customer Success to declare from credible customer reports immediately, without waiting for engineering confirmation.
- Triage the 53 open historical action items into complete, re-plan, or formally risk-accept, prioritising ledger integrity, duplicate payments, regional failover and detection gaps.
- Run a 15-minute daily operations review until the permanent process is live.
3. Forensic baseline of incidents, alerts and money lost (depends on: 1)
Rebuild the facts before designing anything. This is both the design input and the frozen "before" picture for the executive team and the auditor.
- Re-code all 31 incidents: trigger, services, detection source, timestamps for impact-start, detect, declare, commander-assigned, mitigate, resolve, who led, credits paid, contributing factors.
- For each of the 40% customer-first detections, name the exact signal that was missing. This list becomes the detection backlog.
- Reconstruct the two "nobody in charge" incidents minute by minute. They are the burning-platform narrative and the test case for every design decision.
- Profile the alert estate: 3,400 monthly alerts by tool, team, rule and outcome; the top 50 noisiest rules; every rule with no owner or no runbook.
- Quantify the true cost: credits, failed payment volume, delayed value, reconciliation breaks, support hours, engineering hours lost.
- Freeze the baselines formally: MTTD 22 min, 40% customer-first, MTTM 3 h 10 min, 31 incidents, $1.3M credits, 11/64 actions closed, on-call in 12 of 28 teams.
- Confirm the SOC 2 Type II observation window with the auditor; map the future response controls to applicable Trust Services Criteria; define evidence set and retention requirements.
4. Listening tour, resistance map and the written on-call deal (depends on: 1)
Engineer pushback is the largest delivery risk, not an attitude problem. Diagnose it precisely and answer it with a written, signed deal.
- Interview all 28 teams plus Support, CS and Sales within two weeks. Separate the objections: unpaid work, lost sleep, unfamiliar code, missing runbooks, unfair load, fear of blame. Each has a different remedy.
- Harvest practice from the 12 teams already on-call. They supply the pilot teams and the first commander candidates.
- Publish **the deal** in one sentence: you are paged only for services your team owns, a trained commander runs the room, nights are paid, alert noise is a defect we fix, and your postmortem actions get protected sprint capacity.
- Recruit 12–15 credible engineers and managers as a co-design working group so the process is authored with the teams, not at them.
- Run a baseline sentiment survey on fairness, trust in alerts, willingness and burnout. Re-measure at 60, 120 and 365 days.
5. Service catalog, ownership, tiering and customer-journey map (depends on: 3)
You cannot route a page correctly across 180 services until every service has an owner. Build a machine-readable catalog that becomes the single source of truth for paging, impact, status components and audit evidence.
- One accountable team per service, plus engineering manager, chat channel, escalation policy, dashboards, runbook link and dependency list.
- Tier by business impact, not technology: **Tier 0** (money movement, ledger, shared PostgreSQL, auth, settlement, reconciliation), Tier 1 (customer-facing, degradable), Tier 2 (internal/batch), Tier 3 (non-critical).
- Map every customer journey — initiate, authorise, settle, reconcile, report, onboard — to its services, data stores, regions, sponsor banks and processors.
- Record for each Tier 0/1 service: SLO, RTO, RPO, active-active or region-bound, failover method, and dependencies that make nominal redundancy fake.
- Assign separate but coordinated responders for the ledger application and the PostgreSQL platform; they fail differently and need different hands.
- Every orphan service gets an owner within 30 days or an approved decommission date. An unowned Tier 0 service is an executive escalation and a release blocker.
6. Severity standard, declaration rights and incident lifecycle (depends on: 3, 5)
Replace judgement calls with a lookup table. Severity is set by actual or credible customer and financial impact, never by the seniority of the reporter or the difficulty of the fix.
- **SEV1 (crisis):** money moved wrongly, duplicated or lost; ledger integrity in doubt; confirmed security or data exposure; both regions impaired; payment processing halted or suspended. Triggers: immediate 24x7 page of commander, comms, scribe, owning SMEs and executive duty officer; bridge in 5 minutes; status page in 15; regulatory assessment started within 1 hour; mandatory postmortem.
- **SEV2 (critical):** material degradation of a payment journey, a settlement window at risk, a strategic or large customer fully down, or SLA breach likely. Triggers: commander and SMEs paged, status page in 30 minutes, account-manager briefing pack, mandatory postmortem.
- **SEV3 (contained):** narrow or single-customer impact with a workaround and no integrity or security risk. Owning team leads, business-hours communications, postmortem only if customer-detected, over two hours, or a repeat.
- **SEV4:** no customer impact. Ticket only; never pages a human.
- Payments-specific anchors alongside error rates: value of payments at risk per hour, settlement-deadline proximity, and number of affected named accounts. Percentage thresholds are guardrails, never an excuse to under-classify integrity or settlement risk.
- Auto-escalation: anything touching the ledger cluster, any SEV3 open beyond 2 hours, any cross-team incident, and any incident whose impact is still unknown after 15 minutes becomes at least SEV2.
- **Anyone may declare and nobody is penalised for over-declaring.** Only the commander may downgrade, with recorded evidence.
- Lifecycle: Detected → Declared → Triaged → Mitigated (customer harm ends) → Monitoring → Resolved (backlog drained and ledger reconciled) → Reviewed. Detection time starts at the earliest reliable indication of impact, including customer reports.
- Publish a decision tree plus 12 worked examples taken from the real 31 incidents.
7. Roles, decision authority and handover discipline (depends on: 6)
Solve "nobody in charge for an hour" by making command explicit, single-holder, transferable and logged. Separate coordination from debugging so a commander never needs to know another team's code.
- **Incident Commander:** owns severity, priorities, role assignment, cadence and closure. Does not type in a production terminal. Pre-authorised to freeze deploys, roll back, disable features, shift traffic, invoke regional failover, put the ledger in read-only mode and commit emergency spend.
- **Communications Lead:** the single voice to the status page, account managers, executives and the hand-off to Legal for regulators. Speaks only from commander-approved facts.
- **Scribe:** maintains the timestamped record of observations, decisions, owners and state changes. Tooling assists; a human validates at SEV1/SEV2.
- **Subject-Matter Responders:** engineers of the owning team. They mitigate; they do not run the room.
- **Executive Duty Officer (SEV1):** removes obstacles, absorbs executive questions, owns board and regulator escalation. Does not take command unless a formal transfer is announced.
- Rules: command claimed within 5 minutes and stated in channel ("I am IC"); distinct people for command, comms and technical lead at SEV1/SEV2; every handover announced verbally and in writing with the exact time and open risks.
- Financial controls survive the incident: the commander may coordinate ledger recovery but may **not bypass dual control, reconciliation or privileged-access rules**.
8. Three-layer 24x7 coverage: command corps, domain rotations, triage desk (depends on: 5, 7)
Do not create 28 night rotations; that is precisely what engineers are refusing. Centralise coordination and first-line triage, and keep technical ownership local.
- **Layer A — Incident Command corps:** ~30 certified commanders drawn from across the 28 teams plus managers, giving each person roughly one primary week every six to eight months, with primary and secondary at all times. Paired pools of ~18 Communications Leads (Support, CS, engineering management) and a scribe pool used as the training entry point.
- **Layer B — domain responder rotations:** consolidate the 28 teams into 10–12 coherent response domains (ledger app, PostgreSQL platform, payments orchestration, settlement/reconciliation, auth, API edge, Kubernetes platform, data/reporting, integrations/partners). Tier 0/1 domains carry 24x7 primary plus secondary, minimum six trained responders each.
- **Layer C — everyone else:** business-hours on-call plus a written after-hours escalation list held by the engineering manager. No night pager.
- **Duty Triage Desk:** a small paid overnight first-line rotation that owns the first ten minutes of any page — verify, enrich, apply the runbook's safe steps, and only then wake a domain responder. This is the direct structural answer to "I will not carry a pager for other teams' code."
- Platform on-call is the safety net for unknown-owner pages, never the permanent owner of someone else's defect.
- Acknowledgement discipline: 5 minutes at SEV1/SEV2, auto-failover to secondary, then manager, then director. Delivery to a device is not acknowledgement.
- Evaluate in writing, with cost and dates, a Lisbon or APAC follow-the-sun cell as a 12-month option rather than a year-one dependency.
9. Paid on-call, New York labour compliance and fatigue safeguards (depends on: 4, 8)
Unpaid on-call is both a retention problem and a wage-hour exposure. Pay before asking anyone to sign up, and publish the actual numbers.
- Indicative scheme: ~$1,000 per primary 24x7 week, ~$400 secondary, ~$250 business-hours rotation, a separate ~$1,200 Duty Commander stipend, holiday premiums, ~$150 per out-of-hours page plus hourly beyond the first hour.
- Mandatory recovery: a paid recovery day after more than two hours of overnight work, a SEV1, or a qualifying SEV2. Managers arrange daytime cover instead of expecting normal output.
- HR, Finance and employment counsel publish amounts, eligibility, FLSA exempt/non-exempt treatment, New York wage-hour handling, tax treatment and payroll mechanics within 14 days.
- Load rules: no more than one primary week in six, no consecutive primary and secondary weeks, no person primary on two rotations, no on-call during PTO, and an annual cap per person.
- A rotation averaging more than two out-of-hours pages per person per week for four weeks triggers a mandatory staffing or alert-remediation review.
- **Hard gate: no mandatory night rotation starts before its compensation is live in payroll.** Interim duty is paid retroactively.
- Non-cash terms: on-call counts as ~15% of delivery load, incident leadership enters promotion criteria, a documented exemption path exists for caring or health reasons, and on-call load is published quarterly by team.
10. Alert quality contract and page budget (depends on: 3, 5)
3,400 alerts at 85% noise is why detection takes 22 minutes: nobody reads them. Make alert quality a condition of being allowed to wake a human.
- Every paging alert must declare owning team, affected service, customer or SLO impact, expected action, dashboard, runbook, deduplication key, severity mapping and escalation policy. Anything failing this becomes a ticket.
- Page on symptoms of customer harm — SLO burn rate, payment success rate, queue age against settlement deadlines. Cause-based CPU and memory alerts become dashboards.
- Run new alerts in shadow mode for seven days, testing both firing and recovery, unless an emergency exception is recorded.
- Set a **page budget of two out-of-hours pages per responder per week**. A breach blocks new paging alerts for that team and triggers a tuning sprint.
- Auto-quarantine any rule firing more than five times a month without action, or with over 70% no-action acknowledgements, and return it to its owner with a correction deadline.
- Never silently disable a rule: verify compensating detection, record the decision, name a correction owner, and review missed detections monthly alongside noise so suppression cannot hide incidents.
11. Detection uplift on the money path, validated by incident replay (depends on: 5, 10)
The goal is blunt: stop customers telling us first. Detection must be driven by payment outcomes and ledger truth, not host metrics.
- Define SLIs and SLOs per customer journey: payment initiation success, authorisation latency, settlement file timeliness, reconciliation break rate, API availability, webhook delivery, reporting freshness. Set internal targets stricter than the 99.95% contract.
- Run external synthetic transactions end to end — create payment, ledger write, webhook, confirmation — every 60 seconds, from outside the platform and from both regions, covering both critical payment methods.
- Add ledger assurance: continuous double-entry balance checks, duplicate-identifier detection, replication lag and failover-readiness alarms, backup verification, and settlement-window countdown alerts.
- Add per-customer anomaly detection for the top 100 accounts; five similar support tickets in ten minutes auto-creates a triage incident; partner and processor notifications enter the same declaration path within five minutes.
- Run the **replay test**: for each of the 31 historical incidents, name the detector that would now fire and at which minute. Publish "replay coverage" as a leading metric and close the gaps it exposes.
- Record detection source on every incident. "Customer detected first" becomes a named defect class with a mandatory tracked action.
12. One pager, one incident record, one status page (depends on: 6, 8, 10)
Collapse six alerting tools into a single operational system of record so there is one queue, one timeline and one audit trail — migrated without creating a monitoring gap.
- Select one paging and scheduling platform, one integrated incident record, and one hosted status page. Time-box selection to two weeks.
- Ingest all six sources first, deduplicate and correlate, then retire a legacy path only after its signals have named owners, a successful end-to-end test and two weeks of verified operation.
- One chat command declares an incident: it creates the channel and bridge, pages the duty commander, sets severity, opens the timeline and starts the clock.
- Route every page through the service catalog: service label → owning domain schedule → escalation policy.
- Capture evidence automatically: declaration, acknowledgements, role assignments, severity changes, decisions, communications sent, mitigated and resolved times, postmortem linkage. Retain at least 12 months with role-based access and periodic access review.
- **Survivability:** out-of-band SMS and phone paging, mobile fallback, and a printed offline runbook that works when one AWS region, chat or the identity provider is down. Test paging, escalation, status publication and bridge access weekly.
- Set a hard date after which a page outside this platform creates no on-call obligation.
13. Escalation ladder and the five-minute command rule (depends on: 7, 8, 12)
Write one unskippable path from "something looks wrong" to "someone is in charge", and make the default action never be waiting.
- All entry points converge on one declaration: automated alert, engineer observation, support ticket, account manager, partner bank, customer hotline, security finding.
- SEV1/SEV2 ladder: owning primary immediately → secondary at 5 minutes → domain manager and Duty Commander at 10 → Executive Duty Officer at 15. SEV2 requires command assigned within 15 minutes.
- If nobody claims command within 5 minutes, **the platform assigns it and announces it**. The assignee may hand over but may not decline.
- Cross-team pull: the commander may page any domain's on-call directly with a 10-minute acknowledgement obligation. This reciprocity is what makes single-team ownership viable.
- If ownership is unclear after 10 minutes, the commander keeps the incident and names a temporary owner; the missing catalog entry is logged as a control defect.
- Pre-authorised triggers for regional failover, ledger read-only mode, payment suspension and partner-bank notification so the commander never waits for an executive.
- Maintain and test quarterly the external escalation contacts: AWS and database premium support, processors, sponsor banks, card networks.
14. Live execution doctrine and payments safety rules (depends on: 7, 13)
Give responders one short operating procedure from the first minute to handback. The priority is limiting customer and financial harm, not proving root cause.
- The commander opens with a scripted first message: severity, known impact, current hypothesis, immediate objective, role assignments, next update time.
- Freeze unrelated production changes during SEV1 and SEV2; exceptions are approved by the commander and recorded.
- Prefer reversible mitigations: rollback, feature flag, traffic shift, rate limit, load shed, partner reroute — before deep diagnosis.
- Split diagnosis and mitigation workstreams once staffing allows. Decisions are stated aloud and written down; no command by direct message.
- **Payments guards:** protect against split-brain, replay, duplication and out-of-order processing during regional or database recovery; drain backlogs under control; complete reconciliation before any payment or ledger incident is declared resolved.
- Long incidents: a fatigue rule and a formal commander handover checklist beyond four hours, with a second shift staffed.
- Closure requires a stability observation window by severity, an explicit handback to the owning team, Support and CS, and reopening if impact recurs.
- Readiness bar for Tier 0/1: architecture diagram, dependency map, dashboards, alert-to-runbook mapping, rollback procedure, kill switches, escalation contacts, and a stated data-loss and latency impact. No readiness sign-off, no paging alerts at night unless the manager accepts the gap in writing with an expiry date.
- Write major-incident playbooks for the top failure modes from the baseline: shared PostgreSQL failure or corruption, cross-region failover, Kubernetes control-plane loss, processor or sponsor-bank outage, settlement-window breach, queue backlog, credential compromise, suspected duplicate payments.
- Treat the **shared ledger cluster as the largest structural risk**: documented and rehearsed failover, read-only degraded mode, reconciliation-after-recovery procedure, and executive-signed RPO/RTO. Open a parallel architecture workstream on blast-radius reduction — partitioning, read replicas, isolation of non-critical readers — with board-visible milestones.
15. Internal communications protocol (depends on: 7, 12)
Standardise the internal picture so executives, Support and Sales are informed without ever interrupting the person running the incident.
- One working channel and bridge per incident, plus one read-only broadcast channel for executives, Support and Sales.
- Cadence by severity: SEV1 every 15–30 minutes even when nothing has changed, SEV2 every 60 minutes, SEV3 on state change.
- Fixed template: what is happening, customer impact in plain language, what we are doing, next update time, current commander and comms lead.
- Executives address questions only to the Executive Duty Officer. Publish this as a **behavioural commitment signed by the executive team**.
- Support and CS receive a holding statement and a live affected-customer list within 15 minutes of SEV1/SEV2 so the front line never improvises.
- Every material decision and its owner is echoed into the incident record for the postmortem and the auditor.
16. Customer communications and status page policy (depends on: 6, 12, 15)
Replace "whoever is around" with a timed, owned, pre-approved process. The Communications Lead is the single author and never writes from scratch under pressure.
- Timings from declaration: status page within 15 minutes for SEV1 and 30 for SEV2; updates every 30 and 60 minutes; resolution notice within 30 minutes of verified recovery; customer-facing summary within five business days for SEV1 and qualifying SEV2.
- Pre-approve 12–15 templates with Legal: degradation, settlement delay, API errors, third-party failure, regional failure, data-integrity investigation, security event, resolution.
- Component-level status page mapped to customer journeys, with email, webhook and RSS subscriptions; all 2,100 customers subscribed by default.
- Tiered outreach: the top 100 accounts receive a named call or email from their account manager within 30 minutes of SEV1, using one approved briefing pack. The long tail gets the status page plus proactive email.
- Language rules: state impact and the next update time; **never speculate on cause, recovery time, data integrity or blame**; state affected regions and workarounds.
- Honour bespoke contractual notice windows for enterprise accounts, encoded in the customer tier list, and record every public update in the incident timeline.
17. Regulatory, partner and legal notification playbook (depends on: 6, 16)
In payments, some incidents start a legal clock at detection. Build the assessment into the process so it is never remembered late.
- Legal and Compliance produce the obligation matrix: NYDFS 23 NYCRR 500 (72-hour cybersecurity event), state breach laws, GLBA/FTC Safeguards, PCI DSS where in scope, money-transmitter rules, FinCEN/OFAC implications, sponsor-bank and card-network contractual windows, and cyber-insurer notice.
- Mandatory reportability checkpoint on every SEV1 and every security-related SEV2, completed within one hour of declaration and **recorded even when the answer is "not reportable"**, with evidence, decision-maker and timestamp.
- Maintain a tested 24x7 contact matrix with named backups: regulators, sponsor banks, networks, outside counsel, insurer.
- Pre-draft notification letters under privilege review. Legal owns outbound regulatory text; the commander owns the facts.
- Security or Legal may restrict public detail during an active threat, but must record the reason and an alternative stakeholder plan.
- Rehearse the playbook quarterly as part of the exercise programme.
- Agree with Legal and Finance the availability measurement method per contract and per component; compute affected minutes per customer automatically from the incident record and journey telemetry; produce a proposed credit schedule within five business days of resolution.
- Decide credit posture explicitly: proactive credits for top-tier accounts, claims-based for the long tail, with a documented approval chain; attribute credits to root-cause families for investment decisions.
18. Blameless postmortem standard and Incident Review Board (depends on: 6, 7, 12)
Replace "some incidents, various formats" with one mandatory format, fixed deadlines and a forum with teeth. The discipline lives in the deadlines and the review, not the template.
- Mandatory for: every SEV1 and SEV2; any incident a customer detected first; any incident over two hours; any repeat of a known cause; any credit-generating or contract-breaching event; any near-miss touching the ledger.
- Deadlines: factual draft in three business days, review meeting in five, published company-wide in ten. The commander owns delivery; the owning manager is accountable.
- One template: executive summary, customer and financial impact, detection source and detection gap, timeline, response analysis, contributing conditions across technical and organisational factors, control performance, what worked, actions.
- Two mandatory questions in every review: **why did a customer see this first, and why did mitigation take as long as it did?**
- Blameless in writing: systems and decisions in the context available at the time, no individual named as a cause, never used in performance reviews. HR handles misconduct separately and exclusively.
- Weekly 60-minute Incident Review Board chaired by the sponsor with mandatory director attendance: ratifies severity, challenges quality, approves or rejects actions, reviews overdue items.
- Facilitators are trained and never the commander of the incident under review. Publish a searchable library and a quarterly "top five recurring causes" analysis, with restricted versions for security or privileged content.
19. Action ownership, reserved capacity and enforcement (depends on: 12, 18)
Eleven of 64 closed is the clearest symptom of a process nobody enforces. Give actions the status of customer commitments and give them real capacity.
- Every action gets one named individual owner, priority class, due date, expected risk reduction, verification method and a ticket created automatically from the postmortem. If it is not on the board, it does not exist.
- Classes with hard SLAs: containment in 7 days, corrective in 30, strategic in 90 with milestones. Actions preventing recurrence of a SEV1 enter the next sprint ahead of roadmap work.
- Reserve a standing **15% of engineering capacity** for reliability and incident actions, protected by the sponsor.
- Overdue ladder: manager at +7 days, director at +14, CTO dashboard at +30. Overdue high-risk items need director approval and written residual-risk acceptance, and can block related feature releases.
- Verify effectiveness before closing. A closed ticket without evidence does not close the action.
- Report closure rate by team monthly and include it in engineering manager objectives.
- Triage all 53 open historical actions within 30 days. Complete, re-plan, or formally risk-accept.
20. Training and certification academy (depends on: 7, 13, 15, 18)
Command is a skill, not a title. Certify before assigning duty, and use paid working time for all of it.
- **All employees (30 min):** how to recognise impact, declare, and find the incident channel and status page.
- **Responder (half day):** severity, declaration, escalation, runbook use, evidence preservation, financial-integrity precautions. Mandatory before joining any rotation.
- **Scribe (2 hours):** timeline discipline, separating fact from hypothesis. The entry point for everyone.
- **Incident Commander (2 days plus shadowing):** command presence, delegation, decisions under uncertainty, severity calls, running a room of twenty, 2 a.m. handover, executive management. Certification requires two shadowed incidents and one simulated SEV1.
- **Comms Lead (1 day):** status writing, customer tiering, legal boundaries, regulator triggers, what never to say.
- Domain responders demonstrate dashboard, runbook, rollback, failover and access competence before independent primary duty; two shadow shifts minimum; never primary in the first 90 days.
- Certification is valid 12 months, renewed by simulation. The register is an audit artefact. Nominate an incident-management champion in each of the 28 teams.
21. Exercise programme: tabletops, game days and unannounced drills (depends on: 12, 14, 20)
The process must meet a simulated SEV1 before it meets a real one. Exercises also build the commander pool and expose runbook gaps cheaply.
- Monthly tabletop per engineering group, replaying a real incident from the 31, rotating who plays commander.
- Quarterly game day in staging or a tightly governed production drill: region loss, PostgreSQL replica promotion, processor failure, queue backlog, Kubernetes control-plane degradation.
- Twice-yearly **unannounced overnight paging drill** to measure real acknowledgement times.
- Annually, one combined operational-and-security scenario testing command boundaries and disclosure control, and one regulator-notification exercise with Legal and the CISO.
- Exercise failure of the tools themselves: status page down, chat or paging provider unavailable, identity provider down.
- Never inject uncontrolled change into the production ledger; validate backups, restore, RPO/RTO and post-recovery reconciliation on replicas.
- Every exercise produces observations tracked on the same board and under the same due-date rules as real incidents, with published time-to-commander, time-to-first-update and time-to-mitigation-decision.
22. Publish Incident Management Policy v1 and the exception register (depends on: 6, 7, 8, 9, 10, 13, 15, 16, 17, 18, 19)
Collapse the design into a document people will actually open mid-outage, and make it official.
- Ten pages maximum, plus laminated one-page cards for severity, roles and authority, and communications timings. Linked directly from the incident tool.
- Include the compensation summary and the explicit rule: **you are not on-call for other teams' services**.
- Companion documents, versioned and signed: Severity Standard, On-Call Policy, Communications Policy, Postmortem Standard, Alert Quality Standard, Notification Matrix.
- Signed by the sponsor, HR, Legal and Compliance; announced at an all-hands, not only in chat.
- Create an exception register: every waiver has an owner, a compensating control, an expiry date and executive approval. No indefinite verbal exceptions.
23. Pilot on the payments critical path (depends on: 9, 11, 12, 14, 20, 22)
Prove the process on the highest-risk surface with the most willing teams before asking 28 teams to adopt it. Six weeks, tightly measured, with a public verdict.
- Scope: ledger, payments orchestration, settlement/reconciliation, PostgreSQL platform, Kubernetes platform, API edge and Support intake — including two of the 12 teams already on-call.
- Activate everything at once for them: severity scale, command corps, triage desk, single paging tool, page budget, status-page policy, mandatory postmortems, and paid rotations processed through payroll.
- Run parallel with old paths for one week, then cut over. Real incidents use the new process only.
- The program lead attends every SEV2+ as a coach, never as a shadow commander. Review every pilot page within one business day for routing, actionability and load.
- Expect and log 20–40 process defects; fix wording and tooling within 48 hours and fold the fixes into the standard before rollout.
- **Exit gate:** commander named within 5 minutes in 95% of cases, first status update on time, MTTD under 10 minutes on pilot services, pages halved, no unpaid page, every required postmortem filed on time, sentiment not worse. Publish a one-page result to the whole company.
24. Metrics, dashboards, review cadence and anti-gaming (depends on: 12, 18, 23)
Instrument the process itself so improvement is visible and the auditor sees evidence of monitoring and review. Report median and 90th percentile, never averages alone.
- **Response:** time to detect from first impact, to declare, to commander, to acknowledge, to mitigate, to resolve — split by severity, tier, journey, region and detection source.
- **Quality:** customer-first detection rate, replay coverage, status-page timeliness, update-cadence compliance, missed escalations, incomplete timelines, role conflicts.
- **Alerts:** volume, actionability, duplicates, out-of-hours pages per person, missed pages, budget breaches, missed-detection reviews.
- **Learning:** postmortems on time, actions closed by due date, action age, verified effectiveness, repeat contributing factors.
- **Business:** availability per journey, error-budget burn, failed payment volume, reconciliation breaks, SLA credits.
- **People:** rotation size, on-call frequency, swaps, recovery days, sentiment, attrition among responders.
- Cadence: weekly Incident Review Board, monthly reliability review per team, monthly executive review to the CEO, quarterly control review with Security, Compliance, Risk and Internal Audit, quarterly board summary.
- Anti-gaming: reconcile incident counts monthly against support tickets, credits and status-page history; scorecards direct investment and help, never individual penalties; declaring an incident is never held against anyone.
25. Alert noise burn-down campaign (depends on: 10, 12, 23)
Run noise reduction as a visible, quota-driven campaign in parallel with rollout, not as a background hope.
- Give every team a ranked list of its noisiest rules with counts and outcomes from the baseline.
- Each sprint, critical-path teams must delete, debounce, group or convert to ticket a fixed quota. Platform supplies burn-rate and grouping libraries and reviews the changes.
- Publish a weekly noise leaderboard that names systems, never people.
- After eight weeks, any rule without a runbook or with over 30% false pages in 14 days is automatically downgraded from paging until fixed.
- Pair every suppression with a compensating-detection check so **noise reduction never becomes blindness**; review missed detections monthly with the same seriousness as noise.
- Milestones: 3,400 → 1,500 pages a month by day 90, under 500 and below 15% noise by month 6.
26. Wave rollout to all 28 teams with readiness gates (depends on: 23, 24, 25)
Roll out in four waves of six to eight teams every two to three weeks, ordered by customer risk. Each wave passes an explicit gate rather than a date.
- Wave 1: remaining Tier 0 teams. Wave 2: Tier 1. Wave 3: Tier 2. Wave 4: Tier 3 and internal platforms.
- Per-team onboarding kit: catalog entry complete, alerts migrated and pruned to budget, runbooks at the readiness bar, rotation staffed with six trained responders or a time-limited exception, one commander candidate nominated, one tabletop passed, payroll set up.
- Each wave gets a named coach from the program office for three weeks.
- Gates are signed by the director. **A failed gate is rescheduled, never waived.** Legacy alert paths are disabled per wave, not kept as a comfort fallback.
- Layer C teams get business-hours schedules plus the manager escalation list; nobody is surprised with a pager.
- Finish all 28 teams at least four months before the audit report date so the observation window covers the whole company. Publish a live adoption scoreboard.
27. Change management, fairness and pager culture (depends on: 4, 9, 23)
Run this from day one in parallel with design. Engineers judge the process on fairness; executives judge it on visible results. Both need constant, honest communication.
- Repeat the deal in every forum: you carry a pager for **your** services, command is coordination, nights are paid, noise is a defect, actions get real capacity.
- Publicly close the two historical "nobody in charge" incidents with a written account of what would be different now.
- Weekly office hours for the first eight weeks, a support channel with a four-hour answer SLA, a per-team champion, and a short weekly newsletter that includes the bad news.
- Managers who cannot staff a fair rotation get headcount or have services reassigned. No two-person 24x7 rotations, ever.
- Recognition: incident leadership in promotion criteria, quarterly awards for best postmortem and biggest noise reduction, public thanks after every SEV1, explicit non-retaliation for good-faith declaration.
- Pulse-survey at 60, 120 and 365 days. If fairness or load is red, pause expansion until it is fixed.
28. SOC 2 evidence by design, internal testing and mock audit (depends on: 22, 24, 26)
Design the evidence as a by-product of doing the work. Never build a parallel audit process; the real process is the evidence.
- Map the process with Compliance and the auditor to the Trust Services Criteria — notably CC7.2–CC7.5 for monitoring, identification, response and recovery, CC2.2/CC2.3 for internal and external communication, CC5 for control activities and A1.2 for availability — and confirm the observation window early.
- Evidence set, retained and indexed automatically: versioned policies and exceptions, rotation schedules, compensation activation, incident records with timestamps and roles, paging and acknowledgement logs, status-page history, notification decisions including "not reportable", postmortems, action closure with verification, training register, drill records, access reviews.
- Internal Audit or an independent control owner tests design and operation in month 5, sampling incidents end to end from first signal to verified action closure.
- Run a formal mock audit in month 7 using the populations and interviews the auditor will request: a commander, a random engineer, Support and Compliance.
- **Correct control failures through tracked actions, never by editing history.** A documented exception log beats a claim of perfection.
- Freeze process wording for the remainder of the observation window once month 7 closes; log any change as an exception.
29. Risk register and pre-committed contingencies (depends on: 1)
Name the ways this programme fails and decide the response in advance. Review it monthly with the sponsor.
- **Too few commander volunteers:** command becomes a rostered duty for managers and staff engineers until the pool reaches 30.
- **Compensation not approved in time:** fall back to time-off-in-lieu plus a phased stipend, but never launch mandatory night on-call with nothing.
- **Tool migration slips:** cut scope on the incident-record layer, never on paging consolidation.
- **Noise pruning hides a real failure:** demote to ticket first, observe 30 days, then delete; keep a recovery path and a monthly missed-detection review.
- **Burnout in the 12 experienced teams during transition:** weekly load monitoring with hard per-person page caps.
- **A major SEV1 mid-rollout:** the program lead becomes a full-time responder, the wave schedule slips by one wave, and the sponsor is told the same day.
- **Shared ledger concentration:** if the blast-radius workstream slips, escalate to the board as a formally accepted risk with a dated plan.
30. Ninety-day inspect-and-adapt, then year-two sustainability (depends on: 24, 26, 28)
Guard against the classic failure: the process decays once the audit is signed. Revise on data, then build the second-year plan before the first year ends.
- At 90 days live, revise the policy using measurements, not opinions: severity calibration if teams inflate or deflate, Layer B versus Layer C membership from real page data, uncovered shifts, commander burnout, missed status updates, action closure, survey results.
- Cut steps nobody follows; add only what the last 90 days proved missing. Freeze v2 as the described process for the audit.
- Assign permanent owners for the policy, paging platform, status page, service catalog, training programme and metrics, independent of the audit cycle.
- Re-baseline targets every six months and shift emphasis from lagging to **leading indicators**: error-budget burn, near-miss rate, replay coverage, drill performance.
- Year-two candidates: follow-the-sun coverage cell, automated mitigation for the top three recurring causes, error-budget policy gating releases, per-customer real-time impact reporting, and completion of ledger blast-radius reduction.
- Keep the annual policy review, certification renewal, exercise calendar and board reporting as permanent commitments owned by the Head of Reliability.
You have just assessed the final proposals, blind to the vote; your first choice was proposal 1. Now evaluate the PROCESS itself, not the plans: a judgement, not a narration. Your answer has these parts:
- "process_evaluation": a Markdown list with one point per question: whether the convergence between agents was earned by better arguments or was mere imitation; whether the agents criticised each other's ideas or just copied them; whether anyone questioned the task's premises; what was lost along the way. Never empty.
- "process_issues": a list of concrete problems you observed in the process (empty only if you saw none).
- "suggestions": a list of concrete changes that would make the process produce a better plan.
[VOTE COMPARISON]
[SYSTEM]
You are an expert reviewer of multi-agent planning processes.
Several LLM agents drafted plans for a task, refined them over a number of rounds while seeing each other's proposals, and finally voted for the best one.
Be exhaustive but precise: name concrete steps, ideas and metrics, never generalities. Judge plans by their fitness for the task as stated, their realism, their completeness, the soundness of their order and dependencies, how measurable their success is and how they handle things going wrong.
You are an impartial evaluator, not a chronicler: assess the proposals and the process on their merits, never rationalise what happened or assume that the outcome was right.
After your analysis, answer in the requested structure.
Every text field you write will be read by a busy person who skims. Make it easy to skim: short sentences and short paragraphs; when you name several things, prefer a list to a paragraph, with sub-items when an item has parts, but keep a single fact as a sentence; lead with the point and then the evidence; name proposals and steps by number (P2, step 4); no preamble, no repetition of the question, no closing summary; bold at most one key phrase per item or paragraph. Text fields accept Markdown: a blank line between paragraphs, "- " for lists, **bold**.
[HUMAN]
Task given to the agents: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
You assessed the proposals of the final round without knowing the vote, and ranked them from best to worst as: 1, 2, 4, 3, 5. Your reasons were:
**P1 first**: it matches every rival on the shared skeleton and then adds three things that bear directly on this company's numbers — change correlation (step 13) against a 3 h 10 min MTTM, Support/AMs as a measured detection tier (step 20) against 40% customer-first detection, and a concurrency doctrine (step 8) that a 30-person command corps facing 31 incidents a year will need. It also has the earliest standalone risk register (step 34) and the deepest week-one legal/evidence scoping (step 4). The price is 35 steps, one control test instead of two, and a "zero exceptions" claim it contradicts itself.
**P2 second**: the most disciplined plan — it alone defines what a page and noise are before promising to cut them, tests controls twice (months 4 and 6), finishes company-wide by week 18 for a longer observation window, and refuses to hide overnight risk behind a Layer C manager list. It loses to P1 on two counts: no programme risk register at all, and a four-week design window for 180 services that leaves no slack.
**P4 third**: equal coverage in 26 steps and the two sharpest operational rules in the field (the triage-desk carve-out, the overnight status-page owner) plus the vendor-incident class. It falls below P2 because contingencies sit in the last step, and steps 14 and 23 bundle readiness, playbooks, ledger architecture, noise burn-down and waves into single units that are harder to gate.
**P3 fourth**: complete and well written but purely a merge, with two dating inconsistencies (month-5 vs months-4-and-6 testing; "five months before the report date" against a week-24 finish) and two mega-steps covering escalation/execution and all communications.
**P5 last**: the same content as P3 with less of it original, an extra overloaded step (14), and a headline metric that contradicts its own day-7 floor. Nothing in it is wrong; nothing in it is its own.
The vote is now revealed. Each voter received the same final proposals and chose one:
opus5_voter_1 (anthropic/claude-opus-5) voted for proposal 1. Justification: Proposal 1 covers every element of the brief with the most operational depth, and it uniquely adds items the others omit or bury: week-one audit/legal and evidence scoping, change-intelligence and deploy-safety on the critical path (the main lever on a 3h10 MTTM), Support/account management as an instrumented detection tier, a multi-incident concurrency doctrine, and an incentive guardrail warning that credits may rise before they fall.
Its staffing answer to the pager objection is the strongest: three coverage layers plus a paid overnight Duty Triage Desk, with published arithmetic showing why 28 night rotations are rejected and roughly 100 of 260 engineers carry a night duty at one week in six, gated on compensation being live in payroll.
The success metrics are specific, baseline-anchored and mostly leading rather than lagging (replay coverage, page budget, action closure, mock-audit outcomes), and the 35-step dependency graph runs coherently from a seven-day operating floor through pilot, waves, control test and mock audit before the SOC 2 window. The only real cost is length; the competing proposals are near-identical in structure but consistently thinner on these differentiators.
gpt5.6-sol_voter_2 (openai/gpt-5.6-sol) voted for proposal 2. Justification: Proposal 2 provides the best balance of completeness, operational clarity, and feasibility. It covers every required control with coherent dependencies, sustainable paid domain-based coverage, careful alert-suppression safeguards, precise communication rules, and an evidence-producing rollout that begins immediately rather than waiting for the audit.
qwen3.8-max_voter_3 (alibaba/qwen3.8-max) voted for proposal 1. Justification: Proposal 1 is the most comprehensive plan, with 35 steps that uniquely include a dedicated **change intelligence** step, a **concurrency doctrine** for simultaneous incidents, and an explicit **Support-as-detection-tier** step, while also carrying the most detailed success metrics (including change-correlated incident tracking and top-100 account outreach). Its dependencies are internally consistent and it addresses every element of the main objective with the greatest operational specificity.
grok4.6_voter_4 (xai/grok-4.6) voted for proposal 1. Justification: Proposal 1 is the most complete match to the brief: it defines severity, roles, 24x7 staffing, paid on-call, detection, communications, postmortems, and metrics in operational detail, and it puts a live command floor in place on day 7.
It uniquely adds a **Duty Triage Desk**, change intelligence, concurrency rules, and Support as a detection path, which directly attack unclear ownership, customer-first detection, and pager resistance.
deepseek-v4-pro_voter_5 (deepseek/deepseek-v4-pro) voted for proposal 1. Justification: Proposal 1 is the most comprehensive and detailed, explicitly defining every required element—severity, roles, 24x7 coverage, paid on-call, alert quality, detection/replay, communications, postmortems, metrics, and SOC 2 readiness—with clear dependencies and measurable success metrics. It uniquely adds high-value components such as change intelligence, SLA credit guardrails, and a Duty Triage Desk, which directly target the stated pain points.
The selected plan is proposal 1 (4 of 5 votes).
Your answer has these parts:
- "agrees": true if the selected plan is your first choice, false otherwise.
- "comment": two to four short sentences confronting the vote with your ranking. If they differ, say on which criteria the voters and you weigh the proposals differently, and whether the voters' justifications reveal a bias (length, confidence, style, position in the list). If they agree, say whether the voters' justifications rest on the same merits you found or on others. Do not revise your ranking.
[BRIEF]
[SYSTEM]
You are an expert reviewer of multi-agent planning processes.
Several LLM agents drafted plans for a task, refined them over a number of rounds while seeing each other's proposals, and finally voted for the best one.
Be exhaustive but precise: name concrete steps, ideas and metrics, never generalities. Judge plans by their fitness for the task as stated, their realism, their completeness, the soundness of their order and dependencies, how measurable their success is and how they handle things going wrong.
You are an impartial evaluator, not a chronicler: assess the proposals and the process on their merits, never rationalise what happened or assume that the outcome was right.
After your analysis, answer in the requested structure.
Every text field you write will be read by a busy person who skims. Make it easy to skim: short sentences and short paragraphs; when you name several things, prefer a list to a paragraph, with sub-items when an item has parts, but keep a single fact as a sentence; lead with the point and then the evidence; name proposals and steps by number (P2, step 4); no preamble, no repetition of the question, no closing summary; bold at most one key phrase per item or paragraph. Text fields accept Markdown: a blank line between paragraphs, "- " for lists, **bold**.
[HUMAN]
Task given to the agents: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
You have analysed a deliberation between 5 agents over 4 rounds and its vote. Here is everything you wrote, in order:
ROUNDS:
Round 0: All five agents converge on the same skeleton — exec mandate, baseline of the 31 incidents, a 4–5 level severity scale, separated IC/comms/scribe/SME roles, paid on-call, one paging tool, synthetic payment probes, timed status-page updates, mandatory blameless postmortems with tracked actions, pilot-then-waves, SOC 2 evidence as a by-product — so they differ mainly in staffing design, pace, and operational depth. P1 and P4 are the most granular (29–30 steps), P2 is the most payments-literate and the only one with day-1 interim controls, P3 is mid-weight with some internal inconsistencies, and P5 is the thinnest and least dated.
shared: **Command separated from debugging**: every plan defines IC / comms lead / scribe / SMEs, forbids the IC from typing, allows anyone to declare, and lets only the IC downgrade (P1 S6, P2 S5, P3 S4, P4 S7, P5 S3).
shared: **Pay for on-call and page only for owned code** as the explicit answer to the pushback, with a service-ownership catalog and criticality tiers as the routing foundation (P1 S4/S9, P2 S3/S6, P3 S4/S5, P4 S5/S18, P5 S4/S5).
shared: **Alert-quality gate plus customer-journey detection**: every page needs owner, runbook, severity and SLO linkage; synthetic end-to-end payment probes from both regions; consolidation of the six tools into one pager (P1 S10/S11/S12, P2 S7/S8, P3 S6/S9, P4 S13/S16/S20, P5 S5/S6/S10).
shared: **Mandatory postmortems with enforced action tracking**: SEV1/SEV2 always, 3–5 day drafts, single template, named individual owners, due dates in one tracker, escalation of overdue items, and SOC 2 evidence produced from the live process rather than reconstructed.
differences: **24x7 staffing design** is the biggest split: P1 S7/S8 uses a 25–35-person Duty IC pool plus tiered team obligations; P2 S5 collapses 28 teams into 8–12 domain rotations of ≥6; P3 S4 adds an SRE Tier 2 that absorbs infra pages and floats a managed overnight triage vendor; P4 S8 puts Layer C teams on business hours only with an EM night list; P5 S4 invokes "follow-the-sun between the two AWS regions", which does not create time-zone coverage.
differences: **Pace and interim cover**: P2 S2 installs interim IC staffing, one declaration path and a provisional severity guide within 7 days and triages the 53 open actions immediately; P4 S1 targets live in ~10 weeks and S30 freezes the process before the observation window; P1 runs design 1–6, pilot 7–12, rollout 13–24; P3 ends rollout at week 24 with a month-5 mock audit; P5 gives almost no dates and pilots only at step 11 of 14.
differences: **Payments/ledger rigour**: P2 S10 guards against split-brain, duplication and replay, requires reconciliation and backlog processing before "resolved", and separates mitigation from resolution; P1 adds ledger double-entry assurance (S11), a regulator matrix incl. NYDFS/FinCEN/sponsor banks (S16) and an automated SLA-credit workflow (S17); P4 S4 extracts notification clocks from contracts before setting severity triggers; P5 S7 defers regulator timing entirely to legal with no clock.
differences: **How failure and resistance are handled**: P1 S28 is a standing risk register with pre-committed fallbacks (volunteer shortfall, comp not approved, pruning a real alert); P4 S26 stops wave expansion if the 60/120-day pulse is red; P2 S17 requires exception records with owners and expiry dates; P3 and P5 have no comparable contingency mechanism.
Proposal 1: 29-step program run by a Head of Reliability with exec steering, starting from a forensic re-coding of all 31 incidents and the six-tool alert estate. Builds a service catalog with Tier 0–3, a five-level severity scale with auto-escalation on any ledger-touching incident, a 25–35-person Duty IC rotation plus tiered team on-call, and an on-call readiness bar that blocks a service from paging anyone until runbooks exist. Unusually complete on money and law: stipend ranges, FLSA/NY review, a regulator playbook, and an automated SLA-credit workflow targeting $1.3M → <$400k. Ends with a month-6 dry-run audit over 15 sampled incidents and a year-two maturity roadmap.
Proposal 2: 19 dense steps that put interim controls live in the first 7 days — one declaration path, staffed interim commanders, and triage of the 53 open historical actions — before any tool consolidation. Uses SEV0–SEV3, domain rotations of ≥6 instead of 28 fragile ones, compensation and fatigue rules (one week in six, paid recovery day, self-declared unfitness) and shadow-mode alerts for 7 days before they can page. Strongest on payments correctness: dual-control preserved during ledger recovery, split-brain/duplication guards, reconciliation required before "resolved". Milestones are day-anchored (7/14/30/60/90/120/180) with a month-7 mock audit and metric-gaming checks.
Proposal 3: 13 steps run by a new Incident Management Office and a 28-manager Council, starting with a baseline plus peer benchmarking against Stripe/Adyen. Distinctive is the three-tier staffing where a 6-person SRE Tier 2 absorbs Kubernetes/PostgreSQL pages and a Ledger Duty Officer joins any ledger-touching incident, explicitly to defuse the cross-team pager objection. Compensation is fully costed ($500/$250 weekly, $75–150 per page, $300 SEV1 bonus, ~$420K/yr). Weak spots: a per-team budget of 100 alerts/month contradicts the <600/month target, "56 certified ICs and all 260 engineers trained in 14 weeks" is optimistic, and step 12 funnels all eleven prior steps into one rollout.
Proposal 4: 30 short, opinionated steps that time-box design to four weeks and aim to be live in ~10 weeks so five-plus months of SOC 2 evidence accumulate. Organises coverage as Layer A (24 trained company ICs, 24x7), Layer B (critical path 24x7), Layer C (business hours plus an EM escalation list), with mechanical rotation rules and a page-load SLO of p50 ≤4 per night shift. Maps contractual and regulatory notification clocks (S4) before setting severity triggers, runs a forced noise burn-down with a weekly leaderboard, and gates wave expansion on a red on-call pulse survey. Closes with a mock auditor interview and a deliberate wording freeze for the observation window.
Proposal 5: 14 conventional steps covering the full brief: SEV1–SEV4, RACI roles, service on-call plus a central IC guild, stipends and TOIL, one paging platform, comms timings, mandatory postmortems, metrics and a month-6 readiness assessment. Correct in outline but the thinnest on specifics — no dates until months 6–8, no wave plan, no page budget, no risk register, and no targeted answer to the "other teams' code" objection beyond a survey. Two realism problems: a 5-minute SEV1 status-page post, and a 24x7 model built on "follow-the-sun between the two AWS regions" for a single New York organisation. The pilot lands at step 11 with training and full rollout after it, leaving little room to iterate before the audit.
Round 1: All five plans converged on a common architecture — 7-day interim command, forensic baseline, service catalog with tiers, 4–5 level severity, IC/Comms/Scribe/SME roles, a central command pool plus 8–12 domain rotations, paid on-call before mandatory nights, page budgets, money-path detection, one pager, 3/5/10-day postmortems with tracked actions, critical-path pilot, gated waves, month-6/7 audit rehearsal. The real remaining spread is in severity thresholds, compensation concreteness, and how much program scaffolding (policy v1, risk register, 90-day revision) each keeps.
differences: **Severity design.** P2 (S5) alone gives quantitative declaration guardrails (>10% payment failures for 5 min = SEV1; 1–10% = SEV2) plus a SEV0 crisis tier and no time-based auto-escalation; P5 (S5) uses five levels with aggressive clocks (SEV2 unresolved >1h becomes SEV1); P1/P3/P4 (S6) use SEV1–SEV4 with ledger-touch and 2-hour auto-escalation.
differences: **Compensation and load caps.** P1 (S9) and P5 (S9) publish dollar figures (~$800–1,200/primary week, $150/night page); P4 (S9) budgets $600k–1M with structure only; P2 (S8) and P3 (S7) defer amounts to HR in 14 days. Caps diverge: one primary week in six (P1, P2, P4) versus one in four (P5 S8; P3 S7 says "four or six depending on staffing").
differences: **Program scaffolding.** P1 keeps a standalone Policy v1 with exception register (S24), risk register (S31) and a 90-day inspect-and-adapt (S32); P3 keeps a risk register (S26). P4 dropped its own Policy v1 and 90-day revision; P2 and P5 have no contingency/risk step at all.
differences: **Interim protection and density.** P1 (S2), P2 (S2), P4 (S2) and P5 (S1) stand up interim command in 48h–7 days; P3 only lists "interim controls week 1" in S1 with no step behind it. On density, P4 is tightest at 24 steps, P1 heaviest at 32 with overlapping alert work (S10 standard vs S28 burn-down) and an 11-dependency join at S24.
influences: P2's 7-day interim minimum controls, retroactive pay for interim duty, triage of the 53 open actions and the daily 15-minute ops review (R0 S2) were taken by P1 (S2), P4 (S2) and P5 (S1); P3 left it as a single timeline bullet.
influences: P4's three-layer coverage and the "deal" fairness contract (R0 S8, S26) were taken by P1 (S8, S29), P3 (S8, S24) and P5 (S7–S8); P2 reached the same result via 8–12 responder domains (S7) without the layer naming.
influences: P1's forensic baseline, page budget, SLA-credit workflow, Incident Review Board and gated waves (R0 S2, S10, S17, S18, S25) were adopted almost wholesale by P3 (S2, S9, S14, S16, S22) and selectively by P4 (S3, S10, S16, S17, S22) and P2 (S3, S11, S15, S18, S21).
influences: P2's safety rails travelled widely: shadow-mode alerts and "never silently disable", two weeks of verified operation before retiring a legacy path, human acknowledgement required, mitigation-vs-resolution semantics and ledger dual-control (R0 S1, S4, S7, S8, S9) now appear in P1 (S6, S7, S10, S12), P4 (S7, S10, S12, S14), P3 (S9, S11) and P5 (S13).
influences: Nobody took P3's outsourced overnight triage vendor or its per-page pay table (R0 S4, S5) — P3 dropped both itself — and nobody kept P5's 5-minute SEV1 status-page deadline (R0 S7); all five settled on 15 minutes for the top availability tier.
Proposal 1 (improved): Closed its biggest R0 gap — six weeks of design with no protection — by adding a 7-day interim command bridge (S2). Replaced "14–18 teams at 24x7" with a Layer A/B/C model over 10–12 domains (S8), added live-execution doctrine (S14), Policy v1 (S24), a noise burn-down campaign (S28) and staged day-90/month-6 metric targets. Now the most complete plan, but at 32 steps it is also the heaviest.
Proposal 2 (improved): Grew from 19 to 24 steps by adding the pieces it was missing — readiness/runbook standard (S9), detection uplift (S12), pilot (S20), waves (S21), exercises (S22) — while keeping its distinctive strengths: compliance designed on day one (S4, before severity in S5), a balanced anti-gaming scorecard (S18), and now numeric severity guardrails.
Proposal 3 (improved): Doubled from 13 to 27 steps and adopted most of P1's R0 architecture: baseline, listening tour, layered staffing, escalation ladder, regulatory playbook, credits, runbooks, exercises, gated pilot and waves, risk register, SOC 2 dry-run. Far more complete and better sequenced, but largely derivative and it lost its own concrete compensation numbers.
Proposal 4 (improved): Consolidated 30 loose steps into 24 denser ones (postmortems and actions merged into S17, runbooks and readiness folded into S14, all comms into S15) and imported the financial and training machinery it was missing. Remains the most readable plan, but it dropped two of its own best R0 steps.
Proposal 5 (improved): Expanded from 14 thin steps to 26, closing nearly every R0 gap: interim duty officer, forensic baseline, listening tour, escalation ladder, regulatory playbook, credits, runbooks, exercises, pilot, waves and SOC 2 dry-run. The gains come almost entirely from copying P1's R0 structure and P2's SEV0, so it is now complete but the least differentiated plan.
Round 2: All five plans converged onto a near-common template: four-level severity, a central command corps plus 10–12 domain rotations plus business-hours Layer C, paid on-call gated before any mandatory night rotation, a two-page-per-week budget, one paging platform, 3/5/10-day blameless postmortems, wave rollout with gates and a month-7 mock audit. Real differentiation narrowed to severity anchoring (P4's integrity flag, P1's payment anchors), plan weight and audit sequencing, while P3 essentially re-issued P1's round-1 plan.
differences: **Severity anchoring.** P4 step 6 attaches a financial-integrity/security flag to any level (SEV0's forcing function without a fifth level); P1 step 6 adds payments anchors (value at risk per hour, settlement-deadline proximity, named accounts); P2 step 6 demotes percentages to guardrails; P3 step 7 and P5 step 6 stay purely qualitative. P1/P4 also tighten the SEV3 postmortem trigger (customer-detected, >2h, repeat), while P3/P5 keep "postmortem on request".
differences: **Audit sequencing.** P4 confirms the observation window in week 1 (step 1), P2 in step 3, P3 in step 5; P1 leaves it in step 30, which depends on steps 24/26/28, and cuts independent testing from two rounds to one at month 5. P2/P3/P5 keep tests in months 4 and 6.
differences: **Coverage economics.** P1 alone funds an overnight Duty Triage Desk owning the first ten minutes (step 8) but never sizes or budgets it; P4 alone shows the arithmetic (260 engineers sustain 10–12 rotations, not 28, step 8); P2 alone refuses to merge small teams for schedule convenience, scopes the six-responder rule to direct 24x7 rotations (steps 8, 22) and keeps a surge roster for concurrent incidents.
differences: **Weight and pace.** P1/P3 run 32 steps with standalone noise campaign, Policy v1 and risk register; P4 compresses to 26 with an 11-bullet comms mega-step (15) and the risk register demoted to last (26); P5 merges regulatory+credits (16) and postmortems+actions (17) and has no policy-publication step; P2 has 26 steps, no risk register at all, and the fastest schedule (Tier 0/1 by week 10, all teams by week 18 versus week 24 elsewhere).
influences: P2 and P5 abandoned their SEV0 tiers for the SEV1–SEV4 scale of P1/P3/P4 (P2 step 6, P5 step 6); P4 instead converted SEV0 into a financial-integrity/security flag (step 6) and P1 into payments anchors (step 6) — the round's one genuine design argument, resolved three different ways.
influences: P3 imported P1's round-1 plan almost wholesale (21 matching steps: Layer A/B/C 9, noise burn-down 26, Policy v1 24, risk register 31, 90-day inspect 32) and grafted on P2's day-one compliance step (P3 step 5 from P2 step 4) and P2's fake-redundancy check (P3 step 12 from P2 step 12).
influences: P5 took P1's Layer A/B/C coverage (step 8), readiness bar and ledger blast-radius workstream (step 10), risk register (step 25) and culture step (23), and tightened its own 1-in-4 rotation cap to P1's 1-in-6 (step 9).
influences: P2 took the standalone listening tour/fairness contract from P1 step 4, P3 step 3 and P4 step 4 (now P2 step 4) and concrete pay bands from P1/P5 step 9; in the other direction P4 took P2's week-1 observation-window confirmation (P4 step 1) and P1 added P2's 99.95% availability metric and an Internal Audit seat (step 1).
influences: Nobody took P5's named vendor shortlist (PagerDuty/incident.io/FireHydrant, round-1 step 12) — P5 dropped it itself; nobody adopted P2's numeric declaration thresholds (10% / 1–10% failure rates), which P2 also demoted; and only P2 (step 8) still plans for simultaneous incidents with a surge roster.
Proposal 1 (improved): Structure is unchanged at 32 steps; the gains are inside steps. Step 8 adds a paid overnight Duty Triage Desk, step 6 adds payments-specific severity anchors, step 11 adds an incident-replay test with "replay coverage" as a leading metric. Independent control testing shrank from two rounds to one.
Proposal 2 (improved): Renamed the scale to SEV1–SEV4, ending the SEV0/SEV1 ambiguity, and split seven previously buried topics (fairness contract, alert contract, customer comms, actions, policy, metrics, sustainability) into their own steps. Compensation moved from "fixed weekly stipends" to actual dollar bands and an annual budget.
Proposal 3 (improved): Effectively re-issues P1's round-1 plan: 21 of its 32 steps match P1 and only 5 match its own previous version. It gains the machinery it lacked (Policy v1, noise campaign, risk register, 90-day inspect) plus a day-one compliance step from P2, but contributes nothing new and leaves two internal inconsistencies.
Proposal 4 (improved): Compressed to 26 steps while adding the round's neatest severity idea and the only staffing arithmetic. Gains: integrity/security flag (step 6), 260-engineer rotation math (step 8), week-1 audit clock (step 1), postmortems split from action tracking. Losses: the change-management step disappeared and the risk register is demoted to the end.
Proposal 5 (improved): Dropped SEV0 for a four-level scale, adopted P1's Layer A/B/C model, metric set and risk register, and added the readiness bar, simulations and a culture step. Still 26 steps, but two of them are now overloaded and the policy-publication step is missing.
Round 3: Round 3 converged hard: all five plans now run the same skeleton (charter → 7-day floor → forensic baseline → fairness contract → catalog/tiers → severity → roles → three-layer 24x7 → paid on-call → alert budget → detection → single paging platform → escalation → comms → regulatory/credits → postmortems/actions → training → exercises → policy → pilot → gated waves → metrics → SOC 2 → 90-day inspect), with near-identical numbers ($1,000 primary week, 2 out-of-hours pages/responder/week, 15% reserved capacity, 30 ICs / 18 comms leads). Differentiation now comes from a few genuine additions — P1's change intelligence and Support-as-detection tier, P4's Duty Triage Desk carve-out and vendor-incident class, P2's defined noise/actionable metrics and severity modifiers — while P3 and P5 mostly merge material authored by P1 and P4.
differences: **Unique content**: only P1 has change intelligence/deployment safety (step 13), Support and account managers as an instrumented detection tier (step 20) and a concurrency doctrine with a Multi-Incident Coordinator (step 8); only P4 has a vendor-incident class for processor/bank/cloud failures (step 13) and a rule for who writes the 3 a.m. status page (step 8); only P2 defines "actionable page" and "noise" and uses five modifiers instead of flags (steps 12, 6). P3 and P5 contribute nothing the others do not have.
differences: **Overnight first line**: P1 (step 9), P4 (step 8) and P5 (step 8) fund a paid Duty Triage Desk; P4 adds the safety carve-out that it never holds suspected ledger, payment-halt or security pages. P2 (step 8) and P3 (step 9) refuse a triage desk and route unknown-owner pages to Duty Command plus platform, P2 additionally reclassifying lower-tier services that can cause severe overnight harm.
differences: **Schedule**: P2 (step 1) compresses design to weeks 1–4, pilot 5–10, rollout by week 18, control tests in months 4 and 6; P1, P3, P4 and P5 keep design weeks 1–6, pilot 7–12, rollout 13–24 and a month-5 internal test. P2 buys a longer observation window at the cost of a 4-week design for 180 services.
differences: **Failure handling and metric honesty**: P1 (step 34), P3 (step 28) and P5 (step 29) keep a standing monthly risk register; P4 compresses contingencies into bullets in step 26; P2 has none. On claims, P2 targets "no unresolved material exception" and P1 warns credits may rise before they fall (step 22), while P3 and P5 still promise "SOC 2 passes with zero exceptions".
influences: P1's replay test (R2 step 11) went everywhere: P2 step 13 plus a replay-coverage metric, P3 step 12, P4 steps 3 and 11 (as a "replay catalog" with a gap owner), P5 step 11. It is now the field's standard proof that detection actually improved.
influences: P4's severity flag (R2 step 6) was taken by P1 (step 7, financial-integrity/security flag forcing dual control and the reportability checkpoint) and generalised by P2 into five modifiers — FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, VENDOR (step 6). P3 and P5 declined it and kept four levels plus auto-escalation.
influences: P1's Duty Triage Desk (R2 step 8) was adopted by P4 (step 8) and P5 (step 8); P4 improved it with the carve-out that ledger, payment-halt and security events page commander and domains in parallel, plus a half-day triage-desk training module (step 18). P2 and P3 explicitly did not.
influences: P4's "confirm the SOC 2 observation window in week 1" (R2 step 1) was taken by P3 (step 1) and P5 (step 3), and expanded by P1 into a full week-one audit, legal and evidence scoping step (step 4) covering privilege, legal hold and the obligation matrix.
influences: P4's staffing arithmetic (R2 step 8 — 260 engineers cannot sustain 28 night rotas, never a two-person rotation) was taken verbatim by P3 (step 9) and restated by P1 (step 9, ~100 of 260 engineers carry a night obligation at one week in six). Nobody took P2's R2 caution against merging small teams for schedule convenience — all five consolidate into 10–12 domains.
Proposal 1 (improved): Four substantive new steps and three hardening additions, all addressing gaps the field had left open. It is now the most complete plan, at the cost of being the longest (35 steps).
Proposal 2 (improved): It decomposed its round-2 mega-steps into discrete, individually gated steps and added the sharpest measurement definitions in the field. It remains the only plan without a risk register.
Proposal 3 (improved): A competent merge: it absorbed P4's audit clock and staffing math, P1's replay test and P2's accommodations, and split the readiness bar out of the execution doctrine. It adds nothing the other plans do not already have, and it created two new mega-steps.
Proposal 4 (improved): It imported P1's Duty Triage Desk and replay test and then improved both, and added two things nobody else has: a vendor-incident class and a rule for who writes the status page overnight. It stays the most compact serious plan at 26 steps.
Proposal 5 (improved): It closed two real holes from its round-2 version — there was no policy-publication step and no separated postmortem/action or internal/customer comms steps — but it did so by copying P1's round-2 plan almost step for step. It contributes no idea of its own and picks up none of P1's round-3 advances.
OUTCOME (the final round, ranked blind):
- **P1** — the most complete plan (35 steps): the only one with change-intelligence/deploy-safety (step 13), Support and account managers as an instrumented detection tier (step 20), a concurrency doctrine with a Multi-Incident Coordinator (step 8) and a written warning that credits may rise before they fall (step 22).
- **P2** — the most rigorous on measurement and audit: defines "actionable page" and "noise" and counts *notification episodes* (step 12), five severity modifiers instead of a fifth level (step 6), two independent control tests in months 4 and 6, company-wide by week 18, and the only honest audit claim ("no unresolved material exception"); the only plan with no programme risk register.
- **P3** — a competent merge of P1's architecture with P2's day-one compliance step; complete but contributes nothing of its own and carries two internal inconsistencies (month-5 vs months-4-and-6 testing; "five months before the report date" against a week-24 rollout).
- **P4** — the most compact serious plan (26 steps) with the two sharpest operational rules in the field: a Duty Triage Desk that never holds ledger/payment-halt/security pages (step 8) and a named owner for the 3 a.m. status page; weakened by bundling (steps 14, 23) and by burying contingencies in the final step.
- **P5** — complete and correct but almost entirely derivative of P1's earlier draft, with two overloaded steps (14, 17) and a metric that contradicts its own day-7 floor.
- **All five share**: day-7 operating floor, forensic re-coding of the 31 incidents, written fairness contract, service catalog with Tier 0–3, SEV1–SEV4 with ledger auto-escalation, IC/comms/scribe/SME with dual control preserved, ~30-person command corps plus 10–12 domain rotations plus business-hours Layer C, ~$1,000 primary week with a payroll hard gate, two out-of-hours pages/responder/week budget, synthetic money-path probes plus the 31-incident replay test, one paging platform, 15/30-minute status-page clocks, 3/5/10-day blameless postmortems with verified action closure, pilot then gated waves, month-7 mock audit, 90-day inspect-and-adapt.
Ranking: 1, 2, 4, 3, 5
**P1 first**: it matches every rival on the shared skeleton and then adds three things that bear directly on this company's numbers — change correlation (step 13) against a 3 h 10 min MTTM, Support/AMs as a measured detection tier (step 20) against 40% customer-first detection, and a concurrency doctrine (step 8) that a 30-person command corps facing 31 incidents a year will need. It also has the earliest standalone risk register (step 34) and the deepest week-one legal/evidence scoping (step 4). The price is 35 steps, one control test instead of two, and a "zero exceptions" claim it contradicts itself.
**P2 second**: the most disciplined plan — it alone defines what a page and noise are before promising to cut them, tests controls twice (months 4 and 6), finishes company-wide by week 18 for a longer observation window, and refuses to hide overnight risk behind a Layer C manager list. It loses to P1 on two counts: no programme risk register at all, and a four-week design window for 180 services that leaves no slack.
**P4 third**: equal coverage in 26 steps and the two sharpest operational rules in the field (the triage-desk carve-out, the overnight status-page owner) plus the vendor-incident class. It falls below P2 because contingencies sit in the last step, and steps 14 and 23 bundle readiness, playbooks, ledger architecture, noise burn-down and waves into single units that are harder to gate.
**P3 fourth**: complete and well written but purely a merge, with two dating inconsistencies (month-5 vs months-4-and-6 testing; "five months before the report date" against a week-24 finish) and two mega-steps covering escalation/execution and all communications.
**P5 last**: the same content as P3 with less of it original, an extra overloaded step (14), and a headline metric that contradicts its own day-7 floor. Nothing in it is wrong; nothing in it is its own.
Better than every initial proposal: True
- **Gained an operating floor**: every final plan protects the company from day 7 (interim commander roster, one declaration path, Support authorised to declare, 53 legacy actions triaged) instead of leaving six-plus weeks of design unguarded — P2's round-0 idea that spread to all five.
- **Gained staffing realism**: "28 teams on 24x7" was replaced everywhere by a ~30-person command corps plus 10–12 domain rotations plus business-hours Layer C, with published arithmetic and a ban on two-person rotations — the direct, defensible answer to the pager pushback.
- **Gained payments correctness**: mitigate-vs-resolve semantics, reconciliation and controlled backlog drain before closure, split-brain/replay/duplication guards, and preservation of dual control during ledger recovery.
- **Gained proof rather than assertion on detection**: the 31-incident replay test with a named detector and expected detection minute became the field's standard evidence that MTTD work actually worked.
- **Gained safety rails on noise cutting**: shadow mode for seven days, never silently disable, compensating detection before suppression, monthly missed-detection reviews — so the 3,400 → 500 target cannot be met by going blind.
- **Gained honest instrumentation**: medians and p90, detection source on every incident, and monthly reconciliation of incident counts against support tickets, credits and status-page history to catch under-reporting.
- **Gained enforceable money and law**: dollar pay bands with a payroll hard gate, FLSA/New York review, a reportability checkpoint recorded even when "not reportable", and an automated SLA-credit workflow tied to root-cause families.
- **Lost differentiation**: P3 and P5 converged into near-copies of P1's earlier drafts, so five agents produced roughly three distinct plans.
- **Lost some good ideas without a decision**: P2's numeric declaration thresholds (>10% failures = SEV1), its caution against merging small teams into domains for schedule convenience, and P3's explicit per-page pay table were all dropped rather than argued out.
- **Carried an over-claim to the end**: four of five still promise "SOC 2 passes with zero exceptions"; only P2 corrected it to "no unresolved material exception", which is what an auditor can actually deliver.
- **Grew heavier**: step counts rose from 13–30 to 26–35, with several plans creating mega-steps (execution doctrine plus readiness bar plus playbooks plus ledger architecture in one) that are harder to gate than the pieces they replaced.
PROCESS:
- **Convergence: genuine in round 0, imitation thereafter.** Five agents independently landing on the same skeleton (exec mandate → baseline → severity → separated roles → paid on-call → one pager → synthetics → timed comms → postmortems → pilot → waves → audit evidence) is real evidence the skeleton is right. What followed was not argument but transcription: P3 imported 21 of 32 steps from P1's round-1 plan; P5 copied P1's round-2 plan step for step; by round 3 the success-metric lists of P1, P3, P4 and P5 are near-verbatim (same "$1,000 primary week", "2 out-of-hours pages per responder", "30 certified commanders and 18 comms leads", "10–12 domain rotations", "15% reserved capacity"). Two exceptions were earned: P4 improved P1's Duty Triage Desk by carving out ledger/payment-halt/security pages so triage never sits on a crisis, and P2 generalised P4's integrity flag into five modifiers (FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, VENDOR).
- **Almost no criticism; disagreements were resolved by fiat, not by evidence.** The one real design argument — SEV0 tier vs. four levels vs. severity flags vs. modifiers — ended with three different unreconciled answers and nobody testing them against the 31 historical incidents that every plan claims to have re-coded. P2's warning against merging small teams into domains purely for schedule convenience was ignored rather than rebutted; all five consolidated into 10–12 domains anyway. P1's Duty Triage Desk was copied by P4 and P5 and refused by P2 and P3 — but the refusal was never argued, and the desk is still unsized and unbudgeted in P1 and P5 (only P4 prices it at $400–800/week). Nobody attacked P1's 11-dependency join at step 27, the growth to 35 steps, or the arithmetic of a three-person programme office delivering all of this in 24 weeks.
- **Premises were barely questioned, and the best challenges did not spread.** Three genuine ones: P1 step 22 warns credits may *rise* before they fall because better detection surfaces previously unbilled outages; P2 refuses "zero exceptions" and targets "no unresolved material exception"; P4 does the staffing arithmetic showing 260 engineers cannot sustain 28 night rotas. Untouched premises: whether 99.95% is achievable at all on a single shared PostgreSQL ledger (relegated to a parallel "blast-radius workstream" in every plan); whether the business case survives the 15% capacity reservation (~39 FTE of 260, far larger than the $1.3M in credits, and never costed by anyone); whether hiring a dedicated 10–15 person SRE/NOC cohort beats putting ~100 of 260 engineers on nights; and the flat contradiction between "you are paged only for services your team owns" and Layer B consolidation of 28 teams into 10–12 shared domains — only P2 (step 8, "explicit support acceptance") resolves it, and nobody flagged it in the others.
- **Losses: options, precision and independent judgement.** Dropped without argument: P3's managed overnight triage vendor (a live option for a single-timezone company); P2's numeric declaration thresholds (>10% payment failures for 5 min = SEV1), which were the only thing making severity a genuine lookup rather than a judgement call; P5's vendor shortlist; P3's fully derived $420K compensation budget, replaced by an undefended $0.9–1.3M. Regressions propagated as readily as improvements: P1 cut independent control testing from months 4 and 6 to month 5 alone, and P4 and P5 copied the weaker version. Unique good ideas failed to spread because copying flowed one way (P1 → P3/P5): only P1 has change intelligence/deployment safety (step 13) and Support as an instrumented detection tier (step 20), and only P1 and P2 plan for concurrent incidents at all. Most of all, the field lost the ability to discriminate: the vote was between one plan wearing five costumes.
Problems: **Two of five agents became pure copyists.** P3 and P5 contributed no idea the others lacked from round 2 onward, wasting 40% of the parallelism and inflating apparent consensus.; **Success metrics were copied verbatim across plans**, so they stopped being evidence of design quality and became boilerplate — and the boilerplate carried an overclaim ("SOC 2 Type II incident-response controls pass with zero exceptions" in P1, P3, P4, P5) that the company does not control. Only P2 phrased it defensibly.; **Step inflation rewarded length over executability.** P1 29→32→35, P5 14→26→30, P3 13→27→29. Nothing forced removal, so P1 now carries alert quality (11), a separate noise burn-down (29) and a per-wave noise quota (31) covering the same work, and comms split across steps 18/19/20.; **Big dependency joins went unchallenged.** P1 step 27 (Policy v1) has 11 predecessors, P5 step 22 has 11, P3 step 22 has 10 — a single-point schedule risk in the middle of the plan that no reviewer named.; **No plan was ever costed end-to-end.** Every plan reserves 10–15% of 260 engineers plus $0.9–1.3M compensation plus $150–250k tooling plus a three-person office, and every plan anchors the case on $1.3M of credits. Nobody did the subtraction.; **Feasibility of the timeline was asserted, never tested.** Design in 6 weeks for 180 services, 30 certified commanders, synthetics in both regions, six-tool migration, all 28 teams by week 24. P2's alternative (design weeks 1–4, rollout by week 18) was never debated on the merits, only noted as different.; **The severity design argument was decided without data.** Every plan claims to re-code the 31 incidents and to publish 12 worked examples from them, yet nobody used those incidents to show which severity model would have classified them correctly.; **Contradiction survived all rounds unexamined:** the fairness promise "paged only for what your team owns" versus Layer B domain rotations that necessarily merge several of the 28 teams. Only P2 wrote the reconciliation (formally accepted support agreements, training, access).; **Regressions propagated as fast as improvements** — the cut from two independent control tests (months 4 and 6) to one (month 5) spread from P1 to P4 and P5 with no discussion.; **The final vote had almost nothing to discriminate on**, since the five plans share their architecture, numbers, metrics and most step text; preference collapsed onto length and two or three unique steps.
Suggestions: **Assign adversarial roles from round 1.** One agent must produce a minimum-viable plan (≤12 steps, must still pass the audit); one must attack cost and feasibility; one must defend a structurally different coverage model (managed overnight NOC, or hiring a dedicated 12-person SRE first line instead of putting 100 engineers on nights). Forbid these agents from converging.; **Require a written critique artefact each round**: every agent names three specific flaws by proposal and step number, states what it rejected from others and why, and those rejection reasons are carried into the final round as a visible dissent register. Dropping an option silently should be disallowed.; **Ban shared success metrics.** Each plan's metrics must be derivable from its own design choices, capped at ~10, each with a falsification condition and an owner. Penalise any metric the organisation cannot control ("audit passes with zero exceptions").; **Impose a step budget with a removal quota**: to add a step you must delete or merge one, and any step with more than four dependencies must be split or justified. Reward net simplification in scoring.; **Require one costing table per plan**: rotation count × stipend, triage-desk headcount and pay, tooling, programme office, and the opportunity cost of reserved capacity in FTE — set against $1.3M credits plus the quantified true incident cost. No plan advances without it.; **Force premise challenges.** Each agent must list two brief assumptions that may be wrong and say what it would do if they are (e.g., 99.95% is unachievable on a shared ledger without re-architecture; credits will rise before they fall; pager pay does not fix attrition).; **Run a resource-feasibility check as a gate**: a week-by-week load plan for the programme office and a named critical path, plus a written statement of what gets cut first if the mid-rollout SEV1 contingency fires.; **Add two evaluator personas before the vote** — a SOC 2 auditor and a sceptical senior engineer from one of the 16 teams with no on-call — and require every plan to answer their written objections in its final round.; **Score the final round on discriminating dimensions separately** (staffing arithmetic, cost, severity calibration evidence, contingency handling, readability under pressure) rather than an overall preference, so near-identical plans are forced to state trade-offs.; **Detect and penalise derivativeness automatically**: measure step-level overlap with other proposals and require any agent above a threshold to either justify each copied step or be replaced with a fresh brief.
THE VOTE, CONFRONTED WITH YOUR RANKING:
The analyst agrees with the vote. The vote matches my first choice, and largely for the same reasons I gave: change intelligence against the 3h10 MTTM, Support/AMs as an instrumented detection tier against 40% customer-first detection, the concurrency doctrine, and the paid Duty Triage Desk with published arithmetic answering the pager objection. Voter 1 adds a merit I weighed less — the SLA-credit incentive guardrail — and voters 3 and 5 lean partly on step count ("35 steps", "most comprehensive"), which hints at a length bias, since none of them flags P1's single control test, its self-contradictory "zero exceptions" claim, or P2's superior discipline in defining a page and noise before promising to cut them. Voter 2's case for P2 is exactly the runner-up argument I made (feasible paid domain-based coverage, suppression safeguards, evidence-producing rollout), so the top two are ordered as I had them. No voter engaged with P3/P4/P5, so the vote says nothing about the derivative-merge problem I penalised in P3 and P5.
Now write the brief a busy reader will read first, and often only: "brief", a Markdown list of at most eight items, each one sentence or two, covering in this order: what the outcome is (your first choice, the vote, whether the plan beats the initial proposals); what happened in the rounds (genuine convergence or imitation, criticism or copying, who drove it); the two or three problems that matter most; the two or three changes that would help most. Name proposals and steps by number. Nothing new: only what your analysis already says.
[ROUND 0]
{"round_summary": "All five agents converge on the same skeleton — exec mandate, baseline of the 31 incidents, a 4–5 level severity scale, separated IC/comms/scribe/SME roles, paid on-call, one paging tool, synthetic payment probes, timed status-page updates, mandatory blameless postmortems with tracked actions, pilot-then-waves, SOC 2 evidence as a by-product — so they differ mainly in staffing design, pace, and operational depth. P1 and P4 are the most granular (29–30 steps), P2 is the most payments-literate and the only one with day-1 interim controls, P3 is mid-weight with some internal inconsistencies, and P5 is the thinnest and least dated.", "shared": ["**Command separated from debugging**: every plan defines IC / comms lead / scribe / SMEs, forbids the IC from typing, allows anyone to declare, and lets only the IC downgrade (P1 S6, P2 S5, P3 S4, P4 S7, P5 S3).", "**Pay for on-call and page only for owned code** as the explicit answer to the pushback, with a service-ownership catalog and criticality tiers as the routing foundation (P1 S4/S9, P2 S3/S6, P3 S4/S5, P4 S5/S18, P5 S4/S5).", "**Alert-quality gate plus customer-journey detection**: every page needs owner, runbook, severity and SLO linkage; synthetic end-to-end payment probes from both regions; consolidation of the six tools into one pager (P1 S10/S11/S12, P2 S7/S8, P3 S6/S9, P4 S13/S16/S20, P5 S5/S6/S10).", "**Mandatory postmortems with enforced action tracking**: SEV1/SEV2 always, 3–5 day drafts, single template, named individual owners, due dates in one tracker, escalation of overdue items, and SOC 2 evidence produced from the live process rather than reconstructed."], "differences": ["**24x7 staffing design** is the biggest split: P1 S7/S8 uses a 25–35-person Duty IC pool plus tiered team obligations; P2 S5 collapses 28 teams into 8–12 domain rotations of ≥6; P3 S4 adds an SRE Tier 2 that absorbs infra pages and floats a managed overnight triage vendor; P4 S8 puts Layer C teams on business hours only with an EM night list; P5 S4 invokes \"follow-the-sun between the two AWS regions\", which does not create time-zone coverage.", "**Pace and interim cover**: P2 S2 installs interim IC staffing, one declaration path and a provisional severity guide within 7 days and triages the 53 open actions immediately; P4 S1 targets live in ~10 weeks and S30 freezes the process before the observation window; P1 runs design 1–6, pilot 7–12, rollout 13–24; P3 ends rollout at week 24 with a month-5 mock audit; P5 gives almost no dates and pilots only at step 11 of 14.", "**Payments/ledger rigour**: P2 S10 guards against split-brain, duplication and replay, requires reconciliation and backlog processing before \"resolved\", and separates mitigation from resolution; P1 adds ledger double-entry assurance (S11), a regulator matrix incl. NYDFS/FinCEN/sponsor banks (S16) and an automated SLA-credit workflow (S17); P4 S4 extracts notification clocks from contracts before setting severity triggers; P5 S7 defers regulator timing entirely to legal with no clock.", "**How failure and resistance are handled**: P1 S28 is a standing risk register with pre-committed fallbacks (volunteer shortfall, comp not approved, pruning a real alert); P4 S26 stops wave expansion if the 60/120-day pulse is red; P2 S17 requires exception records with owners and expiry dates; P3 and P5 have no comparable contingency mechanism."], "proposals": [{"proposal": 1, "summary": "29-step program run by a Head of Reliability with exec steering, starting from a forensic re-coding of all 31 incidents and the six-tool alert estate. Builds a service catalog with Tier 0–3, a five-level severity scale with auto-escalation on any ledger-touching incident, a 25–35-person Duty IC rotation plus tiered team on-call, and an on-call readiness bar that blocks a service from paging anyone until runbooks exist. Unusually complete on money and law: stipend ranges, FLSA/NY review, a regulator playbook, and an automated SLA-credit workflow targeting $1.3M → <$400k. Ends with a month-6 dry-run audit over 15 sampled incidents and a year-two maturity roadmap.", "approach": "Exhaustive reliability-program blueprint with cost, legal and audit lines"}, {"proposal": 2, "summary": "19 dense steps that put interim controls live in the first 7 days — one declaration path, staffed interim commanders, and triage of the 53 open historical actions — before any tool consolidation. Uses SEV0–SEV3, domain rotations of ≥6 instead of 28 fragile ones, compensation and fatigue rules (one week in six, paid recovery day, self-declared unfitness) and shadow-mode alerts for 7 days before they can page. Strongest on payments correctness: dual-control preserved during ledger recovery, split-brain/duplication guards, reconciliation required before \"resolved\". Milestones are day-anchored (7/14/30/60/90/120/180) with a month-7 mock audit and metric-gaming checks.", "approach": "Operate-first, day-anchored, with payments-grade safety rules"}, {"proposal": 3, "summary": "13 steps run by a new Incident Management Office and a 28-manager Council, starting with a baseline plus peer benchmarking against Stripe/Adyen. Distinctive is the three-tier staffing where a 6-person SRE Tier 2 absorbs Kubernetes/PostgreSQL pages and a Ledger Duty Officer joins any ledger-touching incident, explicitly to defuse the cross-team pager objection. Compensation is fully costed ($500/$250 weekly, $75–150 per page, $300 SEV1 bonus, ~$420K/yr). Weak spots: a per-team budget of 100 alerts/month contradicts the <600/month target, \"56 certified ICs and all 260 engineers trained in 14 weeks\" is optimistic, and step 12 funnels all eleven prior steps into one rollout.", "approach": "Governance body plus SRE tier that absorbs shared-infrastructure pages"}, {"proposal": 4, "summary": "30 short, opinionated steps that time-box design to four weeks and aim to be live in ~10 weeks so five-plus months of SOC 2 evidence accumulate. Organises coverage as Layer A (24 trained company ICs, 24x7), Layer B (critical path 24x7), Layer C (business hours plus an EM escalation list), with mechanical rotation rules and a page-load SLO of p50 ≤4 per night shift. Maps contractual and regulatory notification clocks (S4) before setting severity triggers, runs a forced noise burn-down with a weekly leaderboard, and gates wave expansion on a red on-call pulse survey. Closes with a mock auditor interview and a deliberate wording freeze for the observation window.", "approach": "Ship v1 fast, three coverage layers, audit-window engineering"}, {"proposal": 5, "summary": "14 conventional steps covering the full brief: SEV1–SEV4, RACI roles, service on-call plus a central IC guild, stipends and TOIL, one paging platform, comms timings, mandatory postmortems, metrics and a month-6 readiness assessment. Correct in outline but the thinnest on specifics — no dates until months 6–8, no wave plan, no page budget, no risk register, and no targeted answer to the \"other teams' code\" objection beyond a survey. Two realism problems: a 5-minute SEV1 status-page post, and a 24x7 model built on \"follow-the-sun between the two AWS regions\" for a single New York organisation. The pilot lands at step 11 with training and full rollout after it, leaving little room to iterate before the audit.", "approach": "Standard incident-management checklist, light on dates and mechanics"}]}
[ROUND 1]
{"round_summary": "All five plans converged on a common architecture — 7-day interim command, forensic baseline, service catalog with tiers, 4–5 level severity, IC/Comms/Scribe/SME roles, a central command pool plus 8–12 domain rotations, paid on-call before mandatory nights, page budgets, money-path detection, one pager, 3/5/10-day postmortems with tracked actions, critical-path pilot, gated waves, month-6/7 audit rehearsal. The real remaining spread is in severity thresholds, compensation concreteness, and how much program scaffolding (policy v1, risk register, 90-day revision) each keeps.", "converging": true, "differences": ["**Severity design.** P2 (S5) alone gives quantitative declaration guardrails (>10% payment failures for 5 min = SEV1; 1–10% = SEV2) plus a SEV0 crisis tier and no time-based auto-escalation; P5 (S5) uses five levels with aggressive clocks (SEV2 unresolved >1h becomes SEV1); P1/P3/P4 (S6) use SEV1–SEV4 with ledger-touch and 2-hour auto-escalation.", "**Compensation and load caps.** P1 (S9) and P5 (S9) publish dollar figures (~$800–1,200/primary week, $150/night page); P4 (S9) budgets $600k–1M with structure only; P2 (S8) and P3 (S7) defer amounts to HR in 14 days. Caps diverge: one primary week in six (P1, P2, P4) versus one in four (P5 S8; P3 S7 says \"four or six depending on staffing\").", "**Program scaffolding.** P1 keeps a standalone Policy v1 with exception register (S24), risk register (S31) and a 90-day inspect-and-adapt (S32); P3 keeps a risk register (S26). P4 dropped its own Policy v1 and 90-day revision; P2 and P5 have no contingency/risk step at all.", "**Interim protection and density.** P1 (S2), P2 (S2), P4 (S2) and P5 (S1) stand up interim command in 48h–7 days; P3 only lists \"interim controls week 1\" in S1 with no step behind it. On density, P4 is tightest at 24 steps, P1 heaviest at 32 with overlapping alert work (S10 standard vs S28 burn-down) and an 11-dependency join at S24."], "influences": ["P2's 7-day interim minimum controls, retroactive pay for interim duty, triage of the 53 open actions and the daily 15-minute ops review (R0 S2) were taken by P1 (S2), P4 (S2) and P5 (S1); P3 left it as a single timeline bullet.", "P4's three-layer coverage and the \"deal\" fairness contract (R0 S8, S26) were taken by P1 (S8, S29), P3 (S8, S24) and P5 (S7–S8); P2 reached the same result via 8–12 responder domains (S7) without the layer naming.", "P1's forensic baseline, page budget, SLA-credit workflow, Incident Review Board and gated waves (R0 S2, S10, S17, S18, S25) were adopted almost wholesale by P3 (S2, S9, S14, S16, S22) and selectively by P4 (S3, S10, S16, S17, S22) and P2 (S3, S11, S15, S18, S21).", "P2's safety rails travelled widely: shadow-mode alerts and \"never silently disable\", two weeks of verified operation before retiring a legacy path, human acknowledgement required, mitigation-vs-resolution semantics and ledger dual-control (R0 S1, S4, S7, S8, S9) now appear in P1 (S6, S7, S10, S12), P4 (S7, S10, S12, S14), P3 (S9, S11) and P5 (S13).", "Nobody took P3's outsourced overnight triage vendor or its per-page pay table (R0 S4, S5) — P3 dropped both itself — and nobody kept P5's 5-minute SEV1 status-page deadline (R0 S7); all five settled on 15 minutes for the top availability tier."], "proposals": [{"proposal": 1, "assessment": "improved", "what_changed": "Closed its biggest R0 gap — six weeks of design with no protection — by adding a 7-day interim command bridge (S2). Replaced \"14–18 teams at 24x7\" with a Layer A/B/C model over 10–12 domains (S8), added live-execution doctrine (S14), Policy v1 (S24), a noise burn-down campaign (S28) and staged day-90/month-6 metric targets. Now the most complete plan, but at 32 steps it is also the heaviest.", "improvements": ["S2 interim bridge: named commander within 10 min, Support escalates without engineering sign-off, 53 legacy actions triaged, daily 15-min review.", "S6 adds a lifecycle (Detected→Reviewed) with **mitigation = customer impact ends, resolution = backlog processed and ledger reconciled**, plus detection clock starting at first reliable impact signal.", "S8 consolidates 28 teams into 10–12 domain rotations with a ~30-person command corps, directly answering the pager objection instead of asserting it away.", "S14 payments-specific execution guards (split-brain, replay, duplication, controlled backlog drain) that R0 lacked.", "Metrics now staged (MTTD <10 min day 90, <5 min month 6; pages 3,400→1,500 by day 90→<500 by month 6) rather than single endpoints.", "S26 anti-gaming: monthly reconciliation of incident counts against support tickets, credits and status-page history.", "S30 adds independent internal control testing in months 4 and 6 before the month-7 mock audit."], "regressions": ["32 steps with overlap: S10 (alert standard) and S28 (burn-down campaign) cover the same ground, as do S14 playbooks and S21 runbooks.", "S24 Policy v1 depends on 11 steps and gates the pilot (S25) — a fragile critical path late in the schedule.", "Internal inconsistency: S27 claims teams finish \"at least five months before the audit report date\" while the S1 timeline ends rollout at week 24 against a month-8 audit (~2.5 months).", "Metric list has grown to 21 items; no prioritisation of which three the CEO sees."], "taken": [{"from_proposal": 2, "steps": [2], "what": "Interim minimum controls within seven days, compensated retroactively, with Support empowered to escalate.", "why": "Became S2, removing the unprotected design window that was R0's clearest flaw."}, {"from_proposal": 2, "steps": [4], "what": "Impact-based lifecycle with mitigation and resolution defined separately, resolution requiring reconciliation.", "why": "Folded into S6 so the MTTM metric cannot be gamed by calling an unreconciled ledger 'resolved'."}, {"from_proposal": 2, "steps": [1, 5], "what": "Preserve dual control and privileged-access rules during incidents; group teams into 8–12 coherent domains.", "why": "S7 bars the commander from bypassing financial controls; S8 builds 10–12 domain rotations instead of 28."}, {"from_proposal": 2, "steps": [8, 13], "what": "Shadow mode for new alerts, never silently disable, and reconcile dashboards against customer cases to catch gaming.", "why": "S10 adds compensating-detection checks and missed-detection reviews; S26 adds the reconciliation."}, {"from_proposal": 4, "steps": [14, 23, 29], "what": "Publish a short Policy v1 with laminated cards, run noise reduction as a quota campaign with a leaderboard, and revise the process at 90 days on data.", "why": "Became S24, S28 and S32 respectively, adding the artefacts and the correction loop R0 lacked."}, {"from_proposal": 4, "steps": [16], "what": "Hard date after which pages outside the chosen tool create no on-call obligation.", "why": "Added to S12 as the enforcement mechanism for tool consolidation."}], "rejected": [{"from_proposal": 2, "steps": [4], "what": "A distinct SEV0 crisis tier above SEV1.", "why": "S6 keeps four levels and labels SEV1 itself 'crisis', covering money-movement and ledger-integrity events there."}, {"from_proposal": 3, "steps": [5], "what": "All 28 teams staffed with 24x7 primary and secondary on-call.", "why": "S8 Layer C explicitly gives Tier 2/3 teams business-hours on-call with a manager escalation list and no night pager."}, {"from_proposal": 5, "steps": [7], "what": "Status page within 5 minutes of SEV1 declaration.", "why": "S16 keeps 15 minutes for SEV1 and 30 for SEV2, trading speed for pre-approved template accuracy."}]}, {"proposal": 2, "assessment": "improved", "what_changed": "Grew from 19 to 24 steps by adding the pieces it was missing — readiness/runbook standard (S9), detection uplift (S12), pilot (S20), waves (S21), exercises (S22) — while keeping its distinctive strengths: compliance designed on day one (S4, before severity in S5), a balanced anti-gaming scorecard (S18), and now numeric severity guardrails.", "improvements": ["S5 is the round's most operational severity definition: >10% payment failures for 5 minutes = SEV1, 1–10% = SEV2, with the explicit caveat that thresholds are guardrails, not permission to under-classify integrity or settlement risk.", "S4 sequences evidence and retention design before the operating process is finalised, so the audit trail is not retrofitted.", "S7 schedules an Engineering Director as 24x7 Executive Duty Officer — the only plan to staff that role on a roster.", "S9 readiness gate links every paging alert to the exact runbook step expected of the responder.", "S17 adds the escalation ladder (manager +7, director +14, CTO +30) and effectiveness verification the R0 postmortem step lacked.", "Most honest paging target in the round: 3,400 → 1,500 by day 90 → 700 by month 6, rather than a uniform <500."], "regressions": ["Dropped the risk register from R0 S1; nothing in R1 names failure modes or pre-commits contingencies.", "Dropped R0 S15's explicit prohibition on contractors or a managed service as sole incident commander or ledger owner — a useful guardrail.", "Still no dollar amounts or budget envelope for compensation (S8 only defines structure and a 14-day publication deadline), making the CFO ask unquantified.", "28 success metrics with no headline set; the executive view is buried.", "No short field-usable policy card equivalent to P1 S24 / P4 R0 S14 — policies exist only as versioned documents in S4."], "taken": [{"from_proposal": 1, "steps": [20], "what": "Service readiness bar and major-incident playbooks as a precondition for paging.", "why": "Became S9, including prevention of new paging alerts for services failing readiness review."}, {"from_proposal": 1, "steps": [19], "what": "Action escalation ladder with capacity reservation and residual-risk acceptance.", "why": "S17 adopts manager/director/CTO escalation at 7/14/30 days and 10–15% reserved capacity."}, {"from_proposal": 1, "steps": [17], "what": "Automated SLA credit computation from the incident record, proactive versus claims-based posture.", "why": "Folded into S15 so Finance produces affected minutes and credit exposure within five business days."}, {"from_proposal": 1, "steps": [23, 25], "what": "Critical-path pilot with numeric exit gates and gated waves by criticality.", "why": "S20 requires 95% command within 5 minutes and 50% noise reduction before expansion; S21 adds per-wave readiness gates to its date-based schedule."}, {"from_proposal": 4, "steps": [5], "what": "Publish the compensation package before any new rotation starts.", "why": "Hardened into a success metric: no mandatory night rotation begins before compensation, training, access and staffing are active."}], "rejected": [{"from_proposal": 1, "steps": [5], "what": "Time-based auto-escalation (SEV3 open >2h becomes SEV2).", "why": "S5 relies on impact guardrails plus 'if impact is materially unknown after 15 minutes, increase posture', avoiding clock-driven severity inflation."}, {"from_proposal": 1, "steps": [9], "what": "Publishing indicative stipend dollar amounts in the plan.", "why": "S8 specifies only the structure and delegates amounts to HR, Legal, Finance and Payroll within 14 days."}, {"from_proposal": 3, "steps": [6], "what": "Automated email/SMS notification to all 2,100 customers for SEV1 and SEV2.", "why": "S14 routes everything through the status page plus account-manager briefs, and limits direct notice to contractually required cases."}]}, {"proposal": 3, "assessment": "improved", "what_changed": "Doubled from 13 to 27 steps and adopted most of P1's R0 architecture: baseline, listening tour, layered staffing, escalation ladder, regulatory playbook, credits, runbooks, exercises, gated pilot and waves, risk register, SOC 2 dry-run. Far more complete and better sequenced, but largely derivative and it lost its own concrete compensation numbers.", "improvements": ["S2 forensic baseline with a specific case review of the two unowned incidents and a freeze of the reference metrics.", "S8 replaces the R0 'all 28 teams on rotation' stance with a layered model over 8–12 domains and a six-person minimum.", "S15 regulatory playbook now requires the reportability decision to be recorded even when the answer is 'not reportable'.", "S18 readiness bar blocks night paging for services without sign-off unless the EM accepts the gap in writing.", "S26 risk register with pre-committed contingencies (commander shortfall, comp delay, tool slip, pruning-caused misses, SEV1 during rollout).", "Metrics are now staged (day 90 / month 6 / month 9) rather than single 6–12 month endpoints."], "regressions": ["Lost the R0 dollar table ($500/$250 stipends, $75/$150 page pay, $300 SEV1 bonus, ~$420k/yr budget) — S7 now only says 'weekly stipend differentiated by tier', weakening the CFO case that R0 made well.", "No interim command step: S1 lists 'interim controls week 1' in the timeline but nothing defines them, while four rivals staff an interim commander in 48 hours to 7 days.", "S7 contradicts itself on load: 'no more than one primary week in four or one week in six depending on staffing'.", "S21 pilot depends on S7, S10, S11, S14, S18, S19 but not on the escalation path (S12) or postmortem standard (S16) it claims to exercise.", "Dropped R0's peer benchmarking against Stripe/Adyen-class platforms and the explicit tooling budget line."], "taken": [{"from_proposal": 1, "steps": [2, 3, 4], "what": "Forensic incident re-coding, resistance mapping with a co-design group, and a machine-readable ownership catalog with tiers.", "why": "Became S2, S3 and S4 and now anchor every downstream design decision."}, {"from_proposal": 1, "steps": [10, 17, 19, 20], "what": "Page budget, SLA credit workflow, action escalation ladder with reserved capacity, and the readiness bar.", "why": "Became S9, S14, S17 and S18 — the enforcement machinery R0 lacked."}, {"from_proposal": 1, "steps": [23, 25, 28], "what": "Critical-path pilot with published exit gates, gated waves with coaches, and a program risk register.", "why": "Became S21, S22 and S26, replacing R0's calendar-driven three-phase rollout."}, {"from_proposal": 2, "steps": [2, 8], "what": "Triage of the 53 open historical actions and shadow-mode testing for new alerts.", "why": "Added to S2 and S9 respectively."}, {"from_proposal": 4, "steps": [16, 23], "what": "Hard date after which pages outside the platform are not valid obligations, plus a weekly noise leaderboard.", "why": "Added to S11 and S9 as enforcement levers for consolidation and noise reduction."}], "rejected": [{"from_proposal": 2, "steps": [2], "what": "A staffed interim incident-command process in the first seven days.", "why": "S1 mentions 'interim controls week 1' only as a timeline entry; no step assigns duty commanders or a declaration path before the pilot in weeks 6–11."}, {"from_proposal": 2, "steps": [4], "what": "SEV0 crisis tier with distinct crisis-management triggers.", "why": "S5 keeps four levels and puts ledger-integrity and security events at SEV1."}, {"from_proposal": 5, "steps": [7], "what": "Status page within 5 minutes of SEV1.", "why": "S14 sets 15 minutes for SEV1 and 30 for SEV2 with pre-approved templates."}]}, {"proposal": 4, "assessment": "improved", "what_changed": "Consolidated 30 loose steps into 24 denser ones (postmortems and actions merged into S17, runbooks and readiness folded into S14, all comms into S15) and imported the financial and training machinery it was missing. Remains the most readable plan, but it dropped two of its own best R0 steps.", "improvements": ["S2 adds a 7-day operating floor with an interim IC, retroactive pay and triage of all 53 open actions — R0 had no interim state.", "S16 adds the SLA credit workflow R0 lacked, tying credits to root-cause families and turning the $1.3M into a standing business case.", "S18 adds a tiered training academy with certification valid 12 months and a register as audit artefact.", "S21 now reports median and p90 by tier/journey/region, keeps error budgets gating features, and reconciles dashboards monthly against customer cases.", "S7 and S13 preserve ledger dual control and reconciliation even under the IC's pre-authorised failover powers.", "Fixed R0's arithmetic: S22 now claims 'roughly three months of operating evidence' after a week-24 finish rather than R0's impossible five months.", "Tightest metric set in the round (16 items), each tied to a date."], "regressions": ["Dropped R0 S14 'Publish Incident Management Policy v1' — the ten-page, signed, all-hands artefact is now only a bullet in S24, losing the mid-outage usable document.", "Dropped R0 S29 'inspect and adapt after 90 days'; there is no data-driven revision point before the annual review in S21, and severity calibration has no correction loop.", "R0 S30's mock auditor interviews and observation-window freeze are compressed to a single line in S24.", "No risk register or contingency step anywhere, unlike P1 S31 and P3 S26.", "S20 pilot does not depend on S15 (communications) although it activates the status page."], "taken": [{"from_proposal": 1, "steps": [2], "what": "Forensic re-coding of the 31 incidents and profiling of the 3,400 alerts before design is locked.", "why": "Became S3 and is cited as the source for playbook and noise priorities."}, {"from_proposal": 1, "steps": [17, 21, 24], "what": "SLA credit workflow, tiered commander academy, and the metrics/review cadence.", "why": "Became S16, S18 and S21, filling R0's thinnest areas."}, {"from_proposal": 2, "steps": [2], "what": "Seven-day interim controls with a named IC in 10 minutes and a daily 15-minute ops review.", "why": "Became S2, explicitly designated as the first audit evidence."}, {"from_proposal": 2, "steps": [1, 4, 5], "what": "Financial dual control during incidents, mitigation-versus-resolution semantics, and 8–12 coherent responder domains.", "why": "S7 and S13 keep dual control, S14 defines resolution as reconciled, S8 names 8–12 domains rather than loose 'critical-path teams'."}, {"from_proposal": 2, "steps": [7, 8], "what": "Ingest all six tools first and retire a legacy path only after two weeks of verified operation; shadow-mode new alerts.", "why": "Added to S12 and S10 so consolidation cannot open a detection hole."}], "rejected": [{"from_proposal": 2, "steps": [4], "what": "SEV0 crisis tier.", "why": "S6 keeps four levels with SEV1 covering money-movement, ledger-integrity and both-region failure."}, {"from_proposal": 1, "steps": [14, 15, 16], "what": "Separate steps for internal comms, customer comms and the regulatory playbook.", "why": "S15 merges all three into one step with a single timing table, consistent with its compression strategy."}, {"from_proposal": 3, "steps": [4], "what": "Managed on-call service for overnight first-response triage.", "why": "S8 staffs Layer A in-house and treats follow-the-sun as a 12-month option, not a year-1 dependency."}]}, {"proposal": 5, "assessment": "improved", "what_changed": "Expanded from 14 thin steps to 26, closing nearly every R0 gap: interim duty officer, forensic baseline, listening tour, escalation ladder, regulatory playbook, credits, runbooks, exercises, pilot, waves and SOC 2 dry-run. The gains come almost entirely from copying P1's R0 structure and P2's SEV0, so it is now complete but the least differentiated plan.", "improvements": ["S1 stands up an interim 24x7 duty officer and single declaration path within 48 hours, compensated retroactively.", "S5 adopts a SEV0 crisis tier for corrupted or duplicated money movement, with payment-pause consideration.", "S13 requires human acknowledgement (delivery does not count) and routes unowned alerts to command as a logged control defect.", "S16 regulatory playbook with NYDFS 72-hour clock and a 24x7 contact matrix, entirely absent in R0.", "S17 SLA credit workflow and S20 readiness bar give the plan the financial and preparedness legs R0 lacked.", "Dropped R0's incoherent 'follow-the-sun between the two AWS regions' idea in favour of a paid NY command rotation (S7)."], "regressions": ["S22 pilot depends on S12, S13, S15, S19, S21 but **not on S9 (compensation)**, contradicting its own metric that paid on-call must be live before any mandatory rotation.", "No risk register or contingency step, and no standalone change-management/culture step — the pager-resistance answer is a bullet in S3 and a survey in S26.", "No anti-gaming or metric-reconciliation control, unlike P1 S26, P2 S18 and P4 S21.", "S8 caps rotation at one primary week in four, harsher than the one-in-six of P1, P2 and P4, weakening the fairness argument.", "S5's clock-driven auto-escalation (SEV2 unresolved >1h becomes SEV1) will inflate SEV1 volume when mitigation is already underway; no exception is defined.", "Success metrics are still anchored to 'within 6 months of full rollout' rather than absolute days, so the dates float with any schedule slip."], "taken": [{"from_proposal": 2, "steps": [4], "what": "SEV0 crisis level for integrity, duplication and dual-region loss, with consideration of pausing payments.", "why": "Adopted as the top of a five-level scale in S5, with executive, Legal and Security engagement attached."}, {"from_proposal": 2, "steps": [2], "what": "Interim command and one declaration path stood up immediately and paid retroactively.", "why": "Compressed into S1 as a 48-hour action rather than a separate step."}, {"from_proposal": 2, "steps": [9], "what": "Time-bound acknowledgement ladder, human acknowledgement required, unowned alerts to the command rotation as a control defect.", "why": "Became S13 almost verbatim, replacing R0's vaguer routing rules."}, {"from_proposal": 1, "steps": [9, 16, 17, 20, 28], "what": "Stipend figures, regulatory playbook, credit workflow, runbooks and readiness bar — with the exception of the risk register.", "why": "Became S9, S16, S17 and S20; the risk register was the one element left behind."}, {"from_proposal": 1, "steps": [23, 25], "what": "Six-week critical-path pilot with numeric exit criteria and four gated waves with coaches.", "why": "Became S22 and S23, replacing R0's vague 'pilot with 2-3 volunteer teams for 2 weeks'."}], "rejected": [{"from_proposal": 2, "steps": [4], "what": "Mitigation-versus-resolution semantics requiring reconciliation before closure.", "why": "S5 defines severities and auto-escalation clocks but never defines resolution; only S14/S15 mention 'resolution notice within 30 minutes of mitigation'."}, {"from_proposal": 4, "steps": [14, 23], "what": "A published Policy v1 and a quota-driven noise burn-down campaign with a leaderboard.", "why": "Policies appear only as SOC 2 artefacts in S25, and noise reduction stays a standing rule in S10 with no campaign or quota."}, {"from_proposal": 2, "steps": [13], "what": "Reconciling dashboards against customer cases to detect metric gaming or hidden incidents.", "why": "S24 lists metrics and cadence only, with no integrity check on the numbers."}]}]}
[ROUND 2]
{"round_summary": "All five plans converged onto a near-common template: four-level severity, a central command corps plus 10–12 domain rotations plus business-hours Layer C, paid on-call gated before any mandatory night rotation, a two-page-per-week budget, one paging platform, 3/5/10-day blameless postmortems, wave rollout with gates and a month-7 mock audit. Real differentiation narrowed to severity anchoring (P4's integrity flag, P1's payment anchors), plan weight and audit sequencing, while P3 essentially re-issued P1's round-1 plan.", "converging": true, "differences": ["**Severity anchoring.** P4 step 6 attaches a financial-integrity/security flag to any level (SEV0's forcing function without a fifth level); P1 step 6 adds payments anchors (value at risk per hour, settlement-deadline proximity, named accounts); P2 step 6 demotes percentages to guardrails; P3 step 7 and P5 step 6 stay purely qualitative. P1/P4 also tighten the SEV3 postmortem trigger (customer-detected, >2h, repeat), while P3/P5 keep \"postmortem on request\".", "**Audit sequencing.** P4 confirms the observation window in week 1 (step 1), P2 in step 3, P3 in step 5; P1 leaves it in step 30, which depends on steps 24/26/28, and cuts independent testing from two rounds to one at month 5. P2/P3/P5 keep tests in months 4 and 6.", "**Coverage economics.** P1 alone funds an overnight Duty Triage Desk owning the first ten minutes (step 8) but never sizes or budgets it; P4 alone shows the arithmetic (260 engineers sustain 10–12 rotations, not 28, step 8); P2 alone refuses to merge small teams for schedule convenience, scopes the six-responder rule to direct 24x7 rotations (steps 8, 22) and keeps a surge roster for concurrent incidents.", "**Weight and pace.** P1/P3 run 32 steps with standalone noise campaign, Policy v1 and risk register; P4 compresses to 26 with an 11-bullet comms mega-step (15) and the risk register demoted to last (26); P5 merges regulatory+credits (16) and postmortems+actions (17) and has no policy-publication step; P2 has 26 steps, no risk register at all, and the fastest schedule (Tier 0/1 by week 10, all teams by week 18 versus week 24 elsewhere)."], "influences": ["P2 and P5 abandoned their SEV0 tiers for the SEV1–SEV4 scale of P1/P3/P4 (P2 step 6, P5 step 6); P4 instead converted SEV0 into a financial-integrity/security flag (step 6) and P1 into payments anchors (step 6) — the round's one genuine design argument, resolved three different ways.", "P3 imported P1's round-1 plan almost wholesale (21 matching steps: Layer A/B/C 9, noise burn-down 26, Policy v1 24, risk register 31, 90-day inspect 32) and grafted on P2's day-one compliance step (P3 step 5 from P2 step 4) and P2's fake-redundancy check (P3 step 12 from P2 step 12).", "P5 took P1's Layer A/B/C coverage (step 8), readiness bar and ledger blast-radius workstream (step 10), risk register (step 25) and culture step (23), and tightened its own 1-in-4 rotation cap to P1's 1-in-6 (step 9).", "P2 took the standalone listening tour/fairness contract from P1 step 4, P3 step 3 and P4 step 4 (now P2 step 4) and concrete pay bands from P1/P5 step 9; in the other direction P4 took P2's week-1 observation-window confirmation (P4 step 1) and P1 added P2's 99.95% availability metric and an Internal Audit seat (step 1).", "Nobody took P5's named vendor shortlist (PagerDuty/incident.io/FireHydrant, round-1 step 12) — P5 dropped it itself; nobody adopted P2's numeric declaration thresholds (10% / 1–10% failure rates), which P2 also demoted; and only P2 (step 8) still plans for simultaneous incidents with a surge roster."], "proposals": [{"proposal": 1, "assessment": "improved", "what_changed": "Structure is unchanged at 32 steps; the gains are inside steps. Step 8 adds a paid overnight Duty Triage Desk, step 6 adds payments-specific severity anchors, step 11 adds an incident-replay test with \"replay coverage\" as a leading metric. Independent control testing shrank from two rounds to one.", "improvements": ["Duty Triage Desk (step 8) owns the first ten minutes of any page — the round's only concrete mechanism for cutting night MTTD without waking domain owners, and a direct answer to the pager objection.", "Payments anchors in step 6 (value at risk per hour, settlement-deadline proximity, number of affected named accounts) make severity quantitative without importing arbitrary failure percentages.", "Replay test in step 11 plus the metric \"a named detector would have caught ≥90% of the 31 historical incidents, with the expected detection minute\" — makes the detection uplift falsifiable rather than aspirational.", "New outcome metrics: availability ≥99.95% per customer journey by month 6; executive-signed ledger RPO/RTO with rehearsed read-only mode by month 6 and board-visible blast-radius milestones (step 21).", "Rollout deadline corrected from five to four months before the audit report date (step 28) — now consistent with a week-24 finish and a month-8 audit.", "Internal Audit added to the decision group (step 1); charter published company-wide on day 2."], "regressions": ["Independent testing cut from months 4 and 6 to a single month-5 test (step 30), where P2/P3/P5 keep two.", "SOC 2 mapping still sits in step 30, which depends on steps 24, 26 and 28 — the auditor conversation is formally gated behind policy publication and full rollout despite the text saying \"confirm the observation window early\".", "The Duty Triage Desk is unsized and unfunded: no headcount, source pool or line in step 1's budget, although it is a new paid overnight rotation.", "Still the heaviest plan — 32 steps and 24 metric lines — with steps 10 and 27 largely overlapping on alert quality."], "taken": [{"from_proposal": 2, "steps": [5], "what": "Percentage thresholds are declaration guardrails that must never justify under-classifying integrity, settlement or security risk.", "why": "P1 step 6 adopts the caveat but replaces P2's 10% / 1–10% numbers with payments anchors (value at risk, settlement proximity, named accounts)."}, {"from_proposal": 2, "steps": [1, 4], "what": "Internal Audit on the governing body and independent control testing before fieldwork.", "why": "Added Internal Audit to the decision group in step 1 and named a month-5 internal control test in the metrics."}, {"from_proposal": 2, "steps": [], "what": "Monthly availability ≥99.95% as an outcome metric (also in P3's round-1 metrics).", "why": "Added to P1's metric set, tying the process back to the contract the SLA credits come from."}], "rejected": [{"from_proposal": 2, "steps": [4], "what": "A standalone day-one step that designs compliance and evidence controls before the process is finalised.", "why": "P1 keeps SOC 2 work in step 30 behind dependencies on steps 24/26/28 instead of pulling it to week 1."}, {"from_proposal": 5, "steps": [5], "what": "A five-level scale with a separate SEV0 money/security crisis tier.", "why": "P1 keeps SEV1–SEV4 and handles the same triggers through ledger auto-escalation and the new payments anchors in step 6."}]}, {"proposal": 2, "assessment": "improved", "what_changed": "Renamed the scale to SEV1–SEV4, ending the SEV0/SEV1 ambiguity, and split seven previously buried topics (fairness contract, alert contract, customer comms, actions, policy, metrics, sustainability) into their own steps. Compensation moved from \"fixed weekly stipends\" to actual dollar bands and an annual budget.", "improvements": ["Step 6 adopts SEV1–SEV4 and keeps the guardrail caveat, so the scale is now comparable with every other plan and no longer has two crisis tiers.", "New step 4 promotes the listening tour and written fairness contract to an early standalone workstream with baseline, day-60, day-120 and quarterly surveys.", "Step 9 gives approvable numbers: $800–1,200 primary, $300–500 secondary, $900–1,300 duty commander, ~$0.9M–1.2M/yr, plus FLSA/NY treatment and a 15% sprint reduction.", "Standalone alert-quality contract (11), customer/status-page comms (15), corrective actions (18), policy-and-certification (20) and metrics (24) make each control separately testable.", "Best audit sequencing in the round: control mapping, retention, legal hold and observation-window confirmation in step 3; monthly sampling and tests in months 4 and 6 (step 25).", "Fastest credible schedule — Tier 0/1 coverage by week 10, all 28 teams by week 18 — giving the longest operating-evidence window before fieldwork.", "Unique operational realism: surge roster for simultaneous incidents and a 24x7 Executive Duty Officer (step 8); six-responder rule scoped to direct 24x7 rotations only (step 22)."], "regressions": ["Still the only plan without a risk register or pre-committed contingencies for compensation delay, commander shortfall, tool-migration slip or a SEV1 mid-rollout.", "No quota-driven noise burn-down campaign; step 11 relies on review cadence plus \"burn down the top 50 first\", which is weaker than P1/P3's leaderboard and per-sprint quota.", "No decision tree with worked examples from the 31 real incidents, so severity calibration rests on prose alone.", "Dropped its round-1 numeric declaration thresholds without substituting P1-style payment anchors, leaving SEV1/SEV2 boundaries more subjective than before."], "taken": [{"from_proposal": 1, "steps": [6], "what": "Four-level SEV1–SEV4 scale with SEV4 as a never-paging ticket class.", "why": "P2 step 6 renamed its SEV0–SEV3 scheme, removing the overlap between SEV0 and SEV1 response."}, {"from_proposal": 1, "steps": [4], "what": "Listening tour, resistance map and a written on-call deal as a first-class early step.", "why": "Became P2 step 4, with objection taxonomy, co-design recruits from the 12 on-call teams, and a repeated survey cadence."}, {"from_proposal": 1, "steps": [9], "what": "Published dollar figures for on-call stipends and a named annual budget.", "why": "P2 step 9 now carries bands and a $0.9–1.2M/yr envelope so HR, Finance and Payroll can approve within 14 days."}, {"from_proposal": 1, "steps": [16], "what": "Separate customer-communication step with timed status-page posts and tiered account-manager outreach.", "why": "Split out as P2 step 15, with pre-approved templates and strategic-account outreach at 30/60 minutes."}, {"from_proposal": 1, "steps": [24], "what": "Versioned policy suite plus one-page severity/role/escalation cards inside the incident tool.", "why": "Became P2 step 20, combining policy publication with role certification."}], "rejected": [{"from_proposal": 1, "steps": [31], "what": "A standalone risk register with pre-committed contingencies reviewed monthly with the sponsor.", "why": "P2 still has no such step; it handles failure through readiness gates, time-limited executive exceptions and honest deviation logging (steps 22, 25)."}, {"from_proposal": 1, "steps": [28], "what": "Quota-driven noise burn-down campaign with a weekly leaderboard and an automatic 8-week downgrade rule.", "why": "P2 step 11 instead sets an alert contract, two-business-day review of bad rules and a top-50 priority list."}, {"from_proposal": 1, "steps": [8], "what": "Consolidating the 28 teams into 10–12 domains as the default coverage design.", "why": "P2 step 8 warns against merging small teams for schedule convenience and step 22 applies the six-responder rule only to direct 24x7 rotations."}]}, {"proposal": 3, "assessment": "improved", "what_changed": "Effectively re-issues P1's round-1 plan: 21 of its 32 steps match P1 and only 5 match its own previous version. It gains the machinery it lacked (Policy v1, noise campaign, risk register, 90-day inspect) plus a day-one compliance step from P2, but contributes nothing new and leaves two internal inconsistencies.", "improvements": ["New step 5 maps TSC criteria, evidence, retention, legal hold and starts monthly sampling in week 1 — better audit sequencing than P1's round-2 step 30.", "Adds machinery its round-1 lacked: Policy v1 with exception register (24), quota-driven noise burn-down (26), risk register (31), 90-day inspect-and-adapt (32), anti-gaming reconciliation (27), separate live-execution doctrine (15).", "Keeps P4's \"missing ownership or runbooks is a release blocker for Tier 0/1\" (step 6) and P2's check that nominal two-region services do not hide single-region dependencies (step 12).", "Cleaner dependency chain than round 1: baseline → catalog → severity → roles → coverage → compensation → tooling → escalation → pilot → waves."], "regressions": ["Dropped the ledger blast-radius reduction workstream as a defined item, yet step 31 still promises to escalate \"if the blast-radius workstream slips\" — a dangling reference.", "Folded the readiness bar into the tenth bullet of step 15 rather than keeping a runbook/readiness step, where it is easy to lose during rollout gates.", "Dropped its own round-1 metrics for monthly 99.95% availability and \"<10% of engineers unwilling to join their owned-service rotation\".", "Step 28 still claims completion \"at least five months before the audit report date\" against a weeks 12–22 rollout and a month-8 audit — arithmetically impossible.", "Keeps \"postmortem on request\" for SEV3, looser than P1/P4's customer-detected/2-hour/repeat trigger.", "Adds no idea the round did not already have; it is now strictly behind P1's round-2 version."], "taken": [{"from_proposal": 1, "steps": [8, 24, 28, 31, 32], "what": "P1's overall architecture: Layer A/B/C coverage, Policy v1 with exception register, noise burn-down campaign, risk register, 90-day inspect-and-adapt.", "why": "Adopted almost verbatim as P3 steps 9, 24, 26, 31 and 32, replacing P3's thinner round-1 equivalents."}, {"from_proposal": 2, "steps": [4], "what": "Design compliance and evidence controls on day one, with retention, legal hold and monthly sampling.", "why": "Became P3 step 5, fixing the round-1 gap where audit work started late."}, {"from_proposal": 2, "steps": [12], "what": "Validate that nominally two-region services do not depend on single-region databases, identities, queues or third parties.", "why": "Added as the closing bullet of P3 step 12 on detection uplift."}, {"from_proposal": 4, "steps": [5], "what": "Missing ownership or runbooks is a release blocker for Tier 0/1 services.", "why": "Kept in P3 step 6 as the enforcement teeth on the catalog."}], "rejected": [{"from_proposal": 2, "steps": [5], "what": "A SEV0 financial/security crisis tier and numeric declaration guardrails.", "why": "P3 step 7 stays with SEV1–SEV4 and offers no quantitative anchor at all, not even P1's round-2 payment anchors."}, {"from_proposal": 1, "steps": [21], "what": "A standalone runbooks/readiness/ledger blast-radius step with a parallel architecture workstream.", "why": "P3 compresses it into step 15 and drops the blast-radius workstream, while step 31 still assumes it exists."}]}, {"proposal": 4, "assessment": "improved", "what_changed": "Compressed to 26 steps while adding the round's neatest severity idea and the only staffing arithmetic. Gains: integrity/security flag (step 6), 260-engineer rotation math (step 8), week-1 audit clock (step 1), postmortems split from action tracking. Losses: the change-management step disappeared and the risk register is demoted to the end.", "improvements": ["Step 6's financial-integrity or security flag attachable to any severity delivers SEV0's forcing function — dual control, Legal, regulatory checkpoint — without inventing a fifth level.", "Step 8 justifies the coverage model numerically: 260 engineers can sustain 10–12 domain rotations plus one command corps, not 28 night rotas; no two-person 24x7 rotation, ever.", "Step 1 confirms the SOC 2 observation window with the auditor in week 1 so the seven-day operating floor already counts as evidence.", "Postmortem standard (17) and action tracking (18) separated, where round 1 fused them; reportability checkpoint tightened from two hours to one (step 15).", "Noise burn-down runs as a per-wave quota inside step 24 rather than a parallel campaign — one fewer concurrent workstream during rollout.", "SEV3 postmortem trigger tightened to customer-detected, over two hours, repeat, or credit-generating."], "regressions": ["The round-1 change-management step is gone: no office hours, no four-hour help-channel SLA, no awards, no non-retaliation clause, and no public closure of the two \"nobody in charge\" incidents — only the fairness contract (4) and a per-wave reminder (24) remain.", "The risk register is folded into step 26, which depends on steps 23–25, so contingencies only exist after rollout and mock audit, and drop from seven pre-committed responses to four.", "Step 15 is an 11-bullet mega-step spanning internal cadence, status page, account-manager outreach and NYDFS/PCI/FinCEN obligations — too much for one operating card.", "Step 24 still promises completion \"at least five months before the audit report date\" against a week-24 finish and a month-8 audit."], "taken": [{"from_proposal": 2, "steps": [4], "what": "Confirm the auditor's expected Type II observation period immediately.", "why": "Moved into P4 step 1 so the interim floor and its artefacts land inside the window."}, {"from_proposal": 2, "steps": [5], "what": "A distinct SEV0 tier for financial-integrity and security crises.", "why": "Converted into a flag attachable to any severity (step 6), keeping four levels while forcing dual control, Legal and the regulatory checkpoint."}, {"from_proposal": 1, "steps": [19, 20], "what": "Separate postmortem standard and action-ownership steps.", "why": "Split into P4 steps 17 and 18, giving actions their own SLAs, escalation ladder and reserved 15% capacity."}, {"from_proposal": 1, "steps": [17], "what": "One-hour reportability assessment for SEV1 and security SEV2, recorded even when not reportable.", "why": "P4 step 15 tightened its round-1 two-hour window to one hour."}, {"from_proposal": 1, "steps": [31], "what": "Pre-committed contingencies for volunteer shortfall, pay delay and a SEV1 mid-rollout.", "why": "Folded into P4 step 26, though late in the chain and reduced in scope."}], "rejected": [{"from_proposal": 1, "steps": [15, 16, 17], "what": "Three separate steps for internal, customer and regulatory communications.", "why": "P4 merges them into one \"communications clock\" (step 15) so all timings sit on a single page."}, {"from_proposal": 1, "steps": [28], "what": "A standalone quota-driven noise burn-down campaign parallel to rollout.", "why": "P4 runs the quota, leaderboard rule and 8-week downgrade inside the wave gates (step 24)."}, {"from_proposal": 2, "steps": [7], "what": "Do not merge small teams merely for schedule convenience.", "why": "P4 step 8 argues the arithmetic for 10–12 merged domains and offers headcount, service reassignment or a time-limited executive exception instead."}]}, {"proposal": 5, "assessment": "improved", "what_changed": "Dropped SEV0 for a four-level scale, adopted P1's Layer A/B/C model, metric set and risk register, and added the readiness bar, simulations and a culture step. Still 26 steps, but two of them are now overloaded and the policy-publication step is missing.", "improvements": ["Step 6 replaces SEV0–SEV4 with SEV1–SEV4; the old SEV0 triggered nearly the same response as SEV1 and carried an unrealistic 10-minute status-page target.", "Step 8 states an explicit Layer A/B/C model with named domains, six-responder minimum and platform on-call as safety net rather than permanent owner.", "Adds what it lacked: readiness bar plus playbooks and the ledger blast-radius workstream (10), exercise programme (19), culture and fairness (23), risk register with pre-committed contingencies (25).", "Rotation cap tightened from one primary week in four to one in six (step 9), with the payroll-before-pager gate stated as a hard condition.", "Most honest schedule arithmetic in the round: week-24 completion leaves \"~3 months of evidence before audit fieldwork\" (step 22)."], "regressions": ["Lost its distinctive time-based auto-escalation \"unresolved SEV2 >1h becomes SEV1\" — nobody else has it and nothing replaced it.", "No policy-publication step: Policy v1, laminated cards and an exception register survive only as evidence bullets in step 24, so the artefact people open mid-outage is never produced.", "Dropped its round-1 vendor candidates (PagerDuty + incident.io/FireHydrant), the only concrete tooling shortlist anyone offered.", "Steps 16 (regulatory + credits) and 17 (postmortems + actions + review board) are overloaded, and step 20's nine dependencies make the pilot a single choke point.", "Keeps \"postmortem on request\" for SEV3 and has no quota-driven noise burn-down campaign, only the standing alert contract in step 11."], "taken": [{"from_proposal": 1, "steps": [8], "what": "Three-layer coverage: ~30-person command corps, 10–12 domain rotations of six, business-hours Layer C with a manager escalation list.", "why": "Adopted as P5 step 8, replacing its looser round-1 tiered-rotation design."}, {"from_proposal": 1, "steps": [21], "what": "Service readiness bar, major-incident playbooks and a parallel ledger blast-radius reduction workstream.", "why": "Became P5 step 10, with \"no readiness sign-off, no night paging\" enforcement."}, {"from_proposal": 1, "steps": [31], "what": "Risk register with pre-committed responses to volunteer shortfall, pay delay, tool slip, noise-induced blindness and a SEV1 mid-rollout.", "why": "Adopted near-verbatim as P5 step 25, reviewed monthly with the sponsor."}, {"from_proposal": 1, "steps": [29], "what": "Change-management step with the repeated deal, public closure of the two historical incidents and a pulse-survey expansion gate.", "why": "Became P5 step 23, running in parallel from day one."}, {"from_proposal": 1, "steps": [9], "what": "One primary week in six and compensation live in payroll before any mandatory rotation pages a human.", "why": "P5 step 9 tightened its round-1 one-in-four cap and made the payroll gate explicit."}], "rejected": [{"from_proposal": 2, "steps": [5], "what": "A SEV0 tier for financial or security crises.", "why": "P5 abandoned its own five-level scheme too and folded those triggers into SEV1 (step 6)."}, {"from_proposal": 1, "steps": [24], "what": "Publishing Incident Management Policy v1 with one-page cards and an exception register as its own step.", "why": "P5 keeps only versioned policy documents as SOC 2 artefacts in step 24, with no publication or launch moment."}, {"from_proposal": 1, "steps": [28], "what": "Quota-driven noise burn-down campaign with leaderboard and automatic downgrade rule.", "why": "P5 relies on the alert-quality standard and page budget in step 11 alone."}]}]}
[ROUND 3]
{"round_summary": "Round 3 converged hard: all five plans now run the same skeleton (charter → 7-day floor → forensic baseline → fairness contract → catalog/tiers → severity → roles → three-layer 24x7 → paid on-call → alert budget → detection → single paging platform → escalation → comms → regulatory/credits → postmortems/actions → training → exercises → policy → pilot → gated waves → metrics → SOC 2 → 90-day inspect), with near-identical numbers ($1,000 primary week, 2 out-of-hours pages/responder/week, 15% reserved capacity, 30 ICs / 18 comms leads). Differentiation now comes from a few genuine additions — P1's change intelligence and Support-as-detection tier, P4's Duty Triage Desk carve-out and vendor-incident class, P2's defined noise/actionable metrics and severity modifiers — while P3 and P5 mostly merge material authored by P1 and P4.", "converging": true, "differences": ["**Unique content**: only P1 has change intelligence/deployment safety (step 13), Support and account managers as an instrumented detection tier (step 20) and a concurrency doctrine with a Multi-Incident Coordinator (step 8); only P4 has a vendor-incident class for processor/bank/cloud failures (step 13) and a rule for who writes the 3 a.m. status page (step 8); only P2 defines \"actionable page\" and \"noise\" and uses five modifiers instead of flags (steps 12, 6). P3 and P5 contribute nothing the others do not have.", "**Overnight first line**: P1 (step 9), P4 (step 8) and P5 (step 8) fund a paid Duty Triage Desk; P4 adds the safety carve-out that it never holds suspected ledger, payment-halt or security pages. P2 (step 8) and P3 (step 9) refuse a triage desk and route unknown-owner pages to Duty Command plus platform, P2 additionally reclassifying lower-tier services that can cause severe overnight harm.", "**Schedule**: P2 (step 1) compresses design to weeks 1–4, pilot 5–10, rollout by week 18, control tests in months 4 and 6; P1, P3, P4 and P5 keep design weeks 1–6, pilot 7–12, rollout 13–24 and a month-5 internal test. P2 buys a longer observation window at the cost of a 4-week design for 180 services.", "**Failure handling and metric honesty**: P1 (step 34), P3 (step 28) and P5 (step 29) keep a standing monthly risk register; P4 compresses contingencies into bullets in step 26; P2 has none. On claims, P2 targets \"no unresolved material exception\" and P1 warns credits may rise before they fall (step 22), while P3 and P5 still promise \"SOC 2 passes with zero exceptions\"."], "influences": ["P1's replay test (R2 step 11) went everywhere: P2 step 13 plus a replay-coverage metric, P3 step 12, P4 steps 3 and 11 (as a \"replay catalog\" with a gap owner), P5 step 11. It is now the field's standard proof that detection actually improved.", "P4's severity flag (R2 step 6) was taken by P1 (step 7, financial-integrity/security flag forcing dual control and the reportability checkpoint) and generalised by P2 into five modifiers — FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, VENDOR (step 6). P3 and P5 declined it and kept four levels plus auto-escalation.", "P1's Duty Triage Desk (R2 step 8) was adopted by P4 (step 8) and P5 (step 8); P4 improved it with the carve-out that ledger, payment-halt and security events page commander and domains in parallel, plus a half-day triage-desk training module (step 18). P2 and P3 explicitly did not.", "P4's \"confirm the SOC 2 observation window in week 1\" (R2 step 1) was taken by P3 (step 1) and P5 (step 3), and expanded by P1 into a full week-one audit, legal and evidence scoping step (step 4) covering privilege, legal hold and the obligation matrix.", "P4's staffing arithmetic (R2 step 8 — 260 engineers cannot sustain 28 night rotas, never a two-person rotation) was taken verbatim by P3 (step 9) and restated by P1 (step 9, ~100 of 260 engineers carry a night obligation at one week in six). Nobody took P2's R2 caution against merging small teams for schedule convenience — all five consolidate into 10–12 domains."], "proposals": [{"proposal": 1, "assessment": "improved", "what_changed": "Four substantive new steps and three hardening additions, all addressing gaps the field had left open. It is now the most complete plan, at the cost of being the longest (35 steps).", "improvements": ["New step 4: week-one audit, legal and evidence scoping — observation window, TSC mapping, retention, legal hold and which postmortem content is privileged, with monthly evidence sampling from month 1.", "New step 13: change intelligence — deployments, config, flags and migrations streamed into the timeline, a \"what changed in the last 60 minutes\" panel at declaration, tested rollback executable by the on-call without the author, no deploys in the settlement window; backed by a new metric (change correlation identified within 10 minutes in 90% of cases).", "New step 20: Support and account management as a detection tier — declaration macro, five-tickets-in-ten-minutes clustering, auto-generated affected-customer list, and \"signal was in Support before monitoring\" published as a detection defect. This attacks the 40% customer-first number directly.", "Step 22 adds an honest incentive guardrail: better detection will raise credits before it lowers them, and no commercial pressure may influence severity.", "Step 8 adds a concurrency doctrine (secondary commander, surge roster, Multi-Incident Coordinator arbitrating ledger/DB/freeze) with a matching concurrency drill in step 26 and in the metrics.", "Step 9 publishes the staffing arithmetic (~100 of 260 engineers on a night obligation) so the rejection of 28 rotations is argued on numbers."], "regressions": ["35 steps with dense dependency chains (step 28 depends on seven prior steps) — the heaviest plan to actually run inside eight months.", "The on-call compensation metric changed to \"before any mandatory night rotation\", which is correct, but step 2's interim 24x7 roster still runs for weeks on retroactive pay only."], "taken": [{"from_proposal": 4, "steps": [6], "what": "Attach a financial-integrity or security flag to any severity instead of inventing a fifth level.", "why": "Used in step 7 to force dual control, Legal engagement and the reportability checkpoint on small-blast-radius integrity events."}, {"from_proposal": 4, "steps": [1], "what": "Confirm the SOC 2 Type II observation window with the auditor in week 1 so the interim process counts as evidence.", "why": "Expanded into a dedicated step 4 covering TSC mapping, evidence set, retention, legal hold and privilege."}, {"from_proposal": 2, "steps": [3, 10], "what": "Legal hold, confidentiality and privileged handling of incident records, plus MFA and access reviews on the incident platform.", "why": "Folded into step 4 (with a rule that privilege must not make ordinary postmortems secret) and step 14."}, {"from_proposal": 2, "steps": [8], "what": "A surge roster for simultaneous incidents.", "why": "Developed in step 8 into a concurrency doctrine with a Multi-Incident Coordinator and a concurrency drill in step 26."}, {"from_proposal": 2, "steps": [14, 15], "what": "Initial internal brief within 10 minutes for SEV1 and 15 for SEV2, and a monitoring notice after mitigation.", "why": "Added to the internal cadence in step 18 and the status-page sequence in step 19."}, {"from_proposal": 4, "steps": [8], "what": "Staffing arithmetic and the ban on two-person 24x7 rotations.", "why": "Published in step 9 as the numerical justification for 10–12 domain rotations instead of 28."}], "rejected": [{"from_proposal": 2, "steps": [8], "what": "Do not merge small teams merely for schedule convenience.", "why": "P1 step 9 still consolidates 28 teams into 10–12 response domains and defends it with staffing math, offering headcount or reassignment only to teams that cannot staff fairly."}, {"from_proposal": 2, "steps": [22], "what": "Finish company rollout by week 18 with control tests in months 4 and 6.", "why": "P1 keeps waves 13–24 with a month-5 internal test (steps 1, 31, 33), requiring completion four months before the audit report date instead."}]}, {"proposal": 2, "assessment": "improved", "what_changed": "It decomposed its round-2 mega-steps into discrete, individually gated steps and added the sharpest measurement definitions in the field. It remains the only plan without a risk register.", "improvements": ["Round-2 compressions were split: fairness contract (step 4), live lifecycle (step 10), escalation ladder (step 14), credits (step 17), postmortems (step 18), corrective actions (step 19), policy (step 23), scorecard (step 24), rollout (step 26), inspect (step 28).", "Step 12 defines an actionable page and defines noise (duplicate, non-urgent, unactionable, stale, test-generated, misrouted) and counts \"human notification episodes\" rather than raw alert events — the only plan whose noise target is unambiguously measurable.", "Step 6 replaces a single flag with five modifiers (FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, VENDOR) that invoke specialist controls without distorting customer-impact severity.", "Step 24 (scorecard) is now instrumented before the pilot (step 25 depends on 24), so pilot failures are visible immediately.", "Step 12 adds \"never disable an existing critical detector solely because its metadata or runbook is incomplete\"; step 19 prefers hazard-removing actions over \"add monitoring\" or retraining.", "Step 16 auto-enrols customers on the status page only where contracts and consent permit, and metrics target \"no unresolved material exception\" rather than the others' absolute \"zero exceptions\"."], "regressions": ["Still no standing risk register or pre-committed contingencies; the only failure handling is per-gate exceptions and step 28.", "Step 16 now carries internal-to-customer comms, the obligation matrix and reportability in 12 bullets — a new compression that undoes some of the decomposition gains.", "Dropped the round-2 aggregate on-call budget estimate ($0.9M–$1.2M/yr), leaving only per-week bands against the $1.3M credit comparison.", "Design compressed to weeks 1–4 for a 180-service catalog, 28-team listening tour and tool selection — the most optimistic schedule of the five."], "taken": [{"from_proposal": 1, "steps": [11], "what": "Replay the 31 historical incidents to name the detector that would now fire and at which minute.", "why": "Added as step 13's final control and as a day-90 metric (90% replay coverage), tying noise reduction to preserved Tier 0/1 replay coverage."}, {"from_proposal": 4, "steps": [6], "what": "A flag attached to any severity for integrity or security events.", "why": "Generalised in step 6 into five modifiers that trigger specialist controls without inflating customer-impact severity."}, {"from_proposal": 4, "steps": [4], "what": "A standalone, published on-call fairness contract before any mandatory night rotation.", "why": "Promoted from a bullet inside co-design to its own step 4, with confidential accommodations and the platform-does-not-inherit rule."}, {"from_proposal": 4, "steps": [21], "what": "Publish a signed Policy v1 plus an exception register as a discrete step.", "why": "Became step 23 with approval by CTO, HR, Legal, Security, Compliance and Internal Audit and versioning of every change."}, {"from_proposal": 3, "steps": [20], "what": "A blameless postmortem standard with fixed deadlines and a weekly Incident Review Board as its own step.", "why": "Split out as step 18, with corrective actions separated into step 19 so ownership and verification get their own gate."}], "rejected": [{"from_proposal": 1, "steps": [8], "what": "A paid overnight Duty Triage Desk that owns the first ten minutes of any page.", "why": "P2 step 8 instead routes unknown-owner incidents to Duty Command and Platform temporarily, logs each as a catalog control defect, and reclassifies lower-tier services capable of severe overnight harm rather than buffering them."}, {"from_proposal": 1, "steps": [16], "what": "Subscribe all 2,100 customers to the status page by default.", "why": "P2 step 16 offers subscriptions to all customers but auto-enrols only where contracts, consent and communication rules permit."}, {"from_proposal": 1, "steps": [30], "what": "Claim SOC 2 incident-response controls will pass with zero exceptions.", "why": "P2's metrics commit only to completing external testing without an unresolved material exception, consistent with its step 27 rule to record deviations honestly."}]}, {"proposal": 3, "assessment": "improved", "what_changed": "A competent merge: it absorbed P4's audit clock and staffing math, P1's replay test and P2's accommodations, and split the readiness bar out of the execution doctrine. It adds nothing the other plans do not already have, and it created two new mega-steps.", "improvements": ["Step 1 now confirms the SOC 2 observation window in week 1 so the day-7 floor counts as evidence.", "Step 9 publishes the staffing arithmetic (260 engineers cannot sustain 28 rotas) and bans two-person 24x7 rotations, making the coverage design arguable rather than asserted.", "Step 12 adds the replay test and the two-region fake-redundancy validation; step 24 reports replay coverage.", "Step 4 adds a non-punitive accommodation process, the platform-does-not-inherit rule, and the hard gate that no new mandatory night rotation starts before contract, compensation and training are live.", "Step 15 pulls the service-readiness bar, playbooks and ledger blast-radius reduction out of the execution doctrine into their own step with board-visible milestones."], "regressions": ["Step 14 now merges the escalation ladder and the live execution doctrine into 12 bullets, undercutting its own claim to be \"one short operating procedure\".", "Step 16 merges internal comms, status page, account-manager outreach, the regulatory obligation matrix and the contact matrix into one step; round 2 had these as three separate steps (16, 17, 18).", "Dropped the round-2 option to evaluate a first-line overnight triage desk, leaving Layer C nights covered only by manager escalation lists.", "Metrics still claim \"SOC 2 Type II incident-response controls pass with zero exceptions\"; replay coverage appears in step 24 but not in the success-metric list."], "taken": [{"from_proposal": 4, "steps": [1], "what": "Confirm the SOC 2 Type II observation period with the auditor in week 1.", "why": "Added to step 1 so the interim floor is treated as audit evidence from day 7."}, {"from_proposal": 4, "steps": [8], "what": "Staffing math justifying 10–12 domains and the prohibition on two-person 24x7 rotations.", "why": "Inserted into step 9 as the argument against 28 night rotations."}, {"from_proposal": 1, "steps": [11], "what": "The replay test against the 31 historical incidents.", "why": "Appended to step 12 as the validation that detection is genuinely better, with replay coverage in the step 24 quality metrics."}, {"from_proposal": 2, "steps": [4], "what": "Confidential, non-punitive accommodations and the rule that platform on-call triages but does not inherit another team's service.", "why": "Added to the fairness contract in step 4 to answer the caregiving/health objection explicitly."}, {"from_proposal": 4, "steps": [4], "what": "No new mandatory night rotation starts until the fairness contract, compensation and training are live.", "why": "Adopted verbatim as the closing bullet of step 4 and reflected in the metrics."}], "rejected": [{"from_proposal": 4, "steps": [6], "what": "Integrity/security flags attached to any severity level.", "why": "P3 step 7 keeps four levels plus auto-escalation rules; a one-customer ledger corruption is handled only by the ledger auto-escalation bullet."}, {"from_proposal": 1, "steps": [8], "what": "A paid overnight Duty Triage Desk owning the first ten minutes of a page.", "why": "P3 step 9 removed even its round-2 line about evaluating one, relying on domain rotations plus platform as the unknown-owner safety net."}, {"from_proposal": 2, "steps": [22], "what": "Complete company rollout by week 18 with control tests in months 4 and 6.", "why": "P3 keeps waves over weeks 13–24 (step 25) with internal tests in months 4 and 6 and completion five months before the audit report date."}]}, {"proposal": 4, "assessment": "improved", "what_changed": "It imported P1's Duty Triage Desk and replay test and then improved both, and added two things nobody else has: a vendor-incident class and a rule for who writes the status page overnight. It stays the most compact serious plan at 26 steps.", "improvements": ["Step 8 adopts the Duty Triage Desk but fixes its latency risk: the desk never sits on suspected ledger, payment-halt or security events, which page commander and likely domains in parallel; step 18 adds a half-day triage-desk training module.", "Step 8 answers a gap every other plan leaves: SEV1 pages a Communications Lead 24x7, while for customer-visible SEV2 the commander publishes the first status from a template and Comms is paged at 30 minutes or when a top-100 account is affected.", "Step 13 adds a vendor-incident class (processor, sponsor bank, card network, cloud control plane) measured by time-to-customer-notice, time-to-failover-decision and queue management rather than root cause at the vendor.", "Step 3 builds a replay catalog with the gap owner, and step 11 makes closing replay gaps a precondition for claiming detection improved; replay coverage enters the metrics.", "Step 6 widens the flag set to integrity, security, settlement and regulatory, with the explicit worked case that one-customer ledger corruption is still a flagged crisis.", "Step 14 adds \"existing detection stays on while gaps are repaired\", closing the readiness-bar loophole that could silence detectors."], "regressions": ["Contingencies remain four bullets inside step 26 rather than a standing risk register reviewed monthly with the sponsor, as in P1, P3 and P5.", "Step 23 still bundles the noise burn-down quota with the wave rollout, so noise reduction has no independent owner or schedule.", "Typo in step 11 (\"replay test against the S3 catalog\") and metrics still claim SOC 2 passes with zero exceptions."], "taken": [{"from_proposal": 1, "steps": [8], "what": "A paid overnight Duty Triage Desk owning the first ten minutes of ambiguous or unowned pages.", "why": "Adopted in step 8 with a safety carve-out for ledger, payment-halt and security events, and trained as its own role in step 18."}, {"from_proposal": 1, "steps": [11], "what": "Replay the 31 incidents to name the detector that would now fire and at which minute.", "why": "Split into a replay catalog built during the baseline (step 3) and the replay test as a detection gate (step 11), plus a replay-coverage metric."}, {"from_proposal": 1, "steps": [30], "what": "A month-5 internal control test before the month-7 mock audit.", "why": "Replaced its round-2 months 4 and 6 testing with a single month-5 test in step 25, aligned to the mock audit and freeze."}, {"from_proposal": 2, "steps": [19], "what": "Keep existing critical detection active while readiness gaps are repaired.", "why": "Added to the readiness bar in step 14 so the bar cannot be satisfied by turning alerts off."}, {"from_proposal": 2, "steps": [9], "what": "A pay band for Communications Lead and scribe duty.", "why": "Step 9 adds $400–$800 for Comms or triage-desk duty alongside the primary/secondary/commander bands."}], "rejected": [{"from_proposal": 2, "steps": [8], "what": "Do not merge small teams merely for schedule convenience; reclassify or add staffing instead.", "why": "P4 step 8 consolidates ownership into 10–12 domains and justifies it with staffing math, offering headcount or reassignment only to small Tier 0 owners (step 5)."}, {"from_proposal": 1, "steps": [31], "what": "A standing programme risk register reviewed monthly with the sponsor.", "why": "P4 keeps pre-committed contingencies as bullets in step 26 with no review cadence."}, {"from_proposal": 2, "steps": [22], "what": "Complete all 28 teams by week 18.", "why": "P4 step 23 keeps four waves over weeks 13–24 with completion four months before the audit report date."}]}, {"proposal": 5, "assessment": "improved", "what_changed": "It closed two real holes from its round-2 version — there was no policy-publication step and no separated postmortem/action or internal/customer comms steps — but it did so by copying P1's round-2 plan almost step for step. It contributes no idea of its own and picks up none of P1's round-3 advances.", "improvements": ["New step 22 publishes Incident Management Policy v1 plus the exception register; round 2 had no policy step at all, which was a SOC 2 gap.", "Split internal communications (step 15) from customer communications and status page (step 16), so cadence, templates and outreach each have an owner.", "Split blameless postmortems (step 18) from action ownership and enforcement (step 19), making the 11-of-64 problem its own enforcement step.", "Added the Duty Triage Desk to the coverage model (step 8) and a standalone alert noise burn-down campaign with quotas and a weekly leaderboard (step 25).", "Step 3 now confirms the SOC 2 observation window and maps controls to the Trust Services Criteria during the baseline."], "regressions": ["The standalone readiness-bar and runbook step disappeared into step 14, which now carries execution doctrine, readiness bar, eight playbooks and the ledger blast-radius workstream in one step.", "SLA credits lost their own step and are now two bullets at the end of the regulatory playbook (step 17), weakening the money-to-reliability feedback loop the others make explicit.", "Metrics still say paid on-call must be live \"before any rotation pages a human\", which contradicts its own step 2 interim 24x7 roster paid retroactively; P1 and P4 fixed this wording to \"any new mandatory night rotation\".", "Replay coverage is used in step 11 but never appears in the success-metric list, and metrics still claim SOC 2 passes with zero exceptions."], "taken": [{"from_proposal": 1, "steps": [8], "what": "Duty Triage Desk as a paid overnight first line owning the first ten minutes of a page.", "why": "Added to the coverage model in step 8 as the structural answer to \"I will not carry a pager for other teams' code\", and to the metrics."}, {"from_proposal": 1, "steps": [11], "what": "The replay test over the 31 historical incidents.", "why": "Adopted verbatim in step 11 as the validation of the detection uplift."}, {"from_proposal": 1, "steps": [24], "what": "Publish Incident Management Policy v1 with companion standards and an exception register.", "why": "Added as step 22, filling the round-2 gap where no policy was ever formally published or signed."}, {"from_proposal": 1, "steps": [15, 16], "what": "Separate internal communications protocol from customer/status-page policy.", "why": "Split into steps 15 and 16 with distinct cadences, templates and account-manager outreach timings."}, {"from_proposal": 1, "steps": [27], "what": "A quota-driven alert noise burn-down campaign run in parallel with rollout.", "why": "Added as step 25 with the noise leaderboard, the 8-week paging downgrade rule and compensating-detection checks."}, {"from_proposal": 4, "steps": [1], "what": "Charter the programme and start the audit clock by confirming the observation window.", "why": "Adopted as the step 1 title and a baseline bullet in step 3."}], "rejected": [{"from_proposal": 4, "steps": [6], "what": "Integrity/security flags on any severity level.", "why": "P5 step 6 keeps four levels with auto-escalation only, so a narrow-blast-radius ledger corruption relies on the ledger auto-escalation rule."}, {"from_proposal": 2, "steps": [24], "what": "Instrument the metrics scorecard before the pilot so failures show immediately.", "why": "P5's metrics step 24 still depends on the pilot (step 23), so the scorecard lands after the pilot is already running."}, {"from_proposal": 1, "steps": [21], "what": "A standalone service-readiness bar, runbook and ledger blast-radius step.", "why": "P5 folds all of it into step 14 alongside live execution doctrine, producing an 11-bullet step that is hard to gate or sign off."}]}]}
[FINAL]
{"summary": "- **P1** — the most complete plan (35 steps): the only one with change-intelligence/deploy-safety (step 13), Support and account managers as an instrumented detection tier (step 20), a concurrency doctrine with a Multi-Incident Coordinator (step 8) and a written warning that credits may rise before they fall (step 22).\n- **P2** — the most rigorous on measurement and audit: defines \"actionable page\" and \"noise\" and counts *notification episodes* (step 12), five severity modifiers instead of a fifth level (step 6), two independent control tests in months 4 and 6, company-wide by week 18, and the only honest audit claim (\"no unresolved material exception\"); the only plan with no programme risk register.\n- **P3** — a competent merge of P1's architecture with P2's day-one compliance step; complete but contributes nothing of its own and carries two internal inconsistencies (month-5 vs months-4-and-6 testing; \"five months before the report date\" against a week-24 rollout).\n- **P4** — the most compact serious plan (26 steps) with the two sharpest operational rules in the field: a Duty Triage Desk that never holds ledger/payment-halt/security pages (step 8) and a named owner for the 3 a.m. status page; weakened by bundling (steps 14, 23) and by burying contingencies in the final step.\n- **P5** — complete and correct but almost entirely derivative of P1's earlier draft, with two overloaded steps (14, 17) and a metric that contradicts its own day-7 floor.\n- **All five share**: day-7 operating floor, forensic re-coding of the 31 incidents, written fairness contract, service catalog with Tier 0–3, SEV1–SEV4 with ledger auto-escalation, IC/comms/scribe/SME with dual control preserved, ~30-person command corps plus 10–12 domain rotations plus business-hours Layer C, ~$1,000 primary week with a payroll hard gate, two out-of-hours pages/responder/week budget, synthetic money-path probes plus the 31-incident replay test, one paging platform, 15/30-minute status-page clocks, 3/5/10-day blameless postmortems with verified action closure, pilot then gated waves, month-7 mock audit, 90-day inspect-and-adapt.", "assessments": [{"proposal": 1, "fitness": "strong", "strengths": ["**Covers the brief most fully and then adds the two highest-leverage things nobody else has**: step 13 streams every deploy/flag/migration into the timeline and gives the commander a \"what changed in the last 60 minutes\" panel — the fastest lever on a 3 h 10 min MTTM; step 20 turns Support/AMs into a measured detection tier with a \"signal was in Support before monitoring\" defect class, aimed straight at the 40% customer-first rate.", "Step 8 is the only concurrency doctrine in the field: secondary commander, surge roster and a Multi-Incident Coordinator arbitrating the ledger, DB platform and deploy freeze — realistic at 31 incidents a year with one command corps.", "Step 4 does audit, legal and evidence scoping in week one, including privilege, legal hold and segregation of privileged postmortem content — deeper than any rival's compliance step.", "Step 9 states the staffing arithmetic (~100 of 260 engineers on a night obligation at one week in six) and rejects 28 rotations on the numbers; step 10 gates every night rotation on payroll being live and pays the day-7 interim duty retroactively.", "Step 22 pre-empts the perverse incentive that better detection surfaces previously unbilled credits, and bans commercial pressure on severity — an honesty the others lack.", "Step 34 is a standalone risk register depending only on step 1, with pre-committed fallbacks (volunteer shortfall, comp delay, pruning hides a failure, SEV1 mid-rollout, ledger workstream slip).", "Metrics are staged and dated (MTTD <10 min by day 90, <5 by month 6; pages 3,400 → <1,500 by day 90 → <500 by month 6; replay coverage ≥90% by day 90) rather than a single end-state."], "weaknesses": ["**35 steps is the heaviest load in the field** for an eight-month window and a three-person programme office; step 27 (Policy v1) joins 11 dependencies.", "Claims \"SOC 2 Type II incident-response controls pass with zero exceptions\" in the metrics while step 33 rightly says a documented exception log beats a claim of perfection — an internal contradiction.", "Only one independent control test (month 5), against P2's two (months 4 and 6); rollout finishing week 24 leaves a thinner company-wide observation window than P2's week 18.", "Step 29 (noise burn-down) depends on the pilot (step 28, ~week 12) yet the target is under 1,500 pages by day 90 — tight sequencing.", "Auto-subscribes all 2,100 customers to the status page by default (step 19) without P2's consent/contract caveat.", "Layer C accepts a manager escalation list overnight without P2's rule to reclassify lower-tier services that can cause severe overnight harm."]}, {"proposal": 2, "fitness": "strong", "strengths": ["**Best measurement hygiene**: step 12 defines an actionable page and defines noise, and the metrics count \"human notification episodes\" rather than raw alert events, so the 3,400 → 500 claim cannot be met by re-labelling.", "Best audit sequencing: observation window confirmed in week 1 (step 3), independent control tests in **months 4 and 6**, mock audit month 7, company-wide by week 18 — the longest operating evidence period, and the only plan whose success criterion is \"no unresolved material exception\" instead of \"zero exceptions\".", "Step 6's five modifiers (FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, VENDOR) attach specialist controls to any level without distorting customer-impact severity.", "Step 8 refuses to hide overnight risk behind manager callback: any lower-tier service capable of severe overnight harm is **reclassified**, not left on a Layer C list — the sharpest treatment of that blind spot.", "Safety rails against self-inflicted blindness: shadow mode for 7 days, compensating detection and central approval before suppressing any Tier 0/1 rule, and \"never disable an existing critical detector solely because its metadata or runbook is incomplete\" (step 12, step 20).", "Legal care others miss: status subscriptions offered to all but auto-enrolled only where contracts and consent permit (step 16); privileged material kept in restricted linked records (step 11).", "Step 24 reconciles incident records monthly against support cases, credits and status history to detect under-reporting; scorecards direct help, never individual penalties."], "weaknesses": ["**No programme risk register or pre-committed contingencies** — the only final plan without one; nothing states what happens if the commander pool is short, compensation slips, or a SEV1 lands mid-rollout.", "Four-week design window (step 1) for 180 services, catalog, tiers, severity, comp policy and tooling selection is the most aggressive in the field; a slip pushes everything.", "No change-correlation tooling: deploys and flag flips are frozen during incidents but never streamed into the timeline as a first hypothesis.", "No support-ticket clustering automation (the \"five similar tickets in ten minutes\" rule the other four carry); Support intake relies on human judgement plus a five-minute conversion rule.", "No explicit concurrency plan beyond prohibiting simultaneous primary assignments — no surge roster or multi-incident arbitration.", "No overnight first-line triage and no rule for who writes the 3 a.m. status page when the comms pool is asleep."]}, {"proposal": 3, "fitness": "adequate", "strengths": ["Complete against every element of the brief: severity with payments anchors (step 7), roles with dual control (step 8), three-layer 24x7 (step 9), costed pay with a payroll hard gate (step 10), alert contract and page budget (step 11), replay-validated detection (step 12), one comms clock covering internal, status page, AMs and regulators (step 16), postmortems and enforced actions (18, 19), gated waves (25), risk register (28).", "Step 5 maps SOC 2 criteria, evidence set, retention, legal hold and policy versioning on day one, so the week-1 interim process already produces evidence.", "Step 12 validates that nominal two-region services do not depend on single-region databases, identities, queues or third parties — catches fake redundancy.", "Step 25 folds the noise-reduction quota into each wave and pauses expansion if the 60/120-day fairness pulse is red."], "weaknesses": ["**Two internal inconsistencies**: step 1 and the metrics say a month-5 internal control test while step 27 says months 4 and 6; step 25 promises all 28 teams \"at least five months before the audit report date\" against a rollout ending week 24.", "Step 14 bundles the escalation ladder, the five-minute command rule and the entire live-execution doctrine (including payments guards and handover) into one 12-bullet step; step 16 does the same for internal, customer and regulatory comms — weak gating for two of the most safety-critical areas.", "Contributes no idea the other plans do not already have; it is a merge of P1's architecture plus P2's compliance step.", "Keeps the \"SOC 2 ... pass with zero exceptions\" over-claim.", "No overnight triage desk, no change-correlation, no concurrency doctrine, no vendor-incident class."]}, {"proposal": 4, "fitness": "strong", "strengths": ["**The best-designed overnight first line**: step 8's Duty Triage Desk owns the first ten minutes of ambiguous or unowned pages but **never sits on a suspected ledger, payment-halt or security event** — those page commander and domains in parallel; step 18 adds a half-day triage-desk module with explicit safe-step limits.", "Only plan that answers who writes the 3 a.m. status page: SEV1 pages a Comms Lead 24x7; on customer-visible SEV2 the commander publishes from a template and Comms is paged at 30 minutes or if a top-100 account is hit (step 8).", "Step 13's vendor-incident class for processor, sponsor-bank, card-network and cloud-control-plane failures, judged on time-to-customer-notice and failover decision rather than root cause — directly relevant to a payments platform.", "Step 6's integrity/security/settlement/regulatory flag with the concrete test \"a one-customer ledger corruption is still a flagged crisis\" prevents under-classification by volume.", "Step 3 builds a replay catalog with a named gap owner during the baseline, which step 11 then closes — detection improvement is proved, not asserted.", "Step 2 replays one of the two nobody-in-charge incidents as a tabletop within 14 days, testing the floor process almost immediately.", "Most readable at 26 steps with equal coverage, and step 9 sizes the budget ($0.9M–$1.2M) then refines from actual rotation count."], "weaknesses": ["**Pre-committed contingencies live as bullets inside step 26, the last step** (depends on 21, 23, 25) — the failure responses are decided after rollout rather than reviewed monthly from the start.", "Step 14 bundles execution doctrine, the readiness bar, all major-incident playbooks and the ledger blast-radius workstream; step 23 bundles the noise burn-down with the wave rollout — harder to gate and to audit as distinct controls.", "No change-correlation panel or deployment-safety class despite freezing changes during incidents.", "No week-one legal scoping of privilege or legal hold for incident records; step 25 maps criteria but later.", "Single internal control test (month 5) and the \"zero exceptions\" over-claim.", "No concurrency handling beyond \"nobody is primary on two rotations in the same week\"."]}, {"proposal": 5, "fitness": "adequate", "strengths": ["Covers every element requested with sound content: severity with payments anchors (step 6), roles and dual control (step 7), three-layer coverage including the Duty Triage Desk (step 8), costed pay and fatigue rules (step 9), alert contract (10), replay-validated detection (11), one platform with survivability testing (12), escalation ladder (13), postmortems and actions (18, 19), waves with gates (26), risk register (29), 90-day inspect (30).", "Keeps the standalone noise burn-down campaign (step 25) with a weekly leaderboard and a compensating-detection check on every suppression.", "Step 29's risk register depends only on step 1, so contingencies exist from the start."], "weaknesses": ["**Wholly derivative** — it reproduces P1's earlier draft nearly step for step and picks up none of the final round's advances (change intelligence, Support detection tier, concurrency, vendor class, triage-desk carve-out, defined noise metrics).", "Step 14 is a ten-bullet mega-step combining execution doctrine, payments guards, closure rules, the readiness bar, all playbooks and the ledger blast-radius programme; step 17 merges the regulatory playbook with the SLA-credit workflow.", "Metric \"paid on-call live in payroll before **any rotation** pages a human\" contradicts step 2, where interim duty from day 7 is paid retroactively.", "Keeps the \"zero exceptions\" over-claim and a single month-5 control test.", "Status page auto-subscribes all 2,100 customers with no consent or contract caveat."]}], "ranking": [1, 2, 4, 3, 5], "ranking_reasons": "**P1 first**: it matches every rival on the shared skeleton and then adds three things that bear directly on this company's numbers — change correlation (step 13) against a 3 h 10 min MTTM, Support/AMs as a measured detection tier (step 20) against 40% customer-first detection, and a concurrency doctrine (step 8) that a 30-person command corps facing 31 incidents a year will need. It also has the earliest standalone risk register (step 34) and the deepest week-one legal/evidence scoping (step 4). The price is 35 steps, one control test instead of two, and a \"zero exceptions\" claim it contradicts itself.\n\n**P2 second**: the most disciplined plan — it alone defines what a page and noise are before promising to cut them, tests controls twice (months 4 and 6), finishes company-wide by week 18 for a longer observation window, and refuses to hide overnight risk behind a Layer C manager list. It loses to P1 on two counts: no programme risk register at all, and a four-week design window for 180 services that leaves no slack.\n\n**P4 third**: equal coverage in 26 steps and the two sharpest operational rules in the field (the triage-desk carve-out, the overnight status-page owner) plus the vendor-incident class. It falls below P2 because contingencies sit in the last step, and steps 14 and 23 bundle readiness, playbooks, ledger architecture, noise burn-down and waves into single units that are harder to gate.\n\n**P3 fourth**: complete and well written but purely a merge, with two dating inconsistencies (month-5 vs months-4-and-6 testing; \"five months before the report date\" against a week-24 finish) and two mega-steps covering escalation/execution and all communications.\n\n**P5 last**: the same content as P3 with less of it original, an extra overloaded step (14), and a headline metric that contradicts its own day-7 floor. Nothing in it is wrong; nothing in it is its own.", "versus_initial": [{"proposal": 1, "verdict": "better", "why": "Same author lineage, but the final plan closes the initial version's two structural holes — six unprotected weeks of design and an unstaffable coverage model — and adds payments-correctness and audit machinery it lacked.", "how": "Adds the 7-day operating floor with an interim commander roster paid retroactively (step 2); replaces \"14–18 teams at 24x7\" with 10–12 domain rotations, a ~30-person command corps, a Duty Triage Desk and published staffing arithmetic (step 9); adds the mitigate-vs-resolve lifecycle with reconciliation before closure and split-brain/duplication guards (steps 7, 16); adds the 31-incident replay test and replay coverage as a leading metric (step 12); adds change intelligence (13), Support-as-detection (20), concurrency (8), week-one legal/evidence scoping including privilege and legal hold (4), anti-gaming reconciliation (30) and the warning that credits may rise before they fall (22); metrics are now staged by day 90 / month 6 / month 12 rather than single end-states."}, {"proposal": 2, "verdict": "better", "why": "It keeps everything the initial P2 got right — day-one interim controls, ledger dual control, shadow-mode alerts, reconciliation before resolution, triage of the 53 open actions — and supplies the delivery machinery that version simply did not have.", "how": "Initial P2 had no pilot, no wave plan with gates, no exercise programme, no readiness bar, no policy publication and no risk register; the final plan has all six (steps 28, 31, 26, 17, 27, 34). It also publishes actual pay bands and a payroll hard gate rather than deferring amounts to HR, adds the detection replay test, the noise burn-down campaign with a leaderboard, the automated SLA-credit workflow and change-correlation tooling."}, {"proposal": 3, "verdict": "better", "why": "Initial P3 was internally inconsistent and structurally fragile; the final plan is coherent and gated.", "how": "Removes the contradiction between a 100-alerts-per-team budget and a <600/month target by using a two-pages-per-responder-per-week budget with auto-quarantine (step 11); drops the unrealistic \"56 certified ICs and all 260 engineers trained in 14 weeks\" for a certification path with shadow shifts and a 12-month renewal (step 25); replaces the single step-12 rollout that depended on all eleven prior steps with a pilot exit gate plus four risk-ordered waves signed off by directors (steps 28, 31); adds the 7-day floor, the risk register, the reportability checkpoint recorded even when \"not reportable\", and drops the unanalysed outsourced overnight triage vendor in favour of a costed internal triage desk."}, {"proposal": 4, "verdict": "better", "why": "Initial P4 had the best time-boxing and layer model but thin content on money, law, learning and concurrency, and no protection during its first ten weeks.", "how": "Adds a day-7 operating floor instead of leaving the company uncovered until the ~week-10 pilot; adds the regulatory obligation matrix with a one-hour reportability checkpoint recorded even when not reportable (step 21), the automated SLA-credit and true-cost model (22), the Incident Review Board with verified action closure and 15% reserved capacity (23, 24), the replay test (12), concurrency doctrine (8), change intelligence (13) and a monthly-reviewed risk register (34); keeps P4's strengths — week-one audit clock, page-load rules, pause-on-red-pulse, wording freeze for the observation window."}, {"proposal": 5, "verdict": "better", "why": "Initial P5 was the thinnest and contained two unworkable design choices.", "how": "Replaces the 5-minute SEV1 status-page rule with a realistic 15-minute clock plus templates and a named Comms Lead (step 19); replaces \"follow-the-sun between the two AWS regions\" — which creates no time-zone coverage for a single New York organisation — with a command corps, domain rotations, a paid triage desk and a written cost-and-date evaluation of a Lisbon/APAC cell as a year-two option (step 9); adds dated milestones throughout, a wave plan with gates, a page budget, alert shadow-mode, the risk register, the legal/regulatory playbook and the anti-gaming reconciliation that the initial plan had none of."}], "improved_over_initial": true, "improvement_summary": "- **Gained an operating floor**: every final plan protects the company from day 7 (interim commander roster, one declaration path, Support authorised to declare, 53 legacy actions triaged) instead of leaving six-plus weeks of design unguarded — P2's round-0 idea that spread to all five.\n- **Gained staffing realism**: \"28 teams on 24x7\" was replaced everywhere by a ~30-person command corps plus 10–12 domain rotations plus business-hours Layer C, with published arithmetic and a ban on two-person rotations — the direct, defensible answer to the pager pushback.\n- **Gained payments correctness**: mitigate-vs-resolve semantics, reconciliation and controlled backlog drain before closure, split-brain/replay/duplication guards, and preservation of dual control during ledger recovery.\n- **Gained proof rather than assertion on detection**: the 31-incident replay test with a named detector and expected detection minute became the field's standard evidence that MTTD work actually worked.\n- **Gained safety rails on noise cutting**: shadow mode for seven days, never silently disable, compensating detection before suppression, monthly missed-detection reviews — so the 3,400 → 500 target cannot be met by going blind.\n- **Gained honest instrumentation**: medians and p90, detection source on every incident, and monthly reconciliation of incident counts against support tickets, credits and status-page history to catch under-reporting.\n- **Gained enforceable money and law**: dollar pay bands with a payroll hard gate, FLSA/New York review, a reportability checkpoint recorded even when \"not reportable\", and an automated SLA-credit workflow tied to root-cause families.\n- **Lost differentiation**: P3 and P5 converged into near-copies of P1's earlier drafts, so five agents produced roughly three distinct plans.\n- **Lost some good ideas without a decision**: P2's numeric declaration thresholds (>10% failures = SEV1), its caution against merging small teams into domains for schedule convenience, and P3's explicit per-page pay table were all dropped rather than argued out.\n- **Carried an over-claim to the end**: four of five still promise \"SOC 2 passes with zero exceptions\"; only P2 corrected it to \"no unresolved material exception\", which is what an auditor can actually deliver.\n- **Grew heavier**: step counts rose from 13–30 to 26–35, with several plans creating mega-steps (execution doctrine plus readiness bar plus playbooks plus ledger architecture in one) that are harder to gate than the pieces they replaced."}
[PROCESS]
{"process_evaluation": "- **Convergence: genuine in round 0, imitation thereafter.** Five agents independently landing on the same skeleton (exec mandate → baseline → severity → separated roles → paid on-call → one pager → synthetics → timed comms → postmortems → pilot → waves → audit evidence) is real evidence the skeleton is right. What followed was not argument but transcription: P3 imported 21 of 32 steps from P1's round-1 plan; P5 copied P1's round-2 plan step for step; by round 3 the success-metric lists of P1, P3, P4 and P5 are near-verbatim (same \"$1,000 primary week\", \"2 out-of-hours pages per responder\", \"30 certified commanders and 18 comms leads\", \"10–12 domain rotations\", \"15% reserved capacity\"). Two exceptions were earned: P4 improved P1's Duty Triage Desk by carving out ledger/payment-halt/security pages so triage never sits on a crisis, and P2 generalised P4's integrity flag into five modifiers (FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, VENDOR).\n- **Almost no criticism; disagreements were resolved by fiat, not by evidence.** The one real design argument — SEV0 tier vs. four levels vs. severity flags vs. modifiers — ended with three different unreconciled answers and nobody testing them against the 31 historical incidents that every plan claims to have re-coded. P2's warning against merging small teams into domains purely for schedule convenience was ignored rather than rebutted; all five consolidated into 10–12 domains anyway. P1's Duty Triage Desk was copied by P4 and P5 and refused by P2 and P3 — but the refusal was never argued, and the desk is still unsized and unbudgeted in P1 and P5 (only P4 prices it at $400–800/week). Nobody attacked P1's 11-dependency join at step 27, the growth to 35 steps, or the arithmetic of a three-person programme office delivering all of this in 24 weeks.\n- **Premises were barely questioned, and the best challenges did not spread.** Three genuine ones: P1 step 22 warns credits may *rise* before they fall because better detection surfaces previously unbilled outages; P2 refuses \"zero exceptions\" and targets \"no unresolved material exception\"; P4 does the staffing arithmetic showing 260 engineers cannot sustain 28 night rotas. Untouched premises: whether 99.95% is achievable at all on a single shared PostgreSQL ledger (relegated to a parallel \"blast-radius workstream\" in every plan); whether the business case survives the 15% capacity reservation (~39 FTE of 260, far larger than the $1.3M in credits, and never costed by anyone); whether hiring a dedicated 10–15 person SRE/NOC cohort beats putting ~100 of 260 engineers on nights; and the flat contradiction between \"you are paged only for services your team owns\" and Layer B consolidation of 28 teams into 10–12 shared domains — only P2 (step 8, \"explicit support acceptance\") resolves it, and nobody flagged it in the others.\n- **Losses: options, precision and independent judgement.** Dropped without argument: P3's managed overnight triage vendor (a live option for a single-timezone company); P2's numeric declaration thresholds (>10% payment failures for 5 min = SEV1), which were the only thing making severity a genuine lookup rather than a judgement call; P5's vendor shortlist; P3's fully derived $420K compensation budget, replaced by an undefended $0.9–1.3M. Regressions propagated as readily as improvements: P1 cut independent control testing from months 4 and 6 to month 5 alone, and P4 and P5 copied the weaker version. Unique good ideas failed to spread because copying flowed one way (P1 → P3/P5): only P1 has change intelligence/deployment safety (step 13) and Support as an instrumented detection tier (step 20), and only P1 and P2 plan for concurrent incidents at all. Most of all, the field lost the ability to discriminate: the vote was between one plan wearing five costumes.", "process_issues": ["**Two of five agents became pure copyists.** P3 and P5 contributed no idea the others lacked from round 2 onward, wasting 40% of the parallelism and inflating apparent consensus.", "**Success metrics were copied verbatim across plans**, so they stopped being evidence of design quality and became boilerplate — and the boilerplate carried an overclaim (\"SOC 2 Type II incident-response controls pass with zero exceptions\" in P1, P3, P4, P5) that the company does not control. Only P2 phrased it defensibly.", "**Step inflation rewarded length over executability.** P1 29→32→35, P5 14→26→30, P3 13→27→29. Nothing forced removal, so P1 now carries alert quality (11), a separate noise burn-down (29) and a per-wave noise quota (31) covering the same work, and comms split across steps 18/19/20.", "**Big dependency joins went unchallenged.** P1 step 27 (Policy v1) has 11 predecessors, P5 step 22 has 11, P3 step 22 has 10 — a single-point schedule risk in the middle of the plan that no reviewer named.", "**No plan was ever costed end-to-end.** Every plan reserves 10–15% of 260 engineers plus $0.9–1.3M compensation plus $150–250k tooling plus a three-person office, and every plan anchors the case on $1.3M of credits. Nobody did the subtraction.", "**Feasibility of the timeline was asserted, never tested.** Design in 6 weeks for 180 services, 30 certified commanders, synthetics in both regions, six-tool migration, all 28 teams by week 24. P2's alternative (design weeks 1–4, rollout by week 18) was never debated on the merits, only noted as different.", "**The severity design argument was decided without data.** Every plan claims to re-code the 31 incidents and to publish 12 worked examples from them, yet nobody used those incidents to show which severity model would have classified them correctly.", "**Contradiction survived all rounds unexamined:** the fairness promise \"paged only for what your team owns\" versus Layer B domain rotations that necessarily merge several of the 28 teams. Only P2 wrote the reconciliation (formally accepted support agreements, training, access).", "**Regressions propagated as fast as improvements** — the cut from two independent control tests (months 4 and 6) to one (month 5) spread from P1 to P4 and P5 with no discussion.", "**The final vote had almost nothing to discriminate on**, since the five plans share their architecture, numbers, metrics and most step text; preference collapsed onto length and two or three unique steps."], "suggestions": ["**Assign adversarial roles from round 1.** One agent must produce a minimum-viable plan (≤12 steps, must still pass the audit); one must attack cost and feasibility; one must defend a structurally different coverage model (managed overnight NOC, or hiring a dedicated 12-person SRE first line instead of putting 100 engineers on nights). Forbid these agents from converging.", "**Require a written critique artefact each round**: every agent names three specific flaws by proposal and step number, states what it rejected from others and why, and those rejection reasons are carried into the final round as a visible dissent register. Dropping an option silently should be disallowed.", "**Ban shared success metrics.** Each plan's metrics must be derivable from its own design choices, capped at ~10, each with a falsification condition and an owner. Penalise any metric the organisation cannot control (\"audit passes with zero exceptions\").", "**Impose a step budget with a removal quota**: to add a step you must delete or merge one, and any step with more than four dependencies must be split or justified. Reward net simplification in scoring.", "**Require one costing table per plan**: rotation count × stipend, triage-desk headcount and pay, tooling, programme office, and the opportunity cost of reserved capacity in FTE — set against $1.3M credits plus the quantified true incident cost. No plan advances without it.", "**Force premise challenges.** Each agent must list two brief assumptions that may be wrong and say what it would do if they are (e.g., 99.95% is unachievable on a shared ledger without re-architecture; credits will rise before they fall; pager pay does not fix attrition).", "**Run a resource-feasibility check as a gate**: a week-by-week load plan for the programme office and a named critical path, plus a written statement of what gets cut first if the mid-rollout SEV1 contingency fires.", "**Add two evaluator personas before the vote** — a SOC 2 auditor and a sceptical senior engineer from one of the 16 teams with no on-call — and require every plan to answer their written objections in its final round.", "**Score the final round on discriminating dimensions separately** (staffing arithmetic, cost, severity calibration evidence, contingency handling, readability under pressure) rather than an overall preference, so near-identical plans are forced to state trade-offs.", "**Detect and penalise derivativeness automatically**: measure step-level overlap with other proposals and require any agent above a threshold to either justify each copied step or be replaced with a fresh brief."]}
[VOTE COMPARISON]
{"agrees": true, "comment": "The vote matches my first choice, and largely for the same reasons I gave: change intelligence against the 3h10 MTTM, Support/AMs as an instrumented detection tier against 40% customer-first detection, the concurrency doctrine, and the paid Duty Triage Desk with published arithmetic answering the pager objection. Voter 1 adds a merit I weighed less — the SLA-credit incentive guardrail — and voters 3 and 5 lean partly on step count (\"35 steps\", \"most comprehensive\"), which hints at a length bias, since none of them flags P1's single control test, its self-contradictory \"zero exceptions\" claim, or P2's superior discipline in defining a page and noise before promising to cut them. Voter 2's case for P2 is exactly the runner-up argument I made (feasible paid domain-based coverage, suppression safeguards, evidence-producing rollout), so the top two are ordered as I had them. No voter engaged with P3/P4/P5, so the vote says nothing about the derivative-merge problem I penalised in P3 and P5."}
[BRIEF]
{"brief": "- **Outcome: P1 wins, and the vote agrees with my ranking.** P1 is the most complete plan (35 steps) and uniquely carries change-intelligence/deploy-safety (step 13) against the 3 h 10 min MTTM, Support/account managers as an instrumented detection tier (step 20) against 40% customer-first detection, and a concurrency doctrine with a Multi-Incident Coordinator (step 8); P2 is the disciplined runner-up (defines \"actionable page\" and \"noise\" in step 12, tests controls twice in months 4 and 6, company-wide by week 18).\n- **The final plans are clearly better than any round-0 draft**, gaining a day-7 operating floor, staffing arithmetic (~30-person command corps + 10–12 domain rotations + business-hours Layer C instead of \"28 teams on 24x7\"), mitigate-vs-resolve and reconciliation semantics for the ledger, the 31-incident replay test as proof detection improved, shadow-mode/never-silently-disable rails on noise cutting, and dollar pay bands with a payroll hard gate.\n- **Convergence was genuine in round 0, imitation afterwards.** All five independently found the same skeleton, then P3 imported 21 of 32 steps from P1's round-1 plan and P5 copied P1's round-2 plan step for step; only two exchanges were earned — P4 carving ledger/payment-halt/security pages out of P1's Duty Triage Desk (step 8) and P2 generalising P4's integrity flag into five modifiers (step 6).\n- **Criticism was nearly absent; disagreements ended by fiat.** The severity argument (SEV0 tier vs four levels vs flags vs modifiers) closed with three unreconciled answers and nobody testing them against the 31 re-coded incidents; P2's caution against merging small teams into domains was ignored, not rebutted; P1's 11-predecessor join at step 27 and its growth to 35 steps went unchallenged.\n- **Problem 1: two of five agents became copyists.** P3 and P5 added nothing the others lacked from round 2 on, wasting 40% of the parallelism and inflating apparent consensus, so the final vote discriminated mostly on length and two or three unique steps.\n- **Problem 2: metrics became boilerplate carrying an overclaim.** \"$1,000 primary week\", \"2 out-of-hours pages/responder\", \"30 commanders / 18 comms leads\" appear near-verbatim in four plans, and P1, P3, P4 and P5 all promise \"SOC 2 passes with zero exceptions\" — an outcome the company does not control; only P2 says \"no unresolved material exception\".\n- **Problem 3: nothing was ever costed or feasibility-tested.** Every plan reserves 10–15% of 260 engineers (~39 FTE) plus $0.9–1.3M compensation plus tooling and a three-person office against $1.3M of credits without doing the subtraction, and regressions spread as easily as improvements (P1's cut from two control tests to one at month 5 was copied by P4 and P5).\n- **Fixes that would help most:** (1) assign non-converging adversarial roles from round 1 — a ≤12-step minimum-viable plan, a cost/feasibility attacker, and a defender of a structurally different coverage model such as a managed overnight NOC or a dedicated 12-person SRE first line; (2) require a written critique artefact each round naming three flaws by proposal and step number plus a visible dissent register, so options like P2's numeric thresholds (>10% failures for 5 min = SEV1) cannot be dropped silently; (3) mandate one costing table per plan and ban shared metrics, capping each plan at ~10 metrics derived from its own design with a falsification condition and an owner."}The final round, ranked blind
Why this order
P1 first: it matches every rival on the shared skeleton and then adds three things that bear directly on this company's numbers — change correlation (step 13) against a 3 h 10 min MTTM, Support/AMs as a measured detection tier (step 20) against 40% customer-first detection, and a concurrency doctrine (step 8) that a 30-person command corps facing 31 incidents a year will need. It also has the earliest standalone risk register (step 34) and the deepest week-one legal/evidence scoping (step 4). The price is 35 steps, one control test instead of two, and a "zero exceptions" claim it contradicts itself.
P2 second: the most disciplined plan — it alone defines what a page and noise are before promising to cut them, tests controls twice (months 4 and 6), finishes company-wide by week 18 for a longer observation window, and refuses to hide overnight risk behind a Layer C manager list. It loses to P1 on two counts: no programme risk register at all, and a four-week design window for 180 services that leaves no slack.
P4 third: equal coverage in 26 steps and the two sharpest operational rules in the field (the triage-desk carve-out, the overnight status-page owner) plus the vendor-incident class. It falls below P2 because contingencies sit in the last step, and steps 14 and 23 bundle readiness, playbooks, ledger architecture, noise burn-down and waves into single units that are harder to gate.
P3 fourth: complete and well written but purely a merge, with two dating inconsistencies (month-5 vs months-4-and-6 testing; "five months before the report date" against a week-24 finish) and two mega-steps covering escalation/execution and all communications.
P5 last: the same content as P3 with less of it original, an extra overloaded step (14), and a headline metric that contradicts its own day-7 floor. Nothing in it is wrong; nothing in it is its own.
The ranking
- Proposal 1 strong selected by the vote Strengths
- Covers the brief most fully and then adds the two highest-leverage things nobody else has: step 13 streams every deploy/flag/migration into the timeline and gives the commander a "what changed in the last 60 minutes" panel — the fastest lever on a 3 h 10 min MTTM; step 20 turns Support/AMs into a measured detection tier with a "signal was in Support before monitoring" defect class, aimed straight at the 40% customer-first rate.
- Step 8 is the only concurrency doctrine in the field: secondary commander, surge roster and a Multi-Incident Coordinator arbitrating the ledger, DB platform and deploy freeze — realistic at 31 incidents a year with one command corps.
- Step 4 does audit, legal and evidence scoping in week one, including privilege, legal hold and segregation of privileged postmortem content — deeper than any rival's compliance step.
- Step 9 states the staffing arithmetic (~100 of 260 engineers on a night obligation at one week in six) and rejects 28 rotations on the numbers; step 10 gates every night rotation on payroll being live and pays the day-7 interim duty retroactively.
- Step 22 pre-empts the perverse incentive that better detection surfaces previously unbilled credits, and bans commercial pressure on severity — an honesty the others lack.
- Step 34 is a standalone risk register depending only on step 1, with pre-committed fallbacks (volunteer shortfall, comp delay, pruning hides a failure, SEV1 mid-rollout, ledger workstream slip).
- Metrics are staged and dated (MTTD <10 min by day 90, <5 by month 6; pages 3,400 → <1,500 by day 90 → <500 by month 6; replay coverage ≥90% by day 90) rather than a single end-state.
Weaknesses- 35 steps is the heaviest load in the field for an eight-month window and a three-person programme office; step 27 (Policy v1) joins 11 dependencies.
- Claims "SOC 2 Type II incident-response controls pass with zero exceptions" in the metrics while step 33 rightly says a documented exception log beats a claim of perfection — an internal contradiction.
- Only one independent control test (month 5), against P2's two (months 4 and 6); rollout finishing week 24 leaves a thinner company-wide observation window than P2's week 18.
- Step 29 (noise burn-down) depends on the pilot (step 28, ~week 12) yet the target is under 1,500 pages by day 90 — tight sequencing.
- Auto-subscribes all 2,100 customers to the status page by default (step 19) without P2's consent/contract caveat.
- Layer C accepts a manager escalation list overnight without P2's rule to reclassify lower-tier services that can cause severe overnight harm.
- Proposal 2 strong Strengths
- Best measurement hygiene: step 12 defines an actionable page and defines noise, and the metrics count "human notification episodes" rather than raw alert events, so the 3,400 → 500 claim cannot be met by re-labelling.
- Best audit sequencing: observation window confirmed in week 1 (step 3), independent control tests in months 4 and 6, mock audit month 7, company-wide by week 18 — the longest operating evidence period, and the only plan whose success criterion is "no unresolved material exception" instead of "zero exceptions".
- Step 6's five modifiers (FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, VENDOR) attach specialist controls to any level without distorting customer-impact severity.
- Step 8 refuses to hide overnight risk behind manager callback: any lower-tier service capable of severe overnight harm is reclassified, not left on a Layer C list — the sharpest treatment of that blind spot.
- Safety rails against self-inflicted blindness: shadow mode for 7 days, compensating detection and central approval before suppressing any Tier 0/1 rule, and "never disable an existing critical detector solely because its metadata or runbook is incomplete" (step 12, step 20).
- Legal care others miss: status subscriptions offered to all but auto-enrolled only where contracts and consent permit (step 16); privileged material kept in restricted linked records (step 11).
- Step 24 reconciles incident records monthly against support cases, credits and status history to detect under-reporting; scorecards direct help, never individual penalties.
Weaknesses- No programme risk register or pre-committed contingencies — the only final plan without one; nothing states what happens if the commander pool is short, compensation slips, or a SEV1 lands mid-rollout.
- Four-week design window (step 1) for 180 services, catalog, tiers, severity, comp policy and tooling selection is the most aggressive in the field; a slip pushes everything.
- No change-correlation tooling: deploys and flag flips are frozen during incidents but never streamed into the timeline as a first hypothesis.
- No support-ticket clustering automation (the "five similar tickets in ten minutes" rule the other four carry); Support intake relies on human judgement plus a five-minute conversion rule.
- No explicit concurrency plan beyond prohibiting simultaneous primary assignments — no surge roster or multi-incident arbitration.
- No overnight first-line triage and no rule for who writes the 3 a.m. status page when the comms pool is asleep.
- Proposal 4 strong Strengths
- The best-designed overnight first line: step 8's Duty Triage Desk owns the first ten minutes of ambiguous or unowned pages but never sits on a suspected ledger, payment-halt or security event — those page commander and domains in parallel; step 18 adds a half-day triage-desk module with explicit safe-step limits.
- Only plan that answers who writes the 3 a.m. status page: SEV1 pages a Comms Lead 24x7; on customer-visible SEV2 the commander publishes from a template and Comms is paged at 30 minutes or if a top-100 account is hit (step 8).
- Step 13's vendor-incident class for processor, sponsor-bank, card-network and cloud-control-plane failures, judged on time-to-customer-notice and failover decision rather than root cause — directly relevant to a payments platform.
- Step 6's integrity/security/settlement/regulatory flag with the concrete test "a one-customer ledger corruption is still a flagged crisis" prevents under-classification by volume.
- Step 3 builds a replay catalog with a named gap owner during the baseline, which step 11 then closes — detection improvement is proved, not asserted.
- Step 2 replays one of the two nobody-in-charge incidents as a tabletop within 14 days, testing the floor process almost immediately.
- Most readable at 26 steps with equal coverage, and step 9 sizes the budget ($0.9M–$1.2M) then refines from actual rotation count.
Weaknesses- Pre-committed contingencies live as bullets inside step 26, the last step (depends on 21, 23, 25) — the failure responses are decided after rollout rather than reviewed monthly from the start.
- Step 14 bundles execution doctrine, the readiness bar, all major-incident playbooks and the ledger blast-radius workstream; step 23 bundles the noise burn-down with the wave rollout — harder to gate and to audit as distinct controls.
- No change-correlation panel or deployment-safety class despite freezing changes during incidents.
- No week-one legal scoping of privilege or legal hold for incident records; step 25 maps criteria but later.
- Single internal control test (month 5) and the "zero exceptions" over-claim.
- No concurrency handling beyond "nobody is primary on two rotations in the same week".
- Proposal 3 adequate Strengths
- Complete against every element of the brief: severity with payments anchors (step 7), roles with dual control (step 8), three-layer 24x7 (step 9), costed pay with a payroll hard gate (step 10), alert contract and page budget (step 11), replay-validated detection (step 12), one comms clock covering internal, status page, AMs and regulators (step 16), postmortems and enforced actions (18, 19), gated waves (25), risk register (28).
- Step 5 maps SOC 2 criteria, evidence set, retention, legal hold and policy versioning on day one, so the week-1 interim process already produces evidence.
- Step 12 validates that nominal two-region services do not depend on single-region databases, identities, queues or third parties — catches fake redundancy.
- Step 25 folds the noise-reduction quota into each wave and pauses expansion if the 60/120-day fairness pulse is red.
Weaknesses- Two internal inconsistencies: step 1 and the metrics say a month-5 internal control test while step 27 says months 4 and 6; step 25 promises all 28 teams "at least five months before the audit report date" against a rollout ending week 24.
- Step 14 bundles the escalation ladder, the five-minute command rule and the entire live-execution doctrine (including payments guards and handover) into one 12-bullet step; step 16 does the same for internal, customer and regulatory comms — weak gating for two of the most safety-critical areas.
- Contributes no idea the other plans do not already have; it is a merge of P1's architecture plus P2's compliance step.
- Keeps the "SOC 2 ... pass with zero exceptions" over-claim.
- No overnight triage desk, no change-correlation, no concurrency doctrine, no vendor-incident class.
- Proposal 5 adequate Strengths
- Covers every element requested with sound content: severity with payments anchors (step 6), roles and dual control (step 7), three-layer coverage including the Duty Triage Desk (step 8), costed pay and fatigue rules (step 9), alert contract (10), replay-validated detection (11), one platform with survivability testing (12), escalation ladder (13), postmortems and actions (18, 19), waves with gates (26), risk register (29), 90-day inspect (30).
- Keeps the standalone noise burn-down campaign (step 25) with a weekly leaderboard and a compensating-detection check on every suppression.
- Step 29's risk register depends only on step 1, so contingencies exist from the start.
Weaknesses- Wholly derivative — it reproduces P1's earlier draft nearly step for step and picks up none of the final round's advances (change intelligence, Support detection tier, concurrency, vendor class, triage-desk carve-out, defined noise metrics).
- Step 14 is a ten-bullet mega-step combining execution doctrine, payments guards, closure rules, the readiness bar, all playbooks and the ledger blast-radius programme; step 17 merges the regulatory playbook with the SLA-credit workflow.
- Metric "paid on-call live in payroll before any rotation pages a human" contradicts step 2, where interim duty from day 7 is paid retroactively.
- Keeps the "zero exceptions" over-claim and a single month-5 control test.
- Status page auto-subscribes all 2,100 customers with no consent or contract caveat.
The vote, confronted with the analyst
the analyst agrees with the vote
The vote matches my first choice, and largely for the same reasons I gave: change intelligence against the 3h10 MTTM, Support/AMs as an instrumented detection tier against 40% customer-first detection, the concurrency doctrine, and the paid Duty Triage Desk with published arithmetic answering the pager objection. Voter 1 adds a merit I weighed less — the SLA-credit incentive guardrail — and voters 3 and 5 lean partly on step count ("35 steps", "most comprehensive"), which hints at a length bias, since none of them flags P1's single control test, its self-contradictory "zero exceptions" claim, or P2's superior discipline in defining a page and noise before promising to cut them. Voter 2's case for P2 is exactly the runner-up argument I made (feasible paid domain-based coverage, suppression safeguards, evidence-producing rollout), so the top two are ordered as I had them.
No voter engaged with P3/P4/P5, so the vote says nothing about the derivative-merge problem I penalised in P3 and P5.
Is the analyst's first choice better than the initial proposals? better than every initial proposal
- Gained an operating floor: every final plan protects the company from day 7 (interim commander roster, one declaration path, Support authorised to declare, 53 legacy actions triaged) instead of leaving six-plus weeks of design unguarded — P2's round-0 idea that spread to all five.
- Gained staffing realism: "28 teams on 24x7" was replaced everywhere by a ~30-person command corps plus 10–12 domain rotations plus business-hours Layer C, with published arithmetic and a ban on two-person rotations — the direct, defensible answer to the pager pushback.
- Gained payments correctness: mitigate-vs-resolve semantics, reconciliation and controlled backlog drain before closure, split-brain/replay/duplication guards, and preservation of dual control during ledger recovery.
- Gained proof rather than assertion on detection: the 31-incident replay test with a named detector and expected detection minute became the field's standard evidence that MTTD work actually worked.
- Gained safety rails on noise cutting: shadow mode for seven days, never silently disable, compensating detection before suppression, monthly missed-detection reviews — so the 3,400 → 500 target cannot be met by going blind.
- Gained honest instrumentation: medians and p90, detection source on every incident, and monthly reconciliation of incident counts against support tickets, credits and status-page history to catch under-reporting.
- Gained enforceable money and law: dollar pay bands with a payroll hard gate, FLSA/New York review, a reportability checkpoint recorded even when "not reportable", and an automated SLA-credit workflow tied to root-cause families.
- Lost differentiation: P3 and P5 converged into near-copies of P1's earlier drafts, so five agents produced roughly three distinct plans.
- Lost some good ideas without a decision: P2's numeric declaration thresholds (>10% failures = SEV1), its caution against merging small teams into domains for schedule convenience, and P3's explicit per-page pay table were all dropped rather than argued out.
- Carried an over-claim to the end: four of five still promise "SOC 2 passes with zero exceptions"; only P2 corrected it to "no unresolved material exception", which is what an auditor can actually deliver.
- Grew heavier: step counts rose from 13–30 to 26–35, with several plans creating mega-steps (execution doctrine plus readiness bar plus playbooks plus ledger architecture in one) that are harder to gate than the pieces they replaced.
| Initial proposal | Verdict | Why | How |
|---|---|---|---|
| Proposal 1 |
better | Same author lineage, but the final plan closes the initial version's two structural holes — six unprotected weeks of design and an unstaffable coverage model — and adds payments-correctness and audit machinery it lacked. |
Adds the 7-day operating floor with an interim commander roster paid retroactively (step 2); replaces "14–18 teams at 24x7" with 10–12 domain rotations, a ~30-person command corps, a Duty Triage Desk and published staffing arithmetic (step 9); adds the mitigate-vs-resolve lifecycle with reconciliation before closure and split-brain/duplication guards (steps 7, 16); adds the 31-incident replay test and replay coverage as a leading metric (step 12); adds change intelligence (13), Support-as-detection (20), concurrency (8), week-one legal/evidence scoping including privilege and legal hold (4), anti-gaming reconciliation (30) and the warning that credits may rise before they fall (22); metrics are now staged by day 90 / month 6 / month 12 rather than single end-states. |
| Proposal 2 |
better | It keeps everything the initial P2 got right — day-one interim controls, ledger dual control, shadow-mode alerts, reconciliation before resolution, triage of the 53 open actions — and supplies the delivery machinery that version simply did not have. |
Initial P2 had no pilot, no wave plan with gates, no exercise programme, no readiness bar, no policy publication and no risk register; the final plan has all six (steps 28, 31, 26, 17, 27, 34). It also publishes actual pay bands and a payroll hard gate rather than deferring amounts to HR, adds the detection replay test, the noise burn-down campaign with a leaderboard, the automated SLA-credit workflow and change-correlation tooling. |
| Proposal 3 |
better | Initial P3 was internally inconsistent and structurally fragile; the final plan is coherent and gated. |
Removes the contradiction between a 100-alerts-per-team budget and a <600/month target by using a two-pages-per-responder-per-week budget with auto-quarantine (step 11); drops the unrealistic "56 certified ICs and all 260 engineers trained in 14 weeks" for a certification path with shadow shifts and a 12-month renewal (step 25); replaces the single step-12 rollout that depended on all eleven prior steps with a pilot exit gate plus four risk-ordered waves signed off by directors (steps 28, 31); adds the 7-day floor, the risk register, the reportability checkpoint recorded even when "not reportable", and drops the unanalysed outsourced overnight triage vendor in favour of a costed internal triage desk. |
| Proposal 4 |
better | Initial P4 had the best time-boxing and layer model but thin content on money, law, learning and concurrency, and no protection during its first ten weeks. |
Adds a day-7 operating floor instead of leaving the company uncovered until the ~week-10 pilot; adds the regulatory obligation matrix with a one-hour reportability checkpoint recorded even when not reportable (step 21), the automated SLA-credit and true-cost model (22), the Incident Review Board with verified action closure and 15% reserved capacity (23, 24), the replay test (12), concurrency doctrine (8), change intelligence (13) and a monthly-reviewed risk register (34); keeps P4's strengths — week-one audit clock, page-load rules, pause-on-red-pulse, wording freeze for the observation window. |
| Proposal 5 |
better | Initial P5 was the thinnest and contained two unworkable design choices. |
Replaces the 5-minute SEV1 status-page rule with a realistic 15-minute clock plus templates and a named Comms Lead (step 19); replaces "follow-the-sun between the two AWS regions" — which creates no time-zone coverage for a single New York organisation — with a command corps, domain rotations, a paid triage desk and a written cost-and-date evaluation of a Lisbon/APAC cell as a year-two option (step 9); adds dated milestones throughout, a wave plan with gates, a page budget, alert shadow-mode, the risk register, the legal/regulatory playbook and the anti-gaming reconciliation that the initial plan had none of. |
Evaluation of the process
- Convergence: genuine in round 0, imitation thereafter. Five agents independently landing on the same skeleton (exec mandate → baseline → severity → separated roles → paid on-call → one pager → synthetics → timed comms → postmortems → pilot → waves → audit evidence) is real evidence the skeleton is right. What followed was not argument but transcription: P3 imported 21 of 32 steps from P1's round-1 plan; P5 copied P1's round-2 plan step for step; by round 3 the success-metric lists of P1, P3, P4 and P5 are near-verbatim (same "$1,000 primary week", "2 out-of-hours pages per responder", "30 certified commanders and 18 comms leads", "10–12 domain rotations", "15% reserved capacity"). Two exceptions were earned: P4 improved P1's Duty Triage Desk by carving out ledger/payment-halt/security pages so triage never sits on a crisis, and P2 generalised P4's integrity flag into five modifiers (FINANCIAL-INTEGRITY, SECURITY, PRIVACY, REGULATORY, VENDOR).
- Almost no criticism; disagreements were resolved by fiat, not by evidence. The one real design argument — SEV0 tier vs. four levels vs. severity flags vs. modifiers — ended with three different unreconciled answers and nobody testing them against the 31 historical incidents that every plan claims to have re-coded. P2's warning against merging small teams into domains purely for schedule convenience was ignored rather than rebutted; all five consolidated into 10–12 domains anyway. P1's Duty Triage Desk was copied by P4 and P5 and refused by P2 and P3 — but the refusal was never argued, and the desk is still unsized and unbudgeted in P1 and P5 (only P4 prices it at $400–800/week). Nobody attacked P1's 11-dependency join at step 27, the growth to 35 steps, or the arithmetic of a three-person programme office delivering all of this in 24 weeks.
- Premises were barely questioned, and the best challenges did not spread. Three genuine ones: P1 step 22 warns credits may rise before they fall because better detection surfaces previously unbilled outages; P2 refuses "zero exceptions" and targets "no unresolved material exception"; P4 does the staffing arithmetic showing 260 engineers cannot sustain 28 night rotas. Untouched premises: whether 99.95% is achievable at all on a single shared PostgreSQL ledger (relegated to a parallel "blast-radius workstream" in every plan); whether the business case survives the 15% capacity reservation (~39 FTE of 260, far larger than the $1.3M in credits, and never costed by anyone); whether hiring a dedicated 10–15 person SRE/NOC cohort beats putting ~100 of 260 engineers on nights; and the flat contradiction between "you are paged only for services your team owns" and Layer B consolidation of 28 teams into 10–12 shared domains — only P2 (step 8, "explicit support acceptance") resolves it, and nobody flagged it in the others.
- Losses: options, precision and independent judgement. Dropped without argument: P3's managed overnight triage vendor (a live option for a single-timezone company); P2's numeric declaration thresholds (>10% payment failures for 5 min = SEV1), which were the only thing making severity a genuine lookup rather than a judgement call; P5's vendor shortlist; P3's fully derived $420K compensation budget, replaced by an undefended $0.9–1.3M. Regressions propagated as readily as improvements: P1 cut independent control testing from months 4 and 6 to month 5 alone, and P4 and P5 copied the weaker version. Unique good ideas failed to spread because copying flowed one way (P1 → P3/P5): only P1 has change intelligence/deployment safety (step 13) and Support as an instrumented detection tier (step 20), and only P1 and P2 plan for concurrent incidents at all. Most of all, the field lost the ability to discriminate: the vote was between one plan wearing five costumes.
- Two of five agents became pure copyists. P3 and P5 contributed no idea the others lacked from round 2 onward, wasting 40% of the parallelism and inflating apparent consensus.
- Success metrics were copied verbatim across plans, so they stopped being evidence of design quality and became boilerplate — and the boilerplate carried an overclaim ("SOC 2 Type II incident-response controls pass with zero exceptions" in P1, P3, P4, P5) that the company does not control. Only P2 phrased it defensibly.
- Step inflation rewarded length over executability. P1 29→32→35, P5 14→26→30, P3 13→27→29. Nothing forced removal, so P1 now carries alert quality (11), a separate noise burn-down (29) and a per-wave noise quota (31) covering the same work, and comms split across steps 18/19/20.
- Big dependency joins went unchallenged. P1 step 27 (Policy v1) has 11 predecessors, P5 step 22 has 11, P3 step 22 has 10 — a single-point schedule risk in the middle of the plan that no reviewer named.
- No plan was ever costed end-to-end. Every plan reserves 10–15% of 260 engineers plus $0.9–1.3M compensation plus $150–250k tooling plus a three-person office, and every plan anchors the case on $1.3M of credits. Nobody did the subtraction.
- Feasibility of the timeline was asserted, never tested. Design in 6 weeks for 180 services, 30 certified commanders, synthetics in both regions, six-tool migration, all 28 teams by week 24. P2's alternative (design weeks 1–4, rollout by week 18) was never debated on the merits, only noted as different.
- The severity design argument was decided without data. Every plan claims to re-code the 31 incidents and to publish 12 worked examples from them, yet nobody used those incidents to show which severity model would have classified them correctly.
- Contradiction survived all rounds unexamined: the fairness promise "paged only for what your team owns" versus Layer B domain rotations that necessarily merge several of the 28 teams. Only P2 wrote the reconciliation (formally accepted support agreements, training, access).
- Regressions propagated as fast as improvements — the cut from two independent control tests (months 4 and 6) to one (month 5) spread from P1 to P4 and P5 with no discussion.
- The final vote had almost nothing to discriminate on, since the five plans share their architecture, numbers, metrics and most step text; preference collapsed onto length and two or three unique steps.
- Assign adversarial roles from round 1. One agent must produce a minimum-viable plan (≤12 steps, must still pass the audit); one must attack cost and feasibility; one must defend a structurally different coverage model (managed overnight NOC, or hiring a dedicated 12-person SRE first line instead of putting 100 engineers on nights). Forbid these agents from converging.
- Require a written critique artefact each round: every agent names three specific flaws by proposal and step number, states what it rejected from others and why, and those rejection reasons are carried into the final round as a visible dissent register. Dropping an option silently should be disallowed.
- Ban shared success metrics. Each plan's metrics must be derivable from its own design choices, capped at ~10, each with a falsification condition and an owner. Penalise any metric the organisation cannot control ("audit passes with zero exceptions").
- Impose a step budget with a removal quota: to add a step you must delete or merge one, and any step with more than four dependencies must be split or justified. Reward net simplification in scoring.
- Require one costing table per plan: rotation count × stipend, triage-desk headcount and pay, tooling, programme office, and the opportunity cost of reserved capacity in FTE — set against $1.3M credits plus the quantified true incident cost. No plan advances without it.
- Force premise challenges. Each agent must list two brief assumptions that may be wrong and say what it would do if they are (e.g., 99.95% is unachievable on a shared ledger without re-architecture; credits will rise before they fall; pager pay does not fix attrition).
- Run a resource-feasibility check as a gate: a week-by-week load plan for the programme office and a named critical path, plus a written statement of what gets cut first if the mid-rollout SEV1 contingency fires.
- Add two evaluator personas before the vote — a SOC 2 auditor and a sceptical senior engineer from one of the 16 teams with no on-call — and require every plan to answer their written objections in its final round.
- Score the final round on discriminating dimensions separately (staffing arithmetic, cost, severity calibration evidence, contingency handling, readability under pressure) rather than an overall preference, so near-identical plans are forced to state trade-offs.
- Detect and penalise derivativeness automatically: measure step-level overlap with other proposals and require any agent above a threshold to either justify each copied step or be replaced with a fresh brief.
Convergence: steps changed per round
- opus5_refine_1 opus5 · anthropic/claude-opus-5
- gpt5.6-sol_refine_2 gpt5.6-sol · openai/gpt-5.6-sol
- qwen3.8-max_refine_3 qwen3.8-max · alibaba/qwen3.8-max
- grok4.6_refine_4 grok4.6 · xai/grok-4.6
- deepseek-v4-pro_refine_5 deepseek-v4-pro · deepseek/deepseek-v4-pro
- mean of the agents
| Steps kept, added and removed | Round 1 | Round 2 | Round 3 |
|---|---|---|---|
| opus5_refine_1 |
18 | 28 | 31 |
| gpt5.6-sol_refine_2 |
7 | 17 | 18 |
| qwen3.8-max_refine_3 |
2 | 16 | 18 |
| grok4.6_refine_4 |
4 | 7 | 18 |
| deepseek-v4-pro_refine_5 |
6 | 13 | 15 |
Contributions of each agent
- opus5_refine_1 opus5 · anthropic/claude-opus-5 · adaptive thinking, effort high
- gpt5.6-sol_refine_2 gpt5.6-sol · openai/gpt-5.6-sol · reasoning effort high
- qwen3.8-max_refine_3 qwen3.8-max · alibaba/qwen3.8-max · thinking on, budget 16.0k tokens
- grok4.6_refine_4 grok4.6 · xai/grok-4.6 · reasoning effort high
- deepseek-v4-pro_refine_5 deepseek-v4-pro · deepseek/deepseek-v4-pro · thinking on, effort high
| Agent | Steps of the selected plan it introduced | Steps introduced | Copied by others | Survived to the final round | Ideas taken from it | Ideas rejected | Declared adopted | Declared rejected | Votes received |
|---|---|---|---|---|---|---|---|---|---|
| opus5_refine_1 selected plan |
31 | 47 | 102 | 33 | 35 | 17 | 4 | ||
| grok4.6_refine_4 |
3 | 50 | 36 | 17 | 15 | 3 | 0 | ||
| qwen3.8-max_refine_3 |
1 | 22 | 10 | 1 | 1 | 3 | 0 | ||
| gpt5.6-sol_refine_2 |
0 | 39 | 20 | 16 | 24 | 16 | 1 | ||
| deepseek-v4-pro_refine_5 |
0 | 24 | 1 | 0 | 0 | 3 | 0 |
Efficiency of each agent
Cost per step
Steps per minute of model time
Timeline
Costs
Two separate things were paid for in this run:
- The planning process itself ("Deliberation" below): the 25 LLM calls that produced the plan — 5 agents drafting and refining over 4 rounds, then 5 voters. This is the cost of obtaining the plan. Running
slow-thinkercosts exactly this. - The optional evaluation ("Analysis" below, on yellow like everything the analysis adds): the 8 calls made afterwards by
slow-thinker-report --analyzeso that a reviewing model explains how the proposals evolved and judges the outcome. It does not change the plan and is only paid if you ask for it.
6.46
6.49
12.95
Deliberation (the planning process), by model
| Model | Thinking | Calls | Input tokens | Output tokens | Reasoning | Model time | Cost |
|---|---|---|---|---|---|---|---|
| opus5 · anthropic/claude-opus-5 | adaptive thinking, effort high | 5 | 256.1k | 77.1k | 12.3k | 14 min 16 s | 3.21 |
| gpt5.6-sol · openai/gpt-5.6-sol | reasoning effort high | 5 | 157.9k | 51.7k | 20.7k | 12 min 35 s | 1.66 |
| qwen3.8-max · alibaba/qwen3.8-max | thinking on, budget 16.0k tokens | 5 | 169.2k | 45.0k | 8.1k | 14 min 55 s | 0.61 |
| grok4.6 · xai/grok-4.6 | reasoning effort high | 5 | 163.3k | 32.2k | 42.3k | 18 min 25 s | 0.52 |
| deepseek-v4-pro · deepseek/deepseek-v4-pro | thinking on, effort high | 5 | 161.1k | 63.4k | 36.5k | 11 min 28 s | 0.46 |
| Total | 25 | 907.6k | 269.2k | 119.9k | 1 h 11 min | 6.46 |
Deliberation (the planning process), by call
| Phase | Call | Model | Input tokens | Output tokens | Reasoning | Time | Cost |
|---|---|---|---|---|---|---|---|
| Round 0 | deepseek-v4-pro_initial_5 | deepseek-v4-pro · deepseek/deepseek-v4-pro | 1.1k | 9.2k | 6.2k | 1 min 52 s | 0.037 |
| Round 0 | gpt5.6-sol_initial_2 | gpt5.6-sol · openai/gpt-5.6-sol | 926 | 13.4k | 6.5k | 3 min 56 s | 0.27 |
| Round 0 | grok4.6_initial_4 | grok4.6 · xai/grok-4.6 | 1.6k | 5.5k | 10.4k | 4 min 23 s | 0.036 |
| Round 0 | opus5_initial_1 | opus5 · anthropic/claude-opus-5 | 1.7k | 16.5k | 3.1k | 3 min 46 s | 0.42 |
| Round 0 | qwen3.8-max_initial_3 | qwen3.8-max · alibaba/qwen3.8-max | 1.1k | 8.1k | 752 | 3 min 1 s | 0.051 |
| Round 1 | deepseek-v4-pro_refine_5 | deepseek-v4-pro · deepseek/deepseek-v4-pro | 30.2k | 14.1k | 8.3k | 2 min 31 s | 0.095 |
| Round 1 | gpt5.6-sol_refine_2 | gpt5.6-sol · openai/gpt-5.6-sol | 29.5k | 12.7k | 5.2k | 2 min 58 s | 0.37 |
| Round 1 | grok4.6_refine_4 | grok4.6 · xai/grok-4.6 | 30.6k | 7.2k | 10.2k | 3 min 53 s | 0.10 |
| Round 1 | opus5_refine_1 | opus5 · anthropic/claude-opus-5 | 47.7k | 18.2k | 2.4k | 3 min 38 s | 0.69 |
| Round 1 | qwen3.8-max_refine_3 | qwen3.8-max · alibaba/qwen3.8-max | 31.6k | 8.4k | 1.1k | 2 min 44 s | 0.11 |
| Round 2 | deepseek-v4-pro_refine_5 | deepseek-v4-pro · deepseek/deepseek-v4-pro | 36.3k | 16.5k | 9.3k | 2 min 41 s | 0.11 |
| Round 2 | gpt5.6-sol_refine_2 | gpt5.6-sol · openai/gpt-5.6-sol | 35.6k | 11.3k | 3.5k | 2 min 30 s | 0.37 |
| Round 2 | grok4.6_refine_4 | grok4.6 · xai/grok-4.6 | 36.8k | 9.3k | 7.3k | 3 min 56 s | 0.13 |
| Round 2 | opus5_refine_1 | opus5 · anthropic/claude-opus-5 | 57.9k | 18.7k | 2.2k | 2 min 59 s | 0.76 |
| Round 2 | qwen3.8-max_refine_3 | qwen3.8-max · alibaba/qwen3.8-max | 38.3k | 12.0k | 1.3k | 3 min 55 s | 0.15 |
| Round 3 | deepseek-v4-pro_refine_5 | deepseek-v4-pro · deepseek/deepseek-v4-pro | 43.3k | 19.3k | 8.5k | 3 min 13 s | 0.13 |
| Round 3 | gpt5.6-sol_refine_2 | gpt5.6-sol · openai/gpt-5.6-sol | 42.5k | 12.1k | 3.3k | 2 min 36 s | 0.41 |
| Round 3 | grok4.6_refine_4 | grok4.6 · xai/grok-4.6 | 43.7k | 10.1k | 7.6k | 4 min 17 s | 0.15 |
| Round 3 | opus5_refine_1 | opus5 · anthropic/claude-opus-5 | 68.9k | 22.4k | 3.7k | 3 min 33 s | 0.90 |
| Round 3 | qwen3.8-max_refine_3 | qwen3.8-max · alibaba/qwen3.8-max | 45.5k | 13.3k | 2.2k | 4 min 8 s | 0.17 |
| Voting | opus5_voter_1 | opus5 · anthropic/claude-opus-5 | 79.9k | 1.3k | 906 | 20 s | 0.43 |
| Voting | gpt5.6-sol_voter_2 | gpt5.6-sol · openai/gpt-5.6-sol | 49.4k | 2.2k | 2.1k | 35 s | 0.24 |
| Voting | qwen3.8-max_voter_3 | qwen3.8-max · alibaba/qwen3.8-max | 52.7k | 3.0k | 2.9k | 1 min 8 s | 0.12 |
| Voting | grok4.6_voter_4 | grok4.6 · xai/grok-4.6 | 50.6k | 135 | 6.9k | 1 min 56 s | 0.10 |
| Voting | deepseek-v4-pro_voter_5 | deepseek-v4-pro · deepseek/deepseek-v4-pro | 50.2k | 4.2k | 4.1k | 1 min 10 s | 0.083 |
Analysis (optional evaluation, not part of the process), by call
| Call | Model | Input tokens | Output tokens | Reasoning | Time | Cost |
|---|---|---|---|---|---|---|
| analysis of round 0 | judge · anthropic/claude-opus-5 | 47.5k | 5.8k | 3.0k | 1 min 16 s | 0.38 |
| analysis of round 1 | judge · anthropic/claude-opus-5 | 108.1k | 25.8k | 16.8k | 5 min 18 s | 1.18 |
| analysis of round 2 | judge · anthropic/claude-opus-5 | 128.6k | 31.0k | 22.4k | 6 min 23 s | 1.42 |
| analysis of round 3 | judge · anthropic/claude-opus-5 | 150.4k | 25.0k | 16.2k | 4 min 42 s | 1.38 |
| analysis of final: outcome | judge · anthropic/claude-opus-5 | 135.6k | 18.1k | 10.4k | 3 min 58 s | 1.13 |
| analysis of final: process | judge · anthropic/claude-opus-5 | 134.6k | 7.5k | 4.4k | 1 min 58 s | 0.86 |
| analysis of vote comparison | judge · anthropic/claude-opus-5 | 3.3k | 372 | 0 | 7 s | 0.026 |
| analysis of brief (added afterwards) | judge · anthropic/claude-opus-5 | 16.2k | 1.3k | 0 | 16 s | 0.11 |
| Total | 724.3k | 114.8k | 73.1k | 24 min 0 s | 6.49 |